How to Reduce OpenAI API Costs — Cache, Route, Trim
API bills are a routing problem before they're a budget problem. Three levers, in order of impact, with the math.
Quick answer: The three OpenAI API cost levers: 1) Model routing — send easy traffic to GPT-5 mini ($0.25/$2 per 1M) instead of GPT-5 ($1.25/$10): an 80/20 split cuts a GPT-5-only bill ~64%. 2) Prompt caching — cached input runs $0.125/M vs $1.25 fresh (−90%) on GPT-5; structure prompts with a stable shared prefix. 3) Output trimming — output tokens cost 8× input; caps and terser instructions directly cut the expensive side.
Lever 1: Route by difficulty (biggest cut)
Most stacks send 100% of traffic to the flagship model when 60–90% of requests are classification, extraction or simple Q&A that mini-tier models handle fine. At an 80/20 split (mini/flagship): 8M input + 2M output on GPT-5 alone costs $30/month; the same split costs $10.80/month — a 64% cut with a fallback rule ('mini answers, flagship verifies').
Lever 2: Cache the stable prefix
- Put system prompts and shared context FIRST — caches match the longest common prefix
- GPT-5 cached input: $0.125/M vs $1.25 fresh (−90%)
- Agents with long tool definitions are the biggest winners — the same prefix re-sends every call
- Keep volatile content (user data) at the END of the prompt
Lever 3: Trim the output side
- Output tokens are 8× input price on GPT-5 — the expensive side of every request
- Set max_tokens caps per endpoint, not globally
- Ask for structured/terse output ('answer in ≤100 words', 'JSON only')
- Kill retry storms: a buggy parser retrying 5× quintuples the bill silently
Verify with your own usage data
Pull /v1/organization/usage (per-model token counts) with an admin key, or paste any usage JSON into our multi-model calculator, and price the before/after of each lever against your real traffic — estimates hide, per-model numbers don't.