AiCost

How to Reduce OpenAI API Costs — Cache, Route, Trim

API bills are a routing problem before they're a budget problem. Three levers, in order of impact, with the math.

Quick answer: The three OpenAI API cost levers: 1) Model routing — send easy traffic to GPT-5 mini ($0.25/$2 per 1M) instead of GPT-5 ($1.25/$10): an 80/20 split cuts a GPT-5-only bill ~64%. 2) Prompt caching — cached input runs $0.125/M vs $1.25 fresh (−90%) on GPT-5; structure prompts with a stable shared prefix. 3) Output trimming — output tokens cost 8× input; caps and terser instructions directly cut the expensive side.

Lever 1: Route by difficulty (biggest cut)

Most stacks send 100% of traffic to the flagship model when 60–90% of requests are classification, extraction or simple Q&A that mini-tier models handle fine. At an 80/20 split (mini/flagship): 8M input + 2M output on GPT-5 alone costs $30/month; the same split costs $10.80/month — a 64% cut with a fallback rule ('mini answers, flagship verifies').

Lever 2: Cache the stable prefix

  • Put system prompts and shared context FIRST — caches match the longest common prefix
  • GPT-5 cached input: $0.125/M vs $1.25 fresh (−90%)
  • Agents with long tool definitions are the biggest winners — the same prefix re-sends every call
  • Keep volatile content (user data) at the END of the prompt

Lever 3: Trim the output side

  • Output tokens are 8× input price on GPT-5 — the expensive side of every request
  • Set max_tokens caps per endpoint, not globally
  • Ask for structured/terse output ('answer in ≤100 words', 'JSON only')
  • Kill retry storms: a buggy parser retrying 5× quintuples the bill silently

Verify with your own usage data

Pull /v1/organization/usage (per-model token counts) with an admin key, or paste any usage JSON into our multi-model calculator, and price the before/after of each lever against your real traffic — estimates hide, per-model numbers don't.

Frequently Asked Questions

How do I cut my OpenAI bill fast?
Route 80% of traffic to mini tiers (−74% typical), enable prompt caching with stable prefixes (−90% on cached input), cap output tokens (8× price side). Together: 60–90% cuts without touching product quality in most stacks.
Does OpenAI prompt caching save money?
Yes — automatically, when prompts share a leading prefix. Cached input on GPT-5 costs $0.125/M vs $1.25 fresh. Restructure so system prompt + shared context lead and user-specific content trails; caching needs zero code changes beyond prompt order.
Is GPT-5 mini good enough for production?
For classification, extraction, summarization and simple chat: usually yes — benchmark it on your own eval set before assuming. Keep the flagship as a fallback for hard requests and quality-sensitive paths; the routing rule costs one if-statement.

More tools used in this guide

More guides