Inference cost is usually fixable, often dramatically, and almost always without touching output quality. Here are the techniques in order of return on effort.
1. Route to the cheapest model that can do the step
By far the biggest lever. Most production workloads are a mix of easy and hard calls, and teams send all of them to a frontier model because that is what the prototype used.
Split the work. Classification, routing, extraction, formatting, and summarization are handled well by small models. Reserve the expensive model for multi-step reasoning and genuinely ambiguous judgement.
The routing decision itself can be made by a small model, or often by a simple heuristic — document length, detected type, a confidence score from the cheap attempt.
Typical result: 50–80% cost reduction. Validate on your eval set rather than assuming quality held.
2. Cache the stable prefix
If your system prompt, tool definitions, or a reference document repeat across calls, prompt caching bills that prefix at a steep discount. This is free savings with zero quality impact — the only requirement is structuring prompts so the stable part comes first and stays byte-identical.
Order your prompt: system instructions, then tool definitions, then reference material, then the variable user input last.
3. Stop sending context you do not need
Retrieved context is usually the largest part of the input bill, and most pipelines retrieve too much "just in case."
- Rerank and truncate. Retrieve 30 candidates, rerank, pass the best 3–5. Quality typically improves — models attend poorly to long contexts.
- Strip boilerplate from retrieved documents: headers, footers, navigation, repeated legal text.
- Do not resend what has not changed if the provider supports caching it instead.
4. Cap and shape output
Output costs several times input. Set max_tokens deliberately. Ask for structured output rather than prose explanation when a machine is consuming it. Explicitly instruct brevity — "answer in at most two sentences" is a cost control.
Where you only need a label, request only the label. A model asked to explain its reasoning produces many more expensive tokens than one asked for a single word.
5. Prune conversation history
In multi-turn work, resending full history means cost grows quadratically with turn count. Summarize turns older than a window, or keep a running structured state object instead of raw transcript.
For agents this is essential — it is the main reason long agent runs get expensive.
6. Use batch processing
If results are not needed immediately, batch endpoints typically carry a large discount. Any nightly or scheduled workload should use them. Teams default to synchronous APIs out of habit and pay a premium for latency nobody needs.
7. Deduplicate
Real workloads contain repeated identical or near-identical requests. A semantic or exact-match result cache in front of the model eliminates them outright. On support and search workloads, hit rates of 20–40% are common.
8. Fine-tune a small model for a narrow high-volume task
If one task runs at very high volume and a small model almost handles it, fine-tuning that small model can match large-model quality at a fraction of the cost. This is a real engineering investment and only pays back above substantial volume — but at that volume it can pay back enormously.
9. Push work out of the model entirely
The cheapest token is the one you never send. A regex, a lookup table, a database query, or a deterministic rule handles a surprising share of what teams route through a model. Audit your calls: any step with a deterministic correct answer should not be an inference call.
Measure before optimizing
Instrument cost per request, per feature, and per customer before changing anything. Cost is almost always concentrated — a single feature or a handful of heavy users typically dominates the bill. Optimizing the wrong thing is common and wasteful.
Frequently asked questions
What is the single biggest inference saving?
Model routing. Sending every call to a frontier model when most calls are easy is the most common and most expensive mistake in production AI, and fixing it routinely halves the bill or better.
Does using a cheaper model hurt quality?
On the steps small models are good at — classification, extraction, routing, summarization — usually not measurably. The way to know is an eval set: run both models against it and compare scores. Without one, you are guessing in both directions.
Is self-hosting cheaper?
Only at high sustained volume, and only once you count engineering time honestly. You trade variable cost for fixed capacity plus ongoing operational work. It can also be the right answer for data-residency reasons regardless of cost.
How much can we realistically save?
On pipelines that have never been optimized, 60–80% reductions are common and do not require quality compromises. Most of that comes from routing, caching, and cutting oversized contexts.
Does a longer context window cost more?
You pay for the tokens you actually send, not for the window size. But a large window invites sending more, which does cost more — and frequently reduces accuracy. Treat a big context window as capacity, not as an instruction to fill it.
