Your AI costs are out of control

Short answer

Find the concentration before optimizing anything — AI spend is almost always dominated by one feature, one workload, or a handful of heavy users. Then apply fixes in order: route easy calls to a smaller model, cache stable prompt prefixes, cut oversized retrieved context, cap output length, and put hard spend limits on any agent loop. Untuned pipelines routinely come down 60–80% without quality loss.

4 min readUpdated 2026-09-28Problems We Solve

An unexpected AI bill is almost never uniformly distributed. Before changing a single prompt, find out where the money actually goes — teams routinely optimize the wrong thing because they guessed.

Triage: find the concentration

Break spend down along four axes:

By feature. One feature usually dominates. Frequently it is the one nobody considered expensive — a background summarization job, an enrichment step running on every record, a "helpful" suggestion generated on every page load whether or not anyone reads it.

By user or tenant. Power users and automated integrations can account for a wildly disproportionate share. If you are selling a product, check whether any customer is unprofitable at their current plan.

By input vs output tokens. Output typically costs 3–5x input. If output dominates, you have a verbosity problem with a straightforward fix.

By retry and loop behaviour. Failed calls that retry, and agent loops that iterate, can multiply cost invisibly. A single runaway loop can produce a startling bill over a weekend.

If you cannot break spend down this way, that is the first thing to fix. Attribute every call to a feature, a user, and a request ID before optimizing.

The emergency stops

If spend is actively out of control right now:

  1. Set hard spend caps at the provider, and alerts well below them.
  2. Cap agent step counts. An agent without a step budget is an unbounded invoice.
  3. Set max_tokens on every call. Many teams never set it.
  4. Rate-limit per user and per tenant.
  5. Kill or gate the worst offender identified in triage, even temporarily, while you fix it properly.

The fixes, in order of return

1. Route to smaller models

Almost always the biggest lever. Most production traffic is easy — classification, extraction, routing, summarization — and small models handle it well. Teams send everything to a frontier model because the prototype did.

Split the traffic. Use a cheap model for the bulk, escalate to an expensive one only for genuinely hard calls. Typical reduction: 50–80%.

Validate on your eval set rather than assuming quality held.

2. Cache the stable prefix

If your system prompt, tool definitions, or a reference document repeat across calls, prompt caching bills them at a steep discount. Free money, no quality tradeoff. Restructure prompts so the stable part comes first and stays byte-identical.

3. Cut retrieved context

Retrieved context is usually the largest part of the input bill, and most pipelines retrieve far more than needed "just in case." Rerank and pass three to five passages instead of twenty. Accuracy typically improves — see RAG diagnosis.

Strip boilerplate from retrieved documents: headers, footers, navigation, repeated legal text.

4. Stop paying for verbosity

Set max_tokens deliberately. Request structured output rather than prose when a machine consumes it. If you only need a label, ask only for the label — a model asked to explain its reasoning produces many times the expensive tokens.

5. Prune conversation history

In multi-turn systems, resending full history makes cost grow quadratically with turn count. Summarize old turns or maintain a compact structured state. This is the main reason long agent runs get expensive.

6. Use batch endpoints

Anything not needed immediately — nightly jobs, bulk enrichment, backfills — should use batch APIs, which typically carry a large discount. Teams default to synchronous out of habit and pay a premium for latency nobody needs.

7. Deduplicate

Real workloads contain repeated identical requests. An exact-match or semantic cache in front of the model eliminates them. On support and search workloads, 20–40% hit rates are common.

8. Take work out of the model

The cheapest token is the one never sent. Audit your calls for steps with a deterministic correct answer — a lookup, a regex, a database query. A surprising share of model calls are doing work that is not a judgement task.

Stop it recurring

Frequently asked questions

Why did our AI bill suddenly increase?

Most common causes, in order: a feature shipped into a high-traffic path, an agent or retry loop without a cap, retrieved context growing as the corpus grew, conversation history accumulating in multi-turn sessions, or a model upgrade to a more expensive tier. Break spend down by feature and date to find the inflection point.

How much can we realistically cut?

Pipelines that have never been optimized commonly come down 60–80% without quality loss, mostly from model routing, caching, and trimming oversized contexts. Already-tuned systems have less headroom, typically 15–30%.

Does using a cheaper model hurt quality?

On the tasks small models are good at, usually not measurably. The only way to know is an eval set — run both against it and compare. Without one you are guessing, and teams guess wrong in both directions.

Should we self-host to control costs?

Only at high sustained volume, and only counting engineering time honestly. Self-hosting converts variable cost into fixed capacity plus real operational work. At moderate volume, optimizing your API usage almost always beats it.

How do we forecast AI costs for a new feature?

Model five numbers: calls per task, average input tokens, average output tokens, expected tasks per month, and cache hit rate. Multiply through with current provider pricing. Do this before building — it frequently changes the architecture, and it takes an hour.

Guardian Robotics is an AI consultancy.

We build the pipelines, agents, and automation this article describes — for commercial teams and federal agencies alike.