LLM API pricing compared

Short answer

LLM APIs price separately for input and output tokens, with output typically costing three to five times more than input. Because list prices change frequently and vary by model tier, the durable skill is not memorizing prices but modelling your own token profile — average input size, average output size, calls per task, and cache hit rate. Those four numbers determine your bill far more than which provider you pick.

4 min readUpdated 2026-09-28Cost, Pricing & ROI

Published per-million-token prices are the least useful number in this decision, because they change often and because two architectures using the same model can differ tenfold in cost. What follows is the method rather than a price table that would be stale within weeks.

How the pricing model works

Input and output are priced separately. Output is consistently more expensive — commonly 3–5x input. This has a direct architectural consequence: verbose responses cost real money, and instructing a model to be concise is a cost control, not a style preference.

Everything in the prompt is input tokens, every time. Your system prompt, tool definitions, retrieved context, and conversation history are re-billed on every single call. A 4,000-token system prompt on 100,000 calls is 400 million input tokens whether or not it changed.

Tiers exist for a reason. Providers offer small, mid, and frontier models with order-of-magnitude price differences. Small models are genuinely capable at classification, extraction, routing, and summarization — the bulk of production workload.

Caching changes the arithmetic. Prompt caching bills repeated prefixes at a steep discount. For any workload with a stable system prompt, this is among the largest available savings and it requires no quality tradeoff.

Batch processing typically carries a substantial discount for work tolerant of delayed completion. If your process is a nightly run, this is free money.

A rough sense of scale: token counts run about 0.75 words per token for English prose. A 500-word document is roughly 650 tokens; a 50-page contract is roughly 25,000.

Model your cost in five numbers

Skip the price comparison until you have these:

  1. Calls per task — 1 for a simple pipeline, 8–15 for an agent.
  2. Average input tokens per call — including system prompt, tools, retrieved context, and history.
  3. Average output tokens per call.
  4. Tasks per month.
  5. Cache hit rate on the stable prefix.

Monthly cost then falls out as: tasks × calls × ((input × input price × (1 − cache discount × hit rate)) + (output × output price)).

Run this before choosing an architecture. It routinely reveals that the expensive part is a large retrieved context being resent on every turn, not the model tier.

The comparisons that actually matter

Rather than headline price, compare on:

Self-hosting

Open-weight models on your own GPUs shift you from variable to fixed cost. The crossover is genuinely high, because you pay for capacity whether or not you use it, plus engineering time for serving, scaling, and upgrades.

Self-hosting tends to win on very high sustained volume, on hard data-residency requirements, or on a narrow fine-tuned task where a small model suffices. It rarely wins on cost alone at moderate volume once engineering time is counted honestly.

Frequently asked questions

Which LLM API is cheapest?

The answer changes frequently enough that any specific claim here would mislead you within a quarter — check current provider pricing pages directly. More usefully: the cheapest option is rarely the cheapest outcome. Measure candidate models on your own eval set and compare cost per successfully completed task, not cost per token.

Why is output more expensive than input?

Output is generated sequentially, one token at a time, so it cannot be parallelized the way reading an input prompt can. It occupies the accelerator for longer per token, and pricing reflects that.

How do I estimate tokens before building?

Take 20 representative real inputs, count their tokens with the provider's tokenizer, and add your system prompt and expected retrieved context. Multiply by expected volume. This takes an hour and is far more reliable than any rule of thumb.

Does prompt caching work for every workload?

It helps whenever a substantial prefix repeats across calls — a stable system prompt, fixed tool definitions, or a shared document being asked about repeatedly. It does not help when every request is entirely distinct. Cache mechanics and minimum sizes vary by provider, so verify against current documentation.

Should we commit to one provider?

Keep an abstraction layer thin enough that switching is a configuration change rather than a rewrite, and maintain an eval set that lets you requalify a new model in a day. Deep coupling to one provider's specific features is a real risk given how fast pricing and capability move.

Guardian Robotics is an AI consultancy.

We build the pipelines, agents, and automation this article describes — for commercial teams and federal agencies alike.