Agents cost more than they look like they should, for a reason specific to their architecture: you are not paying for one model call, you are paying for a loop whose length the model decides.
Build cost by tier
| Tier | What it does | Build range |
|---|---|---|
| Prototype | One or two tools, read-only, internal demo | $10k–$30k |
| Narrow production agent | 3–8 read-only tools, real users, logging, evals | $30k–$80k |
| Acting agent | Write access, approval gates, authorization, audit trail | $100k–$300k |
| Multi-agent system | Coordinated agents, shared state, orchestration | $250k+ |
The step change is write access. The moment an agent can change something, you need identity, per-action policy, approval routing, spend limits, an audit trail, and a rollback path. That governance layer is frequently larger than the agent itself — it is the subject of Guardian Agent Defense.
Where the hours go
On a typical acting-agent engagement:
- Tool development and hardening — 30–40%. Each tool needs validation, error handling, idempotency, and clear documentation the model can actually use.
- Evaluation — 15–25%. Agents are non-deterministic; you need behavioural evals over distributions, not assertions.
- Authorization and audit — 15–25% for anything with write access.
- Prompt and loop design — 10–15%.
- Integration and deployment — 10–15%.
Note that the "AI" part is the smallest slice. This surprises buyers consistently.
Running cost, and why it surprises people
A single model call has a predictable cost. An agent run does not, because the agent decides how many calls to make.
A rough model for a moderate agent task: 8–15 model calls per run, with context growing each turn as history accumulates. That commonly works out to 10–40x the token cost of a single call for the same nominal task.
Concretely, an agent handling 5,000 tasks a month at 12 calls per task and a growing context can easily land in the low thousands of dollars monthly on a frontier model — and a fraction of that on a smaller model for the routine steps.
The controls that matter:
- Step budgets. A hard maximum on loop iterations. Non-negotiable.
- Spend caps per run and per period, enforced outside the model.
- Model routing. Use a small model for classification and routing, escalate to a large one only for genuinely hard steps. This is usually the largest single saving.
- Prompt caching on the stable system prompt and tool definitions.
- Context pruning. Summarize old turns instead of resending full history.
- Early termination when the goal is provably met.
An agent without a step budget is an unbounded invoice. We have seen a single runaway loop produce a five-figure bill over a weekend. Set the cap before the first production run, not after the first incident.
The ongoing costs beyond tokens
- Human approval capacity. If consequential actions require sign-off, someone must be available. This is a staffing cost.
- Monitoring and tracing. Agent observability tooling is not free, and you cannot debug an agent without it.
- Maintenance. Tools break when upstream APIs change. Budget 20–30% of build cost annually — higher than for a fixed pipeline, because there are more moving integrations.
When an agent is the wrong purchase
If you can draw the flowchart, build a pipeline instead. It will cost less to build, far less to run, be more reliable, and be dramatically easier to audit. Agents earn their cost only when the path genuinely cannot be determined in advance.
Frequently asked questions
Is it cheaper to build an agent or buy an agent platform?
Platforms get you to a working prototype much faster and are usually the right call for evaluating whether the use case works at all. Building becomes preferable when you need deep integration with proprietary systems, specific authorization behaviour, or control over cost at high volume. Many teams sensibly prototype on a platform and rebuild the parts that matter.
How much do the tools cost versus the agent logic?
Tools are usually the larger share — 30–40% of build effort. An agent is only as good as its tools, and production-grade tools need input validation, idempotency so retries are safe, sensible error messages the model can act on, and permission checks. Teams that treat tools as thin API wrappers ship unreliable agents.
Can we reduce agent cost by using a smaller model?
Often dramatically, but selectively. Smaller models handle classification, extraction, and routing well, and handle multi-step planning poorly. The effective pattern is a small model for most steps with escalation to a larger model for the hard ones — commonly a 50–80% cost reduction with little quality loss.
What is the cheapest way to start?
A read-only agent over one data source with three or four tools, evaluated against a fixed task set. That is a $10k–$30k prototype that tells you whether the reasoning works before you fund authorization and integration.
Why are agent costs so unpredictable?
Because the model chooses the number of steps, and the context grows with each one. Two runs of the same nominal task can differ several-fold in cost. This is why caps and budgets are architectural requirements rather than optimizations.
