Most AI business cases are wrong in the same two ways, and both inflate the number. Fixing them produces a case that survives a CFO's questions.
Error one: assuming full automation
A system that handles 80% of cases does not deliver 80% of the benefit, because the 20% remainder still needs a person, and that person now handles only the hard cases — which take longer than average.
Model three streams explicitly:
- Automated, at near-zero marginal labour cost.
- Reviewed — the model proposed, a human verified. Faster than manual, not free. Typically 30–60% of original handling time.
- Exception — fully manual, and slower than your old average because these are the difficult ones.
Error two: hours saved treated as money saved
If you save 2,000 hours a year and headcount stays flat, you have not saved payroll. You may have gained real value — more volume handled, faster cycle time, better quality, less overtime, reduced attrition — but it is not a cost reduction, and presenting it as one damages credibility.
Be explicit about which you are claiming:
Hard savings. Reduced headcount, avoided hiring, eliminated vendor spend, captured early-payment discounts, recovered overbilling. These hit the P&L.
Soft savings. Faster cycle time, higher throughput at flat cost, better consistency, improved employee experience. Real, but do not claim them as cash.
The strongest business cases lead with one or two hard savings and list soft benefits separately as upside. Cases built entirely on soft benefits tend not to survive budget scrutiny.
The full cost side
Business cases routinely omit three ongoing costs:
- Inference. Recurring and volume-linked. See reducing inference cost.
- Human review. The single most-omitted line. If 20% of volume needs review, that is a staffing requirement with a real number attached.
- Maintenance. 15–25% of build cost annually for prompts, model updates, integration drift, and monitoring.
Plus one-time costs beyond the build: internal engineering time, subject-matter expert time for the eval set, change management, and training.
A worked structure
Take invoice processing at 10,000 invoices per month, 12 minutes each manually, loaded cost $35/hour — a manual baseline of $70,000/month.
Model a realistic post-deployment state:
| Stream | Share | Time each | Monthly cost |
|---|---|---|---|
| Touchless | 65% | 0 min | $0 |
| Reviewed | 25% | 4 min | ~$5,800 |
| Exception | 10% | 18 min | ~$10,500 |
| Inference | — | — | ~$1,200 |
| Total | ~$17,500 |
Gross monthly saving is roughly $52,500. Subtract maintenance — say $150,000 build amortized with 20% annual upkeep, about $15,000/month in year one — and you are near $37,000 monthly net, with payback inside a year.
That is a defensible case. Note it assumes 65% touchless, not 100%, and it charges for review.
What to measure after launch
Commit to measuring the same quantities you forecast:
- Automation rate, split by stream.
- Actual handling time per stream, measured not estimated.
- Inference cost per unit.
- Quality — error rate versus the manual baseline.
- Cycle time end to end.
Forecast-versus-actual on these is what earns budget for the next project.
Frequently asked questions
What is a good payback period for AI automation?
Twelve to eighteen months is a reasonable target for a back-office process automation. Faster is achievable where the manual baseline is expensive and the process is high volume — freight invoice audit and AP are common examples. Anything projecting under six months deserves a hard look at whether costs were omitted.
Should we count quality improvements?
Yes, but quantify them separately and conservatively. If automation reduces error rate, the value is the avoided cost of those errors — rework, penalties, customer credits, churn. That is a real number if you can source it from historical data, and hand-waving if you cannot.
How do we handle the fact that we do not know the automation rate yet?
Model a range and make the pilot's job to determine it. Run the case at pessimistic, expected, and optimistic automation rates. If the pessimistic case still clears your hurdle rate, you have a robust project. If only the optimistic case works, you have a gamble.
Does AI ROI improve over time?
Usually yes, for two reasons: the exception queue becomes training data that raises the automation rate, and the ingestion and normalization layers get reused by later projects at a fraction of the original cost. Second and third use cases on the same foundation are markedly cheaper.
What if leadership wants headcount reduction specifically?
Then say plainly what the system can and cannot support. Automation of 65% of volume does not mean 65% fewer staff, because the remaining work is the harder work. Overpromising headcount reduction is the fastest route to a cancelled programme and a distrustful organization.
