Three techniques, constantly confused, solving genuinely different problems. Picking wrong costs months.
The one-line distinction
- Prompt engineering changes the instructions. Fastest, cheapest, first thing to try.
- RAG changes what the model knows at answer time by retrieving your content.
- Fine-tuning changes how the model behaves by adjusting its weights on examples.
Knowledge problem, retrieval answer. Behaviour problem, fine-tuning answer. Instruction problem, prompt answer.
Side by side
| Prompt engineering | RAG | Fine-tuning | |
|---|---|---|---|
| Fixes | Format, tone, reasoning approach | Missing or changing knowledge | Consistent behaviour, style, structure |
| Time to first result | Hours | Weeks | Weeks to months |
| Ongoing cost | None beyond tokens | Index + retrieval infra | Retraining on drift |
| Updates | Instant | Re-index a document | Retrain |
| Citations | No | Yes | No |
| Access control | No | Yes, at retrieval | No |
| Needs labelled data | No | No | Yes — typically 500+ examples |
When fine-tuning is genuinely right
Fine-tuning earns its cost in four situations:
- A rigid output format the model keeps drifting from, especially structured extraction where you need the same schema every time across millions of calls.
- A specialized style or voice — a clinical register, a legal drafting convention, a house tone that prompts approximate but never quite nail.
- Cost reduction at scale. A fine-tuned small model matching a large model's quality on one narrow task can cut inference cost by an order of magnitude. At high volume this alone justifies the project.
- A task with no good words for it — classification against a taxonomy so idiosyncratic that describing it in a prompt takes more tokens than the input.
If you are fine-tuning so the model "knows about our products," stop. That is a retrieval problem. Fine-tuning will produce a model that confidently misremembers your catalogue and cannot cite a source.
Why fine-tuning for knowledge fails
Three reasons, and all three surprise teams.
It does not reliably stick. Facts seen a handful of times during fine-tuning are not durably learned. The model becomes stylistically fluent in your domain while remaining factually unreliable — the worst combination, because the output sounds authoritative.
It cannot cite. A fine-tuned model produces an answer with no provenance. For anything regulated, audited, or legally consequential, that alone disqualifies the approach.
It goes stale immediately. Your pricing changes Tuesday. Retraining takes a week. Re-indexing a document takes seconds.
The order to try things
- Write a better prompt. Include a worked example, state the output schema, say explicitly what to do when uncertain. A surprising share of "we need fine-tuning" turns out to be a prompt with no examples in it.
- Add retrieval if the failure is missing knowledge.
- Improve retrieval — hybrid search and reranking — before concluding RAG does not work. Most failed RAG projects failed at retrieval, not generation.
- Fine-tune only once you have an evaluation set, a clear behavioural gap that prompting cannot close, and enough labelled examples.
Notice that step 4 requires the output of step 1–3: you need the eval set to know whether fine-tuning helped, and building it is most of the work regardless.
Combining them
Mature systems use all three. A fine-tuned small model that reliably emits your schema, fed retrieved context from your document store, driven by a carefully engineered prompt. Each technique doing the job it is actually good at.
Frequently asked questions
How much data do I need to fine-tune?
For behavioural and formatting tasks, useful results often start around 500–1,000 high-quality examples, with diminishing returns after a few thousand. Quality matters far more than volume — 500 carefully curated examples consistently beat 10,000 noisy ones. The labelling effort, not the training run, is the real cost.
Is fine-tuning cheaper than RAG?
Cheaper per request, more expensive to build and maintain. Fine-tuning removes retrieval infrastructure and shrinks prompts, cutting per-call cost. But you pay in labelled data, training runs, evaluation, and retraining whenever behaviour needs to change. At low volume RAG almost always wins on total cost; at very high volume on a stable narrow task, fine-tuning can win decisively.
Can I fine-tune and use RAG together?
Yes, and it is often the strongest configuration. Fine-tune the model to follow your output format and reasoning style reliably, then use RAG to supply current facts. The two operate on different axes and do not conflict.
What about prompt caching — does that change the calculus?
It does. Caching a long stable system prompt substantially reduces the cost penalty of large prompts, which weakens one of the main cost arguments for fine-tuning. If your reason to fine-tune was purely prompt length, re-run the numbers with caching first.
Does fine-tuning make a model smarter?
No. It makes a model more consistent at a specific task. Fine-tuning narrows and sharpens behaviour; it does not add reasoning capability. If a task is beyond the base model's reasoning ability, fine-tuning a smaller model will not close that gap — you need a more capable base model.
