The instinct when a RAG system answers badly is to rewrite the prompt. That fixes a minority of cases and wastes weeks on the rest, because the model usually answered correctly given what it was handed.
Diagnose before you change anything
Take twenty questions your system gets wrong. For each, answer three questions in order:
1. Does the correct answer exist in the indexed content at all?
Search the source documents manually. It is remarkably common to discover the answer was never in the corpus, or lives in a document that failed to parse, or was in a table that your extraction dropped.
If no: this is a content problem, not a RAG problem. Stop tuning.
2. Was the correct passage retrieved?
Log the retrieved chunks for each question and read them. Was the right passage in the set handed to the model?
If no: retrieval failure. This is the most common case by a wide margin. Go to the retrieval fixes below.
3. Did the model answer correctly from what it was given?
If the passage was there and the answer is still wrong: now you have a generation problem, and prompt work is appropriate.
Instrument this permanently. A system that logs retrieved chunks alongside answers can be debugged in minutes. One that logs only the final answer cannot be debugged at all.
Retrieval fixes, in order of return
Add keyword search alongside vector search
The single highest-return change in most systems. Vector search misses cases where the question and the answer share meaning but no vocabulary — a user asking about "returns" against a document that says "the Refund Window shall be." Run both, merge the results.
Rerank and pass fewer passages
Retrieve broadly — 30 to 50 candidates — then rerank with a cross-encoder and pass only the top three to five. Accuracy usually improves when you pass less, because models attend poorly to long contexts and a correct passage buried tenth is frequently ignored.
Fix your chunking
Chunking on a fixed character count is the most common structural error. It cuts tables in half, separates headings from their content, and strips the context that gave a passage meaning.
Chunk on document structure — sections, clauses, rows. Prepend the heading path to each chunk so a passage carries the context of where it came from. Overlap adjacent chunks slightly so a sentence at a boundary is not orphaned.
Fix your parsing
If your extraction turns a two-column PDF into interleaved nonsense, or loses table structure, everything downstream inherits it. Read your parsed output directly — not the source PDF, the parsed text. Teams are often shocked by what their pipeline actually produced.
Handle defined terms and amendments
In contracts and policies, a clause means whatever the definitions section says it means, and the operative text is the original plus every amendment. A pipeline that retrieves a clause without its definitions, or reads a superseded version, will be confidently wrong.
Route aggregate questions elsewhere
"How many contracts expire this quarter" is not a retrieval question. RAG finds passages; it does not count across a corpus. Detect these and route them to a structured query instead.
Generation fixes, once retrieval is proven good
- Instruct grounding explicitly: answer only from the provided context, and state plainly when the answer is not there.
- Require citation of which passage supports each claim. This measurably reduces extrapolation and makes errors visible.
- Put the question after the context, not before — models attend better to instructions near the end.
- Lower the temperature for factual retrieval tasks.
Then build the eval set
If you are diagnosing failures ad hoc, you will fix one thing and break another without noticing. Build a fixed set of questions with known correct answers and known correct source passages, and measure two numbers separately:
- Retrieval recall — was the right passage retrieved, at all.
- Answer accuracy — was the final answer correct.
Tracking them separately is what tells you which stage to work on. See AI evals.
Frequently asked questions
Why does my RAG system make things up?
Usually because it retrieved nothing relevant and answered anyway. The fix is twofold: improve retrieval so relevant context is actually found, and instruct the model to say it does not know when the context does not contain the answer. A model given no useful context will fill the gap unless explicitly told not to.
Is a better embedding model the answer?
Occasionally, and it is rarely the biggest lever. Hybrid search and reranking almost always deliver more improvement than swapping embedding models, and they are cheaper to implement. Try those first, and measure before and after rather than assuming.
How do I know if my chunking is bad?
Print fifty random chunks and read them. If a chunk starts mid-sentence, contains half a table, or has no indication of which document or section it came from, your chunking is hurting you. This takes twenty minutes and is the most informative diagnostic available.
Should I use a bigger context window and skip retrieval?
Tempting and usually worse. Large contexts cost more, are slower, and accuracy degrades as irrelevant material crowds the window. Good retrieval with five relevant passages typically outperforms dumping fifty pages in. Retrieval also gives you access control and citations, which a large context does not.
How accurate should RAG be?
On a clean, well-structured corpus, 90–95% answer accuracy on held-out questions is a realistic target. Below roughly 70%, something structural is wrong — usually parsing or chunking — and tuning will not close the gap.
