What is RAG?

Short answer

RAG (retrieval-augmented generation) is a technique where, instead of relying on what a model memorized during training, you search your own content at question time and paste the most relevant passages into the prompt. The model then answers from that supplied context. It is the standard way to make a general-purpose model answer accurately about private, current, or proprietary information — and to cite where the answer came from.

4 min readUpdated 2026-09-28AI Pipelines & Automation

Ask a general model about your company's refund policy and it will invent something plausible. RAG is how you stop that: retrieve the actual policy, put it in the prompt, and require the answer to come from it.

How RAG works

There are two phases, and confusing them is the source of most RAG problems.

Indexing happens ahead of time. You take your documents, split them into passages, convert each passage into a vector — a numeric representation of its meaning — and store those vectors in a searchable index.

Retrieval happens at question time. The user's question is converted into a vector the same way. The system finds the passages whose vectors are closest, pulls the original text, and inserts it into the prompt alongside an instruction like "answer using only the context below, and say so if it is not there."

The model never learned your content. It is reading it, in the moment, like an open-book exam.

Why teams choose RAG

Where RAG actually breaks

In practice, RAG failures are almost never the language model's fault. They are retrieval failures, and they cluster into four causes.

Bad chunking

Splitting a document every 500 characters cuts tables in half, severs headings from their content, and strips the context that made a passage meaningful. A clause that says "this does not apply to Tier 3 customers" is actively harmful when retrieved without the clause it modifies.

Chunk on structure — sections, clauses, rows — not on character count. Keep headings attached to their body.

The question does not look like the answer

A user asks "how long do I have to return something." The document says "the Returns Window shall be thirty (30) calendar days from delivery." Semantically similar, lexically nothing alike. Pure vector search often misses this.

The fix is hybrid retrieval — run both vector search and keyword search, then merge. This is the single highest-return improvement in most RAG systems we audit.

Retrieving enough but ranking badly

Models attend unevenly to long contexts; a correct passage buried tenth in a list is frequently ignored. Retrieve broadly, then rerank with a cross-encoder and pass only the top three to five passages.

The answer spans many documents

"Which of our vendors have contracts expiring this quarter with auto-renewal?" is not a retrieval question, it is a database query. RAG retrieves passages; it does not aggregate across a corpus. Route these to structured queries instead.

A RAG stack that works

  1. Parse with layout preserved. Extract text while keeping headings, tables, and reading order intact. Most RAG quality is won or lost here.
  2. Chunk on structure, 200–800 tokens, with heading context prepended to each chunk.
  3. Embed with a current embedding model; store vectors plus the original text and metadata.
  4. Retrieve hybrid — vector plus keyword, roughly 20–50 candidates.
  5. Rerank down to 3–5 passages.
  6. Generate with an instruction to answer only from context and to state when the answer is absent.
  7. Cite the source passages back to the user.
  8. Evaluate continuously against a fixed question set.
The check most teams skip. Measure retrieval separately from generation. If the right passage was never retrieved, no amount of prompt engineering will fix the answer — and you will waste weeks tuning the wrong stage.

Frequently asked questions

Is RAG better than fine-tuning?

They solve different problems and are frequently combined. RAG supplies knowledge; fine-tuning shapes behaviour, format, and tone. If the model needs to know something, retrieve it. If the model needs to act a certain way consistently, fine-tune it. See RAG vs fine-tuning for the full comparison.

Do I need a vector database for RAG?

Not always. Under roughly 50,000 chunks, vector search extensions in Postgres or even an in-memory index perform well and spare you another piece of infrastructure. Dedicated vector databases earn their place at larger scale, with high query concurrency, or when you need sophisticated metadata filtering.

How accurate is RAG?

Entirely dependent on retrieval quality and your content. Well-built systems over clean, well-structured corpora routinely reach 90–95% answer accuracy on held-out questions. The same architecture over scanned PDFs with no structure can fall below 60%. Accuracy is a property of your pipeline, not of RAG as a technique.

Can RAG eliminate hallucination?

It reduces it substantially but does not eliminate it. A model given correct context still occasionally over-extrapolates. Grounding it further — requiring citations, validating claims against retrieved text, and instructing it to refuse when context is insufficient — closes most of the remaining gap.

How much does a RAG system cost to build?

A focused internal RAG system over an accessible, reasonably clean document set typically runs six to ten weeks. Costs rise sharply when documents need OCR, contain complex tables, or live behind systems with no export path.

Guardian Robotics is an AI consultancy.

We build the pipelines, agents, and automation this article describes — for commercial teams and federal agencies alike.