TT Lab
Get started
Learn Learning paths Courses

LLM Engineering

RAG — First Work Out Which of the Three Parts Broke

Continue in TT Lab

In one line

RAG is a combination of three components, the retriever, the reranker and the generator, and when you get a report that "answer quality is bad", the first thing to do is to work out which component failed. For that, every answer must have a record of what was retrieved and what went into the model.

Why this was needed

An internal policy chatbot answered "the refund request deadline is 14 days", but the policy had changed to 7 days last month. The person in charge first puts the sentence "prefer the latest policy" into the prompt. The answer stays the same. Later they find that the new policy document had not even entered the index, and the model was faithfully answering from the documents it received. The place to fix was the index from the start. The most common waste of time in improving RAG looks like this.

Still, the reason for using RAG is clear. A model has only the knowledge from its training time and does not know internal documents or recent information. You can also put knowledge in through fine-tuning, but it is expensive, slow and hard to update. RAG first finds the relevant documents when a question arrives, puts them in the prompt, and has the model answer on the basis of those documents. When knowledge changes, you only swap the documents. But as in the case above, whether "just swap the documents" actually holds has to be checked separately.

How it works

The pipeline splits into three components.

Retriever. It fetches chunks similar to the query, broadly. There is dense retrieval (embeddings) and sparse retrieval (keyword methods such as BM25), and the hybrid that combines the two is the practical standard. They fail in different ways. Embeddings are strong on paraphrase but weak on exact matches such as product codes, and keywords are the opposite. When you combine the two results, the score scales differ and cannot simply be added, so methods that use rank are widely used. Reciprocal rank fusion (RRF) scores each document by the rank it received in each list and adds them up.

RRF(d) = sum over lists of  1 / (k + rank(d))      k is commonly 60

k is a constant that keeps a few top-ranked items from monopolizing the result. Elasticsearch also uses 60 as its default.

Reranker. It re-orders, precisely, the candidates the retriever fetched broadly. It scores by putting the query and the document into the model together, so it is accurate but slow, and you apply it only to the top few dozen.

Generator. It produces the answer on the basis of the fetched context.

When diagnosing a quality problem, you must look at retrieval and generation separately. If the correct document was never retrieved in the first place, no amount of tinkering with the generator helps, and you should look at retrieval metrics. If the correct document is in the context and the answer is still wrong, then it is a generator problem and you should look at faithfulness.

Metric What it measures
Recall@K Did the correct document come in within the top K?
MRR At what rank did the first correct answer appear? (the mean of reciprocal ranks)
Precision@K The fraction of relevant documents among the top K
Faithfulness Is the answer grounded in the context?
Answer Relevancy Does the answer fit the question?

A small calculation shows the difference between the two metrics. Say the correct answers for three queries appeared at rank 1, rank 3, and outside the top 3. Recall@3 is 0.667, since two out of three came in. MRR is 0.444, the average of the reciprocal ranks 1, 1/3 and 0. Recall looks only at "did it come in", so it does not distinguish rank 1 from rank 3, while MRR penalizes that difference. If your system puts only the top few into the context, you need to look at both.

What goes wrong in the field

Pairing symptoms with the broken component speeds up investigation.

A stale index. The documents changed, but the index is the old one. The symptom is "the content we said we fixed does not show up". The case where you deleted a document but its chunks remain in the index and keep being retrieved is the same. Index updates must be one procedure together with changes to the source, and you can only check them if you attach the source's update time to every chunk.

Permission leakage. Documents the user cannot see get retrieved and mix into the answer. Hiding them after generation is too late, because the model has already read and summarized that content. Apply the permission filter at the retrieval stage without exception.

Version conflicts. If the old policy and the new policy are retrieved together, the model picks one of them or blends them. The symptom is that the answer wobbles for the same question. Filter out retired documents before retrieval, or keep only the latest version using the effective date in the metadata.

An answer cut at a boundary. The sentence with the correct answer is split across two chunks, and neither rises to the top. This is a chunking-strategy problem and is covered in the next reading.

It is in the context but the answer is wrong. Only now is it a generator problem. A common case is that there are too many irrelevant chunks, or the correct chunk is buried in the middle of a long context. Reduce the count with reranking and put the most relevant chunk first.

Plausible fabrication. Even for a question outside the corpus, it builds an answer on the most similar document. It is far better that, when the retrieval score is below a threshold, the system does not answer and says so. The threshold is decided with an evaluation set. If it is too high, even questions that could be answered are refused, and if it is too low, hallucination increases. And attaching sources to the answer is not a feature but a safeguard. If users can check the original, the damage of a wrong answer shrinks, and you can observe in operation which chunks are being retrieved wrongly.

How to check

For every answer, leave at least the following in one line: the query, the retrieved chunk ids and scores, the chunk ids actually put into the model, the answer, and the sources the answer cited. Without this record, there is no way to work out which component is the problem when a report comes in.

When you collect failure cases, classify them into three boxes.

gold chunk not in top-K        -> retriever (index, chunking, query)
in top-K but not in context    -> reranker or context budget
in context but answer wrong    -> generator (prompt, ordering, noise)

An evaluation set does not need to be grand. Just attaching the correct documents to a few dozen real user questions lets you compute Recall@K and MRR, and each time you change chunking or weights you can measure the same numbers again and judge whether things got better or worse. Changing prompts or settings without these numbers relies on feel.

What to read next

The reading right after this covers chunking, which decides the unit of retrieval. In the lab after that, you split documents into chunks, index them with TF-IDF and search 8 queries. You compare with the list of correct documents to compute Recall@3 and MRR, assemble the context, judge by vocabulary overlap whether an answer is grounded in the context, and implement a threshold rule that does not answer questions outside the corpus.