The RAG Pipeline and Its Evaluation
Goal
You build one complete RAG pipeline: split documents into chunks and index them, search with queries, measure retrieval quality with metrics, and filter out answers with no grounding and questions that cannot be answered.
Why it matters
The most common mistake in RAG is not telling retrieval failure from generation failure. When an answer looks strange people fix the prompt first, but if the correct document was never retrieved in the first place, no amount of polishing the prompt helps.
So this lab computes the retrieval metrics first. Recall@3 looks at whether the correct document came in within the top 3, and MRR looks at the rank at which it appeared. If these numbers are low, the place to work on is the retriever, not the generator.
Next you build the grounding judgment and the refusal rule. Not answering when the retrieval score is low is not a missing feature but a safeguard. Saying you do not know is far better than making up a plausible answer on the basis of an irrelevant document.
No model is used. Retrieval and evaluation can all be computed with numpy, and that part decides most of RAG quality.
Steps
The working directory is /root/llm, and the tokenization rule is the same as in the previous lab (lowercase the text, then [가-힣a-z0-9]+).
- Split each
bodyindocson the period (.), strip surrounding whitespace and discard empty pieces, and save the resulting chunks to/root/llm/chunks.tsv. The format isdocs.id<탭>청크번호<탭>본문(docs.id, tab, chunk number, tab, body text), and the chunk number starts at 0 for each document. Write them in document order. - Save the chunk-level TF-IDF matrix to
/root/llm/chunk_tfidf.npy. The shape is (number of chunks, number of terms), and the terms are the vocabulary collected over all chunks, sorted ascending. N in the idf is the number of chunks, and the formula and normalization are the same as in the previous lab. - For the 8 queries below, save the top 3 chunks to
/root/llm/retrieved.tsvas질의번호<탭>순위<탭>docs.id<탭>청크번호<탭>점수(query number, tab, rank, tab, docs.id, tab, chunk number, tab, score). That is 24 lines in total, the score has six decimal places, and on a tie the smaller document id comes first, then the smaller chunk number.인덱스를 만들었는데 왜 순차 스캔인가(Why is it a sequential scan even though I built an index?)복합 인덱스는 컬럼 순서를 어떻게 정하나(How do you decide the column order of a composite index?)격리 수준을 올려도 안 잡히는 이상 현상은 무엇인가(What anomaly is not caught even if you raise the isolation level?)페이지 폴트가 잦으면 왜 느려지나(Why does frequent page faulting make things slow?)가상 주소를 물리 주소로 어떻게 바꾸나(How is a virtual address turned into a physical address?)연결 거부와 타임아웃은 무엇이 다른가(How do a connection refusal and a timeout differ?)큰 요청만 멈추는 이유는 무엇인가(Why do only large requests stall?)GPU 에서 워프란 무엇인가(What is a warp on a GPU?)
- The correct documents for the queries above, in order, are
24, 25, 27, 12, 11, 19, 17, 6. Compute Recall@3 and MRR and save them in two lines to/root/llm/metrics.tsv. The first column isrecall_at_3andmrr, and the second column is a value with three decimal places. Recall@3 counts 1 when the correct document is within the top 3 and divides by the number of queries, and MRR is the average of the reciprocal of the rank at which the correct answer first appeared (0 if it is absent). - Save the bodies of the top 3 chunks for query 1, one per line in rank order, to
/root/llm/context.txt. - Judge whether the 3 candidate answers below are grounded in the context from step 5, and save the result to
/root/llm/grounded.tsvas번호<탭>겹침비율<탭>판정(number, tab, overlap ratio, tab, verdict). The overlap ratio is the fraction, among the answer tokens collected without duplicates, that are in the context token set, with three decimal places. If it is 0.6 or more, the verdict isgrounded, otherwisehallucinated.인덱스는 값으로 정렬된 구조라 컬럼에 함수를 씌우면 그 정렬이 쓸모없어진다(An index is a structure sorted by value, so wrapping the column in a function makes that ordering useless.)옵티마이저가 인덱스를 쓰지 않는 것은 대개 옳은 판단이다(The optimizer not using an index is usually the right call.)인덱스를 만들면 어떤 쿼리든 언제나 세 배 빨라진다(If you create an index, any query always becomes three times faster.)
- For the 3 queries that are not in the corpus, save the top score and the verdict to
/root/llm/refusal.tsvas번호<탭>최고점수<탭>판정(number, tab, top score, tab, verdict). The score has six decimal places, and if it is below 0.15 the verdict isrefuse, otherwiseanswer.김치찌개 맛있게 끓이는 법(How to make kimchi stew taste good)프리미어리그 이번 주 경기 일정(This week's Premier League match schedule)제주도 항공권 최저가 예약(Booking the lowest-priced flight to Jeju Island)
- Leave a summary in three lines in
/root/llm/rag_report.tsv. The format ischunks<탭>청크수,queries<탭>질의수andrecall_at_3<탭>값(name, tab, value; the placeholders stand for the number of chunks, the number of queries and the value).
Notes
- A chunk index has many more rows than a document index even when the number of terms is similar. Check the shape first.
- The score is the dot product of normalized vectors, that is, cosine similarity.
- Common mistake 1: if you compute idf based on the number of documents, all the values go off. Use the number of chunks.
- Common mistake 2: the denominator of the overlap ratio is the number of answer tokens without duplicates.
Split documents into sentence-level chunks
Split on periods, strip surrounding whitespace and discard empty pieces. The chunk number starts again from 0 for each document.
Build an index at the chunk level
When you compute idf, note that N is the number of chunks, not the number of documents.
Retrieve the top 3 chunks for 8 queries
Build the query vectors the same way. Follow the tie-handling rule.
Compute Recall@3 and MRR
It is 1 if the correct document is within the top 3 and 0 if not. MRR averages the reciprocal of the rank at which the correct answer appeared.
Assemble the context
Join the retrieval results for query 1 one line at a time in rank order.
Judge whether the answers are grounded in the context
Collect the answer tokens without duplicates and count the fraction that are in the context. Judge by whether it exceeds the threshold.
Refuse questions outside the corpus
If the top score is below the threshold, mark it as not answering. The score itself must also be recorded.
Write the summary report
Leave three things, the number of chunks, the number of queries and Recall@3, as name and value pairs.