Where we left off
Every earlier lesson judged correction by eyeballing one or two questions. This lesson replaces eyeballing with naive_rag Lesson 17's precision@k, scored twice, once with plain naive retrieval, once with this course's grade-and-filter correction, so the improvement is a number, not an impression. It also adds a check naive_rag Lesson 17 never needed: does the generated answer actually get better, not just the retrieval metric.
This lesson assumes you've read naive_rag's Lesson 17 README, specifically its "Why this doesn't generalize (yet)" section, rather than re-deriving it here. Everything that section says about five questions being too few to trust the exact number, and about the tune/eval contamination trap, applies here without modification.
The labeled set
LABELED_QUESTIONS = [ ("Project Aurora's Raspberry Pi writes sensor readings...", "bookshelf.md"), ("How often does the wind speed sensor need re-oiling?", "weather-station.md"), ("What's the cold ferment time for the pizza dough?", "pizza-dough.md"), ("What piece is being practiced on the cello?", "cello-practice.md"), ("Which vegetables grow in the second raised bed?", "garden.md"),]Four of these five are unambiguous, naive_rag-style questions, naive retrieval already handles them correctly, they exist so the score isn't measuring a single question's fluke. The first is this course's own running example, the one question this whole labeled set actually depends on to demonstrate anything: naive top-1 retrieval confidently returns the wrong source for it (Lesson 2).
The code, piece by piece
def precision_before(store, query_vectors) -> float: for ...: top1 = retrieve_by_vector(qv, store, k=1)[0] hit = top1["source"] == expected_sourcePlain naive retrieval, k=1, no grading, the naive_rag baseline.
def precision_after(store, query_vectors) -> float: for ...: top3 = retrieve_by_vector(qv, store, k=3) relevant_sources = [c["source"] for c in top3 if grade_chunk(question, c["text"]) == "relevant"] hit = expected_source in relevant_sourcesCorrection, exactly Lessons 4-5's pipeline: over-fetch, grade every candidate, keep what's relevant, check whether the expected source survived. Nothing new, this lesson just runs it across five questions instead of one and scores the result.
naive_answer = generate_answer(hard_question, naive_top1)corrected_answer = generate_answer(hard_question, corrected_chunks)The answer-quality check: generate from naive retrieval's context and from correction's context, on the one question where they differ, and compare the actual text. Precision@k alone doesn't prove this course's Lesson 1 premise, "retrieval was more precise" isn't the same claim as "the user got a better answer," this check closes that gap directly, per this course's To-Do List commitment.
Checkpoint
- Precision@k, scored before and after correction, turns "did grading and filtering help?" into a number you can compare, the same idea
naive_ragLesson 17 introduced, applied here to a before/after comparison instead of a single snapshot. - A retrieval-precision improvement doesn't automatically mean a generation-quality improvement, Lesson 16 already showed grading can share the generator's blind spots, this lesson's answer-quality check is what actually confirms the user-facing outcome improved, not just the metric.
- Every caveat in
naive_ragLesson 17's "Why this doesn't generalize (yet)" section still applies: five questions is enough to demonstrate the mechanism, not enough to trust the exact number, and this course's own Lesson 14 rewrite-strategy heuristic was deliberately tuned against a different signal than this lesson's score, not against it.
If anything here still feels unclear, ask before moving to Lesson 18.