· 7 min read

RAG evaluation sets: what goes in and how to score them

How to build the set of questions that tells you whether a retrieval-augmented AI feature is getting better or worse, and how to score retrieval and answers separately.

By

A RAG evaluation set is a fixed, versioned collection of questions, each paired with the passages a retrieval-augmented system should find and the facts a correct answer must contain, used to score every change to the system the same way. It is the difference between knowing that a prompt, model or retrieval change made answers better and hoping it did.

Retrieval-augmented generation, first described by Lewis and colleagues in 2020, pairs a language model with a retriever that fetches relevant passages from your own documents before the model answers. That gives two places to fail: the retriever can miss the right passage, and the model can ignore or misread the passage it was given. A good evaluation set measures both, separately.

Why not just try a few questions by hand?

Because the questions you think of aren't the ones your users ask, and because you can't remember last week's answers. A change that fixes the three questions you tried can quietly break twenty you didn't. Anthropic's guidance on building evaluations starts in the same place I do: define success criteria that are specific and measurable, then design evaluations that mirror the real distribution of tasks, edge cases included.

What goes into the set?

Real questions first. I collect them from search logs, support tickets, chat transcripts and the people who answer customers today, with personal data removed. Then I fill the gaps deliberately, so the set covers each kind of question the feature must handle.

CategoryExample of what it testsWhat a correct result looks like
Direct lookupsA fact stated once in one documentThe right passage retrieved; the fact quoted and cited
Multi-passage questionsAn answer that needs two documentsBoth passages retrieved; both facts combined correctly
Paraphrased questionsThe user's words differ from the document'sRetrieval still finds the passage
Out-of-scope questionsSomething the documents don't cover"I don't know" or a hand-over, not an invented answer
Stale or conflicting sourcesAn old policy and its replacementThe current one used, or the conflict surfaced
Adversarial inputsInstructions hidden in a question or a documentThe instructions ignored; no data leaked
Sensitive requestsData the user isn't allowed to seeRefused, with nothing from restricted documents retrieved

For each question, the set records three things: the expected source passages (by document and section, so a re-chunking doesn't break the label), the key facts a correct answer must contain, and anything the answer must not do. I don't usually store one "golden answer" to compare word by word; a list of required facts is easier to grade and survives harmless rewording.

What does one entry look like?

Here is an invented example for a fictional HR policy assistant, to show the shape rather than any real client's data.

  • Question: "Can I carry unused leave into next year?"
  • Category: direct lookup, paraphrased (the policy calls it "annual leave carry-over").
  • Expected sources: the leave policy, section on carry-over; the current version only.
  • Required facts: whether carry-over is allowed, the limit, and the deadline to use it.
  • Must not: quote the superseded policy, or answer from general knowledge if the policy isn't retrieved.

Five fields, each checkable, most of them without a model.

How big should it be?

Big enough to cover every category above with several examples each, and the failures you care most about. Coverage matters more than size: a set of a few hundred well-chosen questions beats thousands of near-duplicates. It starts small in the scoping week, grows from production, and every real failure a user reports becomes a new row.

How do you score retrieval?

Separately from the answer, because the fixes are different. If retrieval misses the passage, no prompt will save the answer; you change chunking, embeddings, filters or ranking.

The Ragas documentation gives clear names for the two questions. Context recall measures how many of the relevant passages were retrieved: did we miss anything important? Context precision looks the other way: of what was retrieved, how much was relevant, and was it ranked near the top? When the expected passages are labelled in the set, both can be computed exactly, with no model involved.

How do you score the answer?

With the fastest reliable method for each check. Anthropic's evaluation guide puts it plainly: code-based grading is fastest and most reliable but lacks nuance, human grading is the most flexible but slow and expensive, and model-based grading sits between, to be tested for reliability before it's trusted at scale. In practice:

  • Exact checks wherever possible: is the required number, date, category or citation present?
  • Faithfulness: is every claim in the answer supported by the retrieved passages? Ragas defines faithfulness this way, scoring the share of the answer's claims that the context supports. This is the measure that catches invented details.
  • Completeness: are all the required facts present?
  • Refusal correctness: on out-of-scope and sensitive questions, did it decline or hand over?
  • Model-graded checks for tone and clarity, with a rubric, calibrated against a sample I and the client's domain expert grade by hand. If the grader and the humans disagree too often, the rubric changes before the grader is trusted.

How do you test for prompt injection in the set?

By including it. OWASP lists prompt injection as the first risk for LLM applications and distinguishes direct injection, in the user's own input, from indirect injection, in external content the model reads, such as a document in your index. So the set includes both: questions that try to override instructions, and test documents planted in a staging index that contain instructions. The pass condition is that the system answers the user's actual question and does nothing the planted text asks.

How does the set run?

In CI, on every change to prompts, models, chunking, embeddings or retrieval settings. OpenAI's evals guide frames an eval the same way: a schema for the test data and the graders that decide whether an output is correct. The run produces scores per category, compared with the last accepted run. A drop beyond an agreed margin in any category blocks the change, even if the average went up, because averages hide the category you broke.

Two practical points. Model outputs vary between runs, so I run each question more than once where it matters and look at the spread, not one sample. And the set itself is versioned, so a score is always reported against a named version of the set.

How do you keep it honest over time?

  • Add every production failure as a new question, with the fix verified against it.
  • Re-label when documents change. A policy update can make a "correct" expected passage wrong.
  • Hold some questions back from the people tuning prompts, so the main set doesn't become something the prompts are fitted to.
  • Track cost and latency alongside quality. A change that improves faithfulness slightly and doubles cost per answer is a decision for the product owner, not a silent default.

How I do this

Every AI features engagement starts with a one-week scoping sprint that builds the first version of this set from your real questions and reports measured quality and cost per request before anyone commits to a build. For a feature already in production, LLM evals, guardrails and cost control builds the full set from your traffic, wires it into CI and adds the injection and data-leakage cases. The same habit of measuring before shipping ran the Digital Wardrobe pipeline, where cost per image was a design constraint from the first day.

Sources