Newsletter
One email. Every week. Pure signal.
The week in quality engineering — skip an issue, and you'll wish you hadn't.
20K+ engineers already reading
RAG Testing Framework: How to Test Retrieval-Augmented Generation Systems
Oct 6, 2026
RAG testing evaluates a retrieval-augmented generation system in two stages. Retrieval testing checks whether the right chunks are found, using metrics like hit rate, context recall, context precision and MRR. Generation testing checks whether the answer is faithful to that context, relevant and correct, and whether the system refuses when the answer is not in its knowledge base. Tools such as RAGAS, DeepEval and promptfoo automate these checks in CI.

A retrieval-augmented generation (RAG) system can fail in two completely different places. It can fetch the wrong documents and then answer beautifully from them. Or it can fetch exactly the right documents and still make something up. From the outside both look the same: a confident, wrong answer.
That is why RAG testing needs its own framework. This guide gives you one: what to measure in each stage, how to build a test set, which tools to use, and how to run it in CI so a prompt tweak or a re-indexed knowledge base cannot quietly break production.
What is RAG testing?
RAG testing is the practice of evaluating a retrieval-augmented AI system in two parts. Retrieval testing checks whether the system finds the right chunks of your knowledge base for a question. Generation testing checks whether the model’s answer is correct, relevant and supported only by those chunks. An end-to-end test then confirms the two work together for real user questions.

Retrieval metrics: did we find the right context?
Retrieval is tested like search. For each question in your test set you know which chunks contain the answer, so you can score what the retriever returned.
| Metric | Question it answers | Good for |
|---|---|---|
| Hit rate (recall@k) | Was at least one correct chunk in the top k? | Quick health check |
| Context recall | Of all the chunks needed, how many were retrieved? | Multi-part answers |
| Context precision | Of the chunks retrieved, how many were relevant? | Prompt noise and cost |
| MRR (mean reciprocal rank) | How high did the first correct chunk rank? | Reranker quality |
| nDCG | Are the most relevant chunks ranked highest overall? | Graded relevance |
If hit rate is low, no prompt engineering will save the answer. Fix chunking, embeddings or the query first.
Generation metrics: did we answer well?
| Metric | Question it answers | How it is usually scored |
|---|---|---|
| Faithfulness (groundedness) | Is every claim supported by the retrieved context? | LLM judge checks claims against context |
| Answer relevancy | Does the answer address the question asked? | LLM judge or embedding similarity |
| Answer correctness | Does it match the reference answer? | Comparison with a labelled answer |
| Citation accuracy | Do cited sources actually contain the claim? | Automated check plus spot review |
| Refusal behaviour | Does it say “I don’t know” when the context lacks the answer? | Unanswerable questions in the test set |
Faithfulness is the metric that catches hallucinations specific to RAG: answers that sound right but are not in your documents. Track it on every release.

How to build a RAG test set
- Start from real questions. Pull them from support tickets, search logs or user interviews. Fifty well-labelled questions beat five hundred synthetic ones.
- Label the answer and its sources. For each question record the reference answer and the document chunks that contain it.
- Cover question types: single fact, multi-hop (needs two documents), comparison, and “how do I” procedures.
- Add unanswerable questions whose answer is not in the knowledge base. The correct response is a refusal.
- Add adversarial cases: misleading phrasing, prompt injection hidden inside a document, outdated documents next to current ones.
- Version it with the code so every change is evaluated against the same set.
Synthetic generation with an LLM is fine for expanding coverage, but have a person review every generated question and answer before it becomes a test.
Tools for RAG testing
| Tool | What it does well |
|---|---|
| RAGAS | Open-source metrics for RAG: faithfulness, answer relevancy, context precision and recall |
| DeepEval | Pytest-style LLM and RAG evaluations you can run in CI |
| promptfoo | Config-driven test suites and model comparisons, good for regression runs |
| TruLens | Tracing and feedback functions for groundedness and relevance |
| Your own pytest suite | Retrieval metrics are simple to compute yourself once the test set has labelled chunks |
A minimal retrieval check in pytest needs nothing more than your retriever and a labelled test set:
import json, pytest
from myapp.rag import retrieve # your retriever
CASES = json.load(open("tests/rag_cases.json")) # [{"question": ..., "relevant_ids": [...]}]
@pytest.mark.parametrize("case", CASES, ids=lambda c: c["question"][:40])
def test_hit_rate_at_5(case):
ids = [chunk.id for chunk in retrieve(case["question"], k=5)]
assert set(ids) & set(case["relevant_ids"]), f"no relevant chunk in top 5: {ids}"
Running RAG tests in CI
- Run retrieval metrics on every change to chunking, embeddings or the index. They are fast and cheap.
- Run generation metrics on prompt or model changes and on a nightly schedule. LLM-judged metrics cost money, so sample if the set is large.
- Gate releases on thresholds you set from a baseline run, for example hit rate at 5 and faithfulness, and alert on drops rather than chasing a perfect score.
- Re-run the full set whenever the knowledge base is re-indexed. New documents can push old answers out of the top results.
- Pin the judge model version, otherwise scores shift when the judge changes rather than when your system does.
Common RAG testing mistakes
- Testing only end to end, so every failure looks like “the model is wrong”.
- Trusting an LLM judge without spot-checking its verdicts against human labels.
- No unanswerable questions, so the system is never rewarded for refusing.
- A test set that never changes while the knowledge base changes every week.
- Ignoring latency and cost: a reranker that improves precision but doubles response time is a trade-off, not a free win.
RAG is often one component inside a larger agent. To see where it fits, read our guide to AI agent architectures for test automation, and for the wider topic of evaluating model output see LLM testing for QA engineers.
Rate this article
7.9/10 average · 38 ratings
Discussion
Start the conversation
What do you think about this article? Share your experience, ask a question, or add to the discussion.
He’s a builder of communities, a collector of questions, and a relentless challenger of assumptions. While others chase answers, he chases better questions. While others talk about the future of testing, he quietly helps create it.
Frequently asked questions.
What is RAG testing?
RAG testing evaluates a retrieval-augmented generation system by testing retrieval (are the right documents found?) and generation (is the answer faithful, relevant and correct?) separately, then end to end.
What metrics are used to evaluate RAG?
Retrieval metrics include hit rate, context recall, context precision, MRR and nDCG. Generation metrics include faithfulness, answer relevancy, answer correctness, citation accuracy and refusal behaviour.
Related articles

LLM Testing: How QA Teams Test Large Language Model Applications
LLM testing checks that AI features give correct, safe and consistent answers. This guide explains the test…
4 min
How to Evaluate AI Testing Tools: A Scorecard for QA Teams
This guide explains how to evaluate AI testing tools, covering the main categories, a weighted scorecard, a…
3 min