Articles / LLMs & Agents
Evaluate RAG Offline: Retrieval Quality and Evidence Availability
Use a small executable Python harness to measure retrieval and annotated evidence availability before changing a RAG system.

A RAG answer can be fluent and wrong in at least two different ways. The retriever may fail to bring back the evidence, or the evidence may be present and the generator may still add an unsupported claim. A single quality score hides that distinction and sends debugging toward prompt wording when the index is the real problem.
This article defines two offline measurements that require no model call: retrieval recall and availability of human-annotated evidence. They are intentionally simple. Their value is a reproducible baseline that catches regressions before an embedding swap, reranker, chunking change, or prompt rewrite reaches users. For a general RAG pipeline, see How to Build a RAG System with LangChain and Python; this is the evaluation layer it needs.
Make the test cases inspectable
Each case needs a question, one or more relevant chunk IDs, and candidate-answer claims with human-verified supporting chunk IDs. Keep chunk IDs stable across a test run. If a chunking experiment changes them, maintain a mapping or re-annotate deliberately rather than comparing incompatible labels.
from dataclasses import dataclass
@dataclass(frozen=True)
class Case:
question: str
relevant_chunk_ids: set[str]
answer_claims: list[tuple[str, set[str]]]
CASE = Case(
question="How many vacation days carry into next year?",
relevant_chunk_ids={"leave-policy-2026-1", "leave-policy-2026-4"},
answer_claims=[
("Employees may carry five unused days.", {"leave-policy-2026-4"}),
("The policy applies to the 2026 plan year.", {"leave-policy-2026-1"}),
],
)
The two claims are a candidate answer with support annotations, not contradictory expected answers. A reference answer should not be treated as a script the model must mimic. Annotate factual units that matter to the user and have a person verify the source passages that support each unit before using the labels.
Measure retrieval before answer quality
Recall at k is the share of relevant chunks in the first k results. Hit rate at k only asks whether at least one relevant chunk arrived. Mean reciprocal rank rewards putting the first relevant chunk near the top. Cases with no relevant chunks are abstention cases, so the functions return None and those cases should be reported separately instead of averaged into retrieval recall.
def recall_at_k(retrieved: list[str], relevant: set[str], k: int) -> float | None:
if not relevant:
return None
if k <= 0:
return 0.0
return len(set(retrieved[:k]) & relevant) / len(relevant)
def hit_at_k(retrieved: list[str], relevant: set[str], k: int) -> float | None:
if not relevant:
return None
if k <= 0:
return 0.0
return float(bool(set(retrieved[:k]) & relevant))
def reciprocal_rank(retrieved: list[str], relevant: set[str]) -> float:
for position, chunk_id in enumerate(retrieved, start=1):
if chunk_id in relevant:
return 1.0 / position
return 0.0
For relevant_chunk_ids={"a", "b"}, a top-five result containing only a has recall at five of 0.5 and hit rate of 1.0. Both are worth reporting. Recall shows that required evidence is missing; hit rate shows the search found some evidence.
Do not silently increase k until recall looks good. More context can dilute the answer, consume budget, and hide a ranking failure. Compare fixed configurations such as top 3 and top 5, then inspect misses by document type, date, query phrasing, and metadata filter.
Measure annotated evidence availability, not faithfulness
The offline example below uses human-verified claim-to-chunk annotations rather than an LLM-as-judge. It measures whether the retrieved set contains the already-labeled evidence for a candidate claim. It does not establish that a generated answer is true, correctly quoted, or faithful. A model can cite an available source while misrepresenting it. Human review or a calibrated support judge must label the generated claim against the passage before calling that a faithfulness metric.
def claim_evidence_availability(
claims: list[tuple[str, set[str]]], retrieved: list[str]
) -> float:
if not claims:
return 1.0
available = set(retrieved)
supported = sum(bool(sources & available) for _, sources in claims)
return supported / len(claims)
retrieved = ["leave-policy-2026-4", "benefits-2"]
print(recall_at_k(retrieved, CASE.relevant_chunk_ids, k=2))
print(hit_at_k(retrieved, CASE.relevant_chunk_ids, k=2))
print(claim_evidence_availability(CASE.answer_claims, retrieved))
This prints 0.5, 1.0, and 0.5. Retrieval found some answer-bearing evidence, but it did not retrieve evidence for the plan-year candidate claim. This is evidence availability, not a faithfulness verdict. A generated answer still needs claim-level support review; before generating, the system should either retrieve the missing evidence or avoid making that claim.
Add abstention cases and change gates
Include questions the corpus cannot answer. Their expected behavior is abstention, a clarifying question, or a clearly marked limitation. For those cases, retrieval quality alone is insufficient because vector search always returns neighbors. Record whether the system made an unsupported answer and whether it showed sources the user can inspect.
Use a fixed evaluation split for change decisions and a separate development split for tuning. Set a baseline for each metric, then reject a change that materially reduces recall or annotated evidence availability even if one aggregate value improves. The exact threshold depends on harm, traffic mix, and review capacity. Never present an offline score as a production guarantee.
Agent evaluations require the same discipline. Anthropic's guide to agent evals emphasizes task-specific outcomes and reliable graders. For RAG, the task-specific diagnosis is usually: did the evidence arrive, and did the final answer stay within it? Keep those questions separate and regressions become debuggable.
About the author
Rodrigo Arenas is a software architect and machine learning engineer in Medellín, Colombia. He builds products, platforms, and open-source software for AI, Data Engineering, and Machine Learning: creator of Ciaren, sklearn-genetic-opt, and PyWorkforce.
Are you building an AI or data platform?
I design and build critical systems end to end: architecture, data, models, and product. If that is the scale you are working at, let us talk.