Skip to main content

How to evaluate RAG retrieval quality in a SaaS application

Measure RAG retrieval separately from generated answers with representative queries, relevant-document labels, ranking metrics, groundedness checks and tenant-aware regression tests.

In this guide

How do you measure RAG retrieval quality?

Evaluate a retrieval-augmented generation (RAG) system in two stages: first test whether retrieval returns the right authorized evidence, then test whether the answer uses that evidence accurately. A fluent answer can hide a retrieval miss, and a relevant passage can still produce an unsupported response. Use a repeatable set of real user questions with expected relevant sources, and measure the retrieval and answer stages separately before changing embeddings, chunking, filters, ranking or prompts.

Build a query set with expected evidence

Collect representative questions, paraphrases, follow-ups, misspellings, Hindi and English variants, ambiguous queries and questions whose answer is not in the corpus. For each answerable question, label the document or passage that should be found and why; for unanswerable cases, specify that retrieval or generation should abstain. Keep examples privacy-screened and version them with the corpus snapshot.

Measure retrieval before judging generated prose

Track whether expected evidence appears in the top k results and how highly it ranks. Precision@k and recall@k help show irrelevant results and missed evidence; reciprocal-rank or normalized discounted cumulative gain can reflect position. Select k and acceptance levels from the product's user need rather than copying a universal benchmark score.

Measure whether the answer is grounded in retrieved evidence

Check whether important claims follow from retrieved passages, whether the response answers the question and whether citations identify the supporting source. Use exact or rule-based checks where possible, and human-calibrated rubrics for nuanced judgments. RAG evaluation frameworks offer useful measures such as context precision, context recall and faithfulness, but each metric has assumptions and should be validated against your own use case.

RAG query and retrieval evaluation worksheet
Query and languageExpected source/passageRetrieved top-k resultRetrieval/answer scoreFailure and next action

What should a SaaS RAG test set cover?

Test corpus freshness, missing evidence and difficult matches

Include recently changed and expired documents, duplicate or conflicting passages, short queries with several plausible meanings, and questions that require combining more than one source. Add no-answer examples so the system is tested for abstention, not rewarded for always producing a response. Record a corpus version, index build and retrieval configuration for each run.

Test authorization and tenant isolation as hard assertions

For each query, verify that every returned passage belongs to the authorized tenant and is visible to the current user. Include guessed IDs, revoked access, changed roles, cross-tenant collisions and cache hits. A relevance score cannot compensate for unauthorized evidence; keep access-control tests separate from quality metrics and make any boundary violation a release blocker.

Stratify results by language, use case and user impact

Break results down for languages, document types, query intents and high-impact workflows. Ask reviewers who understand the domain to assess whether the expected sources are useful and whether answer wording preserves important qualifications. Report sample sizes and failures; aggregate retrieval scores can conceal a weak experience for one language or a small but important query class.

How do you run RAG evaluations after index or model changes?

Compare one changed layer at a time

When diagnosing a regression, hold the other components steady and compare ingestion, chunking, embedding, filters, reranking and generation separately. A release that changes all of them at once makes a score shift hard to explain. Retain before-and-after ranked results and answer samples for cases that move materially.

Use the same query set for repeatable comparisons

Keep a versioned core query set for release-to-release comparisons and a rotating set for new traffic patterns and recent failures. Provider-specific search evaluation tools may impose limits on supported data-store configurations, so confirm those constraints and keep a portable copy of your expected queries and sources.

Turn production failures into reviewed test cases

Use support reports, low ratings and safe operational signals to find missed or irrelevant evidence. Review the case, remove unnecessary personal data and add it to the appropriate set with an expected source or abstention behavior. Do not automatically treat every user correction as truth; verify the evidence and protect the evaluation set from poisoning.

RAG retrieval quality FAQs