How to evaluate RAG retrieval quality in a SaaS application
Measure RAG retrieval separately from generated answers with representative queries, relevant-document labels, ranking metrics, groundedness checks and tenant-aware regression tests.
In this guide
How do you measure RAG retrieval quality?
Evaluate a retrieval-augmented generation (RAG) system in two stages: first test whether retrieval returns the right authorized evidence, then test whether the answer uses that evidence accurately. A fluent answer can hide a retrieval miss, and a relevant passage can still produce an unsupported response. Use a repeatable set of real user questions with expected relevant sources, and measure the retrieval and answer stages separately before changing embeddings, chunking, filters, ranking or prompts.
Build a query set with expected evidence
Collect representative questions, paraphrases, follow-ups, misspellings, Hindi and English variants, ambiguous queries and questions whose answer is not in the corpus. For each answerable question, label the document or passage that should be found and why; for unanswerable cases, specify that retrieval or generation should abstain. Keep examples privacy-screened and version them with the corpus snapshot.
Measure retrieval before judging generated prose
Track whether expected evidence appears in the top k results and how highly it ranks. Precision@k and recall@k help show irrelevant results and missed evidence; reciprocal-rank or normalized discounted cumulative gain can reflect position. Select k and acceptance levels from the product's user need rather than copying a universal benchmark score.
Measure whether the answer is grounded in retrieved evidence
Check whether important claims follow from retrieved passages, whether the response answers the question and whether citations identify the supporting source. Use exact or rule-based checks where possible, and human-calibrated rubrics for nuanced judgments. RAG evaluation frameworks offer useful measures such as context precision, context recall and faithfulness, but each metric has assumptions and should be validated against your own use case.
| Query and language | Expected source/passage | Retrieved top-k result | Retrieval/answer score | Failure and next action |
|---|---|---|---|---|
What should a SaaS RAG test set cover?
Test corpus freshness, missing evidence and difficult matches
Include recently changed and expired documents, duplicate or conflicting passages, short queries with several plausible meanings, and questions that require combining more than one source. Add no-answer examples so the system is tested for abstention, not rewarded for always producing a response. Record a corpus version, index build and retrieval configuration for each run.
Test authorization and tenant isolation as hard assertions
For each query, verify that every returned passage belongs to the authorized tenant and is visible to the current user. Include guessed IDs, revoked access, changed roles, cross-tenant collisions and cache hits. A relevance score cannot compensate for unauthorized evidence; keep access-control tests separate from quality metrics and make any boundary violation a release blocker.
Stratify results by language, use case and user impact
Break results down for languages, document types, query intents and high-impact workflows. Ask reviewers who understand the domain to assess whether the expected sources are useful and whether answer wording preserves important qualifications. Report sample sizes and failures; aggregate retrieval scores can conceal a weak experience for one language or a small but important query class.
How do you run RAG evaluations after index or model changes?
Compare one changed layer at a time
When diagnosing a regression, hold the other components steady and compare ingestion, chunking, embedding, filters, reranking and generation separately. A release that changes all of them at once makes a score shift hard to explain. Retain before-and-after ranked results and answer samples for cases that move materially.
Use the same query set for repeatable comparisons
Keep a versioned core query set for release-to-release comparisons and a rotating set for new traffic patterns and recent failures. Provider-specific search evaluation tools may impose limits on supported data-store configurations, so confirm those constraints and keep a portable copy of your expected queries and sources.
Turn production failures into reviewed test cases
Use support reports, low ratings and safe operational signals to find missed or irrelevant evidence. Review the case, remove unnecessary personal data and add it to the appropriate set with an expected source or abstention behavior. Do not automatically treat every user correction as truth; verify the evidence and protect the evaluation set from poisoning.
RAG retrieval quality FAQs
What is the difference between retrieval quality and answer quality?
Retrieval quality measures whether the system finds relevant, authorized evidence. Answer quality measures whether the generated response uses that evidence correctly and meets the user's need. Evaluate both stages.
Which metrics are useful for evaluating RAG search?
Precision@k and recall@k measure relevance and coverage in a result set; ranking metrics such as reciprocal rank or nDCG account for result position. Groundedness and answer relevance assess generation. Pick measures that match the task and validate them with reviewed examples.
Is an LLM judge enough to score RAG groundedness?
No. Compare automated scores with human-reviewed examples, inspect disagreements and use exact checks for source identifiers or citation requirements. Treat a judge score as an imperfect signal.
Can a high RAG score prove tenant isolation?
No. Test access control directly with adversarial tenant-boundary cases. A relevance or answer score does not prove that returned documents were authorized.
Related practical guides
Related issue guides
Sources and publication record
Draft prepared 27 September 2026; engineering, security and editorial review pending · Sources checked .
- Evaluate search qualityGoogle Cloud Documentation
- RAGAs: Automated Evaluation of Retrieval Augmented GenerationAssociation for Computational Linguistics Anthology
- Evaluation best practicesOpenAI API documentation
- OpenAI API deprecationsOpenAI API documentation
- Evaluation best practicesOpenAI API documentation
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1)National Institute of Standards and Technology
- AWS SaaS Lens: Testing multi-tenant SaaS reliabilityAmazon Web Services
- LLM04:2025 Data and Model PoisoningOWASP Gen AI Security Project
- OWASP Cheat Sheet: LoggingOWASP Foundation