Skip to main content

SaaS RAG security: isolate tenant retrieval and customer data

Design tenant-safe SaaS retrieval-augmented generation with authorization-aware search, isolated vector data, trusted metadata, source citations and leakage tests.

In this guide

How do you secure a multi-tenant RAG system?

A secure SaaS retrieval-augmented generation (RAG) system authorizes a user's access before retrieving customer content, carries that verified scope into search, and checks the returned chunks again before sending them to a model. A vector filter or tenant label helps implement that boundary, but it is not a substitute for application authorization. Treat retrieved text as untrusted input because it can contain private data, stale permissions or instructions aimed at the model.

Derive retrieval scope from authenticated membership

Resolve the caller, active organization, role and permitted records in trusted server code. Build the retrieval filter from that authorization result; never let a browser supply a tenant ID, document ACL, user group or filter expression that broadens access. If the system cannot establish the caller's tenant and permissions, stop before search and generation.

Enforce access at retrieval time and on every returned chunk

Apply tenant and document permissions in the vector query or a separate authorized retrieval layer, then verify each result's tenant, document state and current ACL before constructing model context. AWS documents metadata filtering and identity-based ACL patterns for particular Bedrock integrations; exact capabilities differ by store. A post-generation check is too late because the model has already seen the content.

Choose storage boundaries that match the risk

Compare separate indexes or collections, tenant namespaces, and a shared index with mandatory server-side metadata filters. Consider administrator blast radius, noisy-neighbor effects, encryption and key options, backup and deletion behavior, operational complexity and provider guarantees. A namespace is useful defense in depth, but broad service credentials can still cross namespaces unless the store enforces that boundary.

Tenant-safe RAG access review
Caller and verified tenantDocument permission sourceIndex/filter boundaryPre-generation checksLeakage test owner
Tenant member asking about a private file
Support user with approved access
Former member after access revocation

How should SaaS teams ingest and update RAG documents?

Attach trustworthy provenance and permission metadata

For each chunk retain a stable document ID, tenant ID, source, version, ingestion time, classification and the permission reference needed for retrieval. Validate metadata at ingestion and reject missing or conflicting tenant ownership. Do not infer an ACL from a filename or user-editable tag; map source-system permissions through a controlled, reviewed process.

Synchronize permission changes, deletion and re-indexing

Treat an ACL change, tenant transfer, document deletion or user offboarding as a security event for every derived chunk, cached answer and embedding. Use a reconciliation job that can find stale copies and retry safely; mark content unavailable while updates are incomplete. Test that deletion and revocation reach replicas, backups and downstream indexes according to the product's documented retention rules.

Keep retrieval context narrow and traceable

Retrieve only the few authorized passages needed for the task, cap context size and preserve chunk-to-source references. Avoid sending an entire tenant corpus to the model for convenience. Record which document versions informed an answer so an incident reviewer can reconstruct exposure without logging the full sensitive prompt or response by default.

How do you test RAG tenant isolation and retrieval quality?

Run adversarial cross-tenant retrieval tests

Seed tenants A and B with distinctive canary phrases, then search as each tenant across direct questions, paraphrases, filters, hybrid search, reranking and fallback paths. Assert that no result, citation, snippet, log or generated answer reveals the other tenant's canary. Include anonymous callers, suspended tenants, revoked memberships and malformed metadata.

Test injection and poisoned-source cases separately

Place instruction-like text in an otherwise ordinary authorized document and confirm the model treats it as quoted source data, not as permission to reveal secrets or call a tool. Check ingestion review, source provenance and removal workflow for poisoned or compromised documents. Filtering search results does not neutralize malicious instructions inside a result that the user is allowed to read.

Monitor access decisions without collecting excess content

Track tenant-scoped retrieval counts, denied requests, missing ACL metadata, stale-index lag, unusual broad searches and citation failures. Use opaque IDs and limited retention; keep prompts and retrieved text out of routine logs unless a reviewed incident workflow requires them. Alert on a filter unexpectedly disappearing or a retrieval service switching to an unscoped fallback.

Multi-tenant RAG security FAQs

Does adding tenant_id metadata to vector chunks prevent data leaks?

No. The application must derive the filter from verified identity and permissions, enforce it in the retrieval path, validate returned chunks, and test every fallback and cache path. Metadata is useful only when the query cannot be widened by an untrusted caller.