Embedding model migration: re-index and switch SaaS search safely
Upgrade vector embeddings with a versioned index, consistent dual writes, idempotent backfill, retrieval evaluations, an atomic traffic switch and a usable rollback window.
In this guide
Why does an embedding model change require a search migration?
An embedding model maps text into a particular vector space. A new model, dimension, preprocessing rule or distance configuration can change that space, so comparing a new query vector to old document vectors may produce poor results even when both arrays have valid numbers. Treat the change as a data migration with a baseline, compatibility plan, backfill, evaluation and rollback rather than swapping a model string in production.
Inventory the complete vector contract
Record the model and version, output dimensions, normalization, distance metric, text preprocessing, chunker, metadata schema and vector-store configuration. Confirm the query path and document-ingestion path use compatible settings. A dimension that happens to match does not mean two models produce comparable vectors.
Measure quality and capacity before backfilling
Build a representative query and source set, then compare the current and candidate models on retrieval relevance, recall, language coverage and downstream answer grounding. Estimate tokens to re-embed, storage, write throughput and temporary duplication. Newer or larger embeddings can cost more and may improve some tasks while reducing quality on others.
Namespace embeddings by their full configuration
Attach a clear model and index version to each collection, vector name or record metadata. Prevent a new query embedder from silently searching the old-only index. Keep the ingestion pipeline able to identify which documents have been processed with each model and configuration revision.
| Current and candidate model | Dimensions and metric | Corpus and backfill volume | Quality gate | Switch and rollback plan |
|---|---|---|---|---|
How do you migrate vectors without serving mixed results?
Choose a blue-green collection or separate named vector
A blue-green migration builds a second collection and switches an alias after it is populated. A vector database with named-vector support may allow a second model vector beside the old one. Choose based on your store's documented atomicity, schema and rollback behavior; these patterns are not available or identical in every database.
Dual-write new and changed source records
Once the candidate path is ready, write updates to both old and new indexes during backfill. Make each update and delete idempotent and carry the same stable source ID, tenant, permission metadata and source version to both paths. Track failures and lag; do not let the backfill restore a document that was deleted after the job began.
Backfill in bounded, restartable batches
Use a checkpointed job with a stable cursor, per-record outcome, retry limit and dead-letter path. Re-read current source data and permissions before indexing it. Protect provider credentials, bound concurrency and support a safe pause or resume so transient errors do not restart the entire corpus or overload dependencies.
How do you switch traffic and verify rollback readiness?
Check completeness and evaluate before the cutover
Compare source and indexed record counts by tenant and version, check for failed or stale records, and run the same retrieval evaluation against both indexes. Shadow or canary the candidate only with approved data handling. Promote after quality, latency, cost and authorization checks meet explicit thresholds.
Switch through one controlled routing point
Use an alias, configuration value or versioned service boundary so the production search path changes atomically. Monitor candidate usage, zero-result rate, relevant-passage recall, answer grounding, latency and tenant-filter enforcement. Keep the old index available until the rollback window closes and the new route has met its stability criteria.
Retire the old index only after data and deletion checks
After promotion, stop old-model writes, verify update and delete reconciliation, retain only what your approved retention policy requires, and then remove obsolete vectors and temporary files. Document the point at which rollback is no longer possible and preserve the migration record without keeping unnecessary source content.
Embedding model migrations: FAQs
Can I query old document vectors with a new embedding model?
Do not assume they are compatible. Even matching vector dimensions do not guarantee a shared semantic space. Keep the query and document embedder configuration aligned for each active index.
Should I pause writes while re-embedding?
Usually you need a plan to keep changes consistent, such as dual-writing or an explicitly controlled write pause. The right approach depends on your vector store and whether deletes or partial updates can be reconciled safely.
How do I know the new model is better?
Use representative retrieval labels and end-to-end answer evaluations, including no-answer and multilingual cases. Compare quality, latency, token cost and storage; model announcements alone do not establish fit for your corpus.
When may I delete the old vectors?
After the new index passes its gates, production traffic is stable, permissions and deletions reconcile, and the agreed rollback window has ended. Record that this is the point where restoring the old route may no longer be possible.
Related practical guides
Related issue guides
Sources and publication record
Draft prepared 27 September 2026; engineering, security and editorial review pending · Sources checked .
- Vector embeddingsOpenAI API documentation
- Migrate to a new embedding modelQdrant documentation
- Evaluation best practicesOpenAI API documentation
- Connect to SharePoint data sources with access control listsAmazon Web Services Bedrock Knowledge Bases
- Rate limitsOpenAI API documentation