explainer / Oct 11, 2026

Changing an Embedding Model Changes the Library Index

A new chat model can use existing retrieved text. A new embedding model changes how that text is found. Understand the migration boundary before rebuilding a private library.

By Stackarr EditorialLocal AI · RAG · Embeddings · Open WebUI
An open book and two bundles of blank index cards sit on a round table beside a bookcase and a small home server.
The source library can stay the same while the index used to find passages needs to be rebuilt.

Changing the embedding model in a self-hosted document assistant is an index migration, not just a different voice in the chat window. The stored document vectors and the vectors made from new questions must be compatible. Replacing only the question-side model can leave a healthy-looking service retrieving the wrong passages.

For a homelab library of appliance manuals, network notes and home-server recovery instructions, that distinction matters more than whether the chatbot still answers. This explainer covers dense-vector retrieval and Open WebUI-managed knowledge bases. It is a planning aid, not a universal upgrade procedure for every vector database or application release.

The answering model and the finding model have different jobs

In retrieval-augmented generation, or RAG, the application finds relevant text and supplies it to the model generating the response. An embedding model converts text into numerical vectors used for similarity search. The answering model then works with the retrieved passages.

Changing the answering model alone does not inherently require recreating document embeddings. It can still change answer quality or how much retrieved text fits into context. Changing the embedding model alters the representation used to find those passages in the first place.

Ollama's embedding documentation explicitly pairs the indexing and querying model. Treat that pairing as part of the stored library's configuration, rather than a preference belonging only to the next chat.

Matching vector length is necessary, not sufficient

A vector database has structural expectations. In Qdrant's collection model, a vector space has a defined dimensionality and distance metric. Named vectors allow separate configurations within a collection; that does not make unrelated models interchangeable.

A different vector length can produce an obvious incompatibility. Equal length is more deceptive: it does not establish that two models assign compatible meanings to their coordinates. Open WebUI's documentation warns that embeddings made by different models occupy incompatible vector spaces. A request completing without an error is therefore not evidence that retrieval remains correct.

For an operator, the useful record includes the embedding model identity, its output dimensions and the collection configuration. A note saying only that the service uses a local model is not enough to reconstruct the search arrangement later.

A rebuild has a scope, and chat attachments may sit outside it

The current Open WebUI RAG documentation describes a knowledge-base reindex that replaces the existing vector collection, splits stored text using the current chunk settings and generates embeddings with the configured model. It also rebuilds the individual file collections for files belonging to those knowledge bases.

Files uploaded directly into chats without joining a knowledge base are a different case. The documented knowledge-base reindex does not refresh those standalone file collections. Those files need re-uploading to receive the new embeddings.

This creates an important inventory question before any change: is the household's document library actually a set of knowledge bases, or a mixture of collections and forgotten chat attachments? A successful rebuild of one does not prove coverage of the other.

Re-embedding cannot recover text that was never extracted

A scanned manual can have an extraction problem before similarity search ever begins. Open WebUI documents that reindexing uses the text already extracted at upload time; it does not reopen and parse the original document. Switching an OCR or extraction engine therefore does not repair existing extracted text through reindexing alone. Re-uploading is needed when those files must be parsed again.

That separates three distinct decisions:

  • Change the answering model when evaluating how responses are produced from available context.
  • Change the embedding model when evaluating how relevant passages are represented and retrieved.
  • Change extraction when the searchable text itself is missing or malformed.

Combining all three changes makes improvement or regression harder to explain. A narrow comparison gives a home-server operator more useful evidence than a wholesale replacement followed by one plausible answer.

Judge the retrieved evidence before judging the prose

A practical evaluation set can be small but specific. Select questions whose answers are known to exist in particular source passages: a maintenance interval in a manual, a port recorded in a network note, or a recovery prerequisite in a runbook. Keep the questions and expected sources unchanged across the comparison.

Inspect whether the retrieved passages contain the necessary facts, not merely whether the final response sounds convincing. Include a question that the library cannot answer and check whether the response invents support. This is a quality check, not an authorization test: access controls still need separate verification.

The expected result is relevant evidence from the intended collection, followed by an answer grounded in that evidence. A missing citation, irrelevant passage or confident unsupported answer warrants investigation even when every container reports healthy.

Recovery must restore a compatible pair

Before authorizing a rebuild, the administrator needs a recoverable copy of the application's persistent state, the underlying documents and a record of the old retrieval configuration. Readers with chat access alone should not assume permission to change shared embedding settings.

Open WebUI's documented replacement of existing collections makes reindexing consequential, not a harmless preview. Plan for indexing work and reduced retrieval availability rather than assuming the operation is instantaneous.

Reverting only the embedding-model setting after rebuilding with another model can recreate the same mismatch in reverse. Recovery needs the old model with its compatible old index, or a deliberate rebuild under the chosen model. The durable rule is simple: preserve the relationship between the search model and the document vectors, not just the documents or the model name in isolation.

Verification ledger

Sources and further reading

  1. EmbeddingsOllama · Primary source
  2. Retrieval Augmented Generation (RAG)Open WebUI · Primary source
  3. CollectionsQdrant · Primary source