How Vector Search Works: Embeddings, Indexes, and Tradeoffs
Understand embeddings, exact and approximate search, metadata filters, and index updates. Choose vector retrieval based on the questions your product must answer.

Vector search finds records whose numerical representations are close to a query's representation. An embedding model turns text or other input into vectors; a search index retrieves nearby vectors under a chosen distance measure. This makes it possible to find related meaning even when the words differ.
An embedding model, a vector index, and an application memory system are different components. The model determines the representation. The index determines how to search it. The application determines which records belong to a user, which version is current, and how retrieved evidence affects an answer.
What is an embedding?
An embedding is a vector produced by a model to represent an input. The useful similarity relationships depend on how that model was trained and what kinds of input it supports. Two vectors having the same number of dimensions does not make them compatible if they came from unrelated models.
When indexing documents, keep the embedding model and revision in your ingestion metadata. Use the matching query representation expected by that model. If the model has separate query and document instructions, follow them consistently.
Switching embedding models can require re-embedding the corpus and rebuilding the index. Test the new representation against a held-out query set before changing the production read path.
How does similarity become a result list?
For cosine similarity, the angle between vectors matters more than their raw length. An exact search scores every eligible vector. An approximate nearest-neighbor index examines a selected part of the space to reduce work, trading some retrieval completeness for speed or resource use.
This Python function shows the calculation for tiny, invented vectors. Using invented vectors keeps the cosine calculation easy to inspect.
from math import sqrt
def cosine(a, b):
if not a or len(a) != len(b):
raise ValueError("Require nonempty vectors with matching dimensions")
na = sqrt(sum(x * x for x in a))
nb = sqrt(sum(x * x for x in b))
if na == 0 or nb == 0:
raise ValueError("Cosine is undefined for a zero vector")
return sum(x * y for x, y in zip(a, b)) / (na * nb)
assert cosine([1, 0], [1, 0]) == 1
assert cosine([1, 0], [0, 1]) == 0
assert cosine([1, 0], [-1, 0]) == -1
A high similarity score says the representations are close under this measure. It does not prove factual agreement, current validity, or access permission. A sentence denying a claim may be semantically close to a sentence asserting it.
What does a vector database add?
A database can provide persistence, indexes, metadata filters, updates, deletion, and operational features around vectors. Those capabilities vary by implementation and deployment.
For example, Pinecone's metadata filters constrain results using stored attributes. Your application still has to populate the right attributes and derive permitted filters from authenticated identity.
Inspect how filtering interacts with approximate retrieval. A filter can leave too few useful candidates if the implementation retrieves a broad shortlist and filters later. Evaluate heavily filtered queries, not only global searches against the full index.
Estimate the storage you can actually account for
An illustrative collection of one million 1,536-dimensional float32 vectors contains about 6.144 billion bytes of raw vector values: one million × 1,536 × four bytes. That is approximately 6.144 decimal GB, before index structures, metadata, source text, replication, and backups.
This arithmetic is a capacity component, not a vendor bill. Quantization may reduce representation size with its own quality tradeoffs. Managed pricing may depend on a different combination of storage, compute, operations, or service tiers.
Use a workload model rather than choosing a database from a single dollars-per-million-vectors number. The memory operating-cost guide covers the additional processing and retrieval costs.
Know the failure modes before changing the model
- Exact IDs may need a lexical path or structured lookup.
- Long chunks can blend several topics into one representation.
- Old and current versions can both look relevant.
- Repeated near-duplicate chunks can crowd out diverse evidence.
- The query may ask for a relationship that is not stated in one passage.
- An unanswerable question can still have a nearest neighbor.
Nearest does not mean sufficient. Permit an empty or insufficient-evidence result rather than forcing an answer from whichever vector ranks first.
When should you add hybrid search or memory logic?
Add hybrid search when exact terms and paraphrases both matter and a comparison shows value. Add graph retrieval when explicit paths answer questions that passage similarity misses.
For cross-session preferences and decisions, define the write, correction, and deletion behavior around the stored records. Those responsibilities can sit on top of a vector database, a relational store, or a managed memory service. The AI memory versus vector databases guide compares that division of work.
Before scaling the index, prove that one authorized question retrieves the right evidence, that an unrelated question can return insufficient evidence, and that a revised source displaces the outdated answer. A fast index is useful only when those behaviors survive growth.
To evaluate the representation behind the index, compare open embedding models on your questions. If vector size is the constraint, test the supported shorter representations explained in the Matryoshka guide.
If your application also needs to carry user context across sessions, try Supermemory alongside your retrieval baseline. Begin with one remembered decision and check how it is retrieved, corrected, and removed before choosing the broader architecture.