Matryoshka Embeddings: Dimensions, Normalization, and Retrieval
Understand nested embedding dimensions, calculate raw storage tradeoffs, and test shortlisting quality before adopting a smaller representation.

Matryoshka Representation Learning trains representations so selected prefixes remain useful at smaller dimensions. It allows a retrieval system to compare different representation sizes from a compatible model. It does not establish that an arbitrary embedding can be truncated safely or that a smaller vector preserves identical recall on every task.
Use it when the storage or retrieval cost of vectors matters enough to justify measuring the quality tradeoff. Keep embedding-model inference, vector storage, candidate retrieval, and reranking as separate costs.
What changes during training
The original Matryoshka paper trains representations at multiple nested sizes. Instead of optimizing only the full vector, training also rewards useful information at selected shorter prefixes. The result is flexibility within one compatible representation family.
A larger dimension count alone does not prove that one model is better than another. Training data, objective, language, task, and preprocessing all affect retrieval. Compare actual model revisions on the intended workload rather than inferring quality from vector length.
Follow the model's preprocessing contract
The Sentence Transformers guide explains truncation and its downstream tradeoffs. Some model recipes require transformations before truncation and normalization afterward. Follow the selected model card rather than copying a generic slice operation into every embedding pipeline.
This helper performs only prefix selection and L2 normalization on a supplied vector. It is useful for testing that step; it does not generate embeddings or replace model-specific preprocessing.
import math
def normalized_prefix(vector, dimensions):
if (not isinstance(dimensions, int) or isinstance(dimensions, bool)
or not 1 <= dimensions <= len(vector)):
raise ValueError("Invalid prefix dimension")
prefix = [float(x) for x in vector[:dimensions]]
if not all(math.isfinite(x) for x in prefix):
raise ValueError("Vector values must be finite")
norm = math.sqrt(sum(x * x for x in prefix))
if norm == 0:
raise ValueError("Cannot normalize a zero prefix")
return [x / norm for x in prefix]
Apply the same model revision, preprocessing, dimension, and normalization to queries and documents. Mixing 128-dimensional query vectors with incompatible document vectors is not a meaningful retrieval comparison.
Calculate the storage part precisely
For one million float32 vectors, 768 dimensions contain 3.072 GB of raw values. At 128 dimensions, they contain 0.512 GB: one-sixth as many values, or about 83.3% less raw vector storage.
That arithmetic excludes index structures, IDs, metadata, replicas, full-vector reranking storage, backups, and application overhead. It is not an 83.3% reduction in total service cost. If the encoder still computes the full representation before truncation, its inference work does not automatically fall by the same factor.
Shortlist and rerank with separate evidence
A two-stage design can use shorter vectors for a candidate search and larger vectors to rescore that shortlist. However, a relevant document excluded from the shortlist cannot be rescued by the second stage. Measure candidate recall before interpreting final ranking improvements.
Storing full vectors for reranking can retain much of the original storage requirement. Record where the full vectors live, whether they require an extra fetch, and what that does to tail latency. A single model does not imply a single index or zero additional storage.
Code retrieval needs code-shaped tests
For code, include exact symbol names, natural-language descriptions, call relationships, version changes, and cases where the requested implementation does not exist. Keep repository and access filters constant across dimension settings.
Compare at least the full representation and one supported smaller size under the same candidate budget. Record model revision, query/document prompts, preprocessing, index settings, corpus commit, and labeled relevant chunks. Report recall changes alongside memory and latency measurements.
For the general mechanics behind the index, see vector search; for ranking evaluation, see contextual reranking.
If the broader decision is whether to operate the retrieval stack yourself, explore Supermemory’s search API alongside your embedding experiments. Compare the end-to-end results and operating work; the API’s documented options determine which controls are available.