Choosing Embedding APIs for Production Retrieval
Compare embedding providers by supported inputs, retrieval quality, compatibility and workload cost. Keep embedding APIs separate from databases and memory services.

Choose an embedding API by the representation your retrieval workload needs. An embedding endpoint returns vectors. A vector database stores and searches records. A memory service adds behavior around retained context. They can be combined, but a memory benchmark is not an embedding-model leaderboard.
The options below cover text, code and multimodal representations. Use the evaluation method to compare them on your documents and queries.
Define the input and output contract
Record supported modalities, input limits, query/document instructions, output dimensions and data types. A text embedding model does not automatically accept a PDF file simply because text can be extracted from that PDF.
Likewise, a large input limit does not mean every document should become one vector. A long document covering several topics may need smaller searchable units for precise answers. Compare evidence coverage using the chunking guide.
OpenAI text embeddings
OpenAI's embedding guide documents text-embedding-3-small and text-embedding-3-large, including configurable output dimensions. Their default dimensions are 1,536 and 3,072 respectively.
Evaluate those text-embedding endpoints for the languages and tasks you actually serve. Keep query and document preprocessing consistent when comparing them.
Follow the selected model's limits and pricing. Do not equate a smaller vector with an identical reduction in encoding cost or guaranteed unchanged recall.
Voyage embeddings
Voyage's model table distinguishes general-purpose, code, finance and legal models. For example, the listed Voyage 4 general-purpose models have a 32,000-token context length and support several output dimensions. The legal model has a different limit; do not apply one row's specifications to every model.
Use the documented query and document input types where appropriate. Test domain-specific candidates on your own corpus instead of assuming a domain label proves superiority. Multimodal models have a separate input contract that should be evaluated directly for image or video workloads.
Cohere embeddings
Cohere's Embed documentation covers embedding models and their supported text or multimodal inputs. Check the specific model, dimension choices and output types rather than treating every Cohere embedding endpoint as interchangeable.
For visually rich source material, evaluate whether the representation retains the information needed for the question. A page containing a chart requires different evidence checks from plain prose. Larger input limits can reduce some preprocessing constraints, but they do not eliminate the need to select useful retrieval units.
Database integrations are a separate choice
Weaviate's model-provider integrations can generate vectors through configured providers.
Pinecone also documents integrated embedding. Confirm which service performs encoding, where data is sent and how the charges appear. A database integration changes who invokes the model; it does not remove the model's compatibility and quality constraints.
Supermemory's search API is evaluated as part of a memory/retrieval service. Evaluate its returned evidence and answers separately from the raw embedding models you might use in a custom stack.
Build a comparison that matches production
Use representative documents and questions, including exact identifiers, paraphrases, rare terms, changed documents and missing answers. Label acceptable supporting passages before comparing models.
Keep corpus revisions, access filters, chunking and candidate budgets fixed. Change one major component at a time. If a model requires different preprocessing, record that difference rather than hiding it in an otherwise identical-looking result table.
Measure:
- Whether supporting evidence appears within the selected candidate set.
- Whether the final answer uses that evidence correctly.
- Query encoding and retrieval latency separately.
- Rate-limit errors and retry volume under expected concurrency.
- Ingestion, storage and serving costs for the same workload.
MTEB results can help identify candidates, but select the relevant tasks and benchmark revision. They do not establish API availability or p99 latency under your traffic.
Calculate raw vector storage
One million float32 vectors at 1,536 dimensions contain 6.144 decimal GB of raw values. At 3,072 dimensions they contain 12.288 GB. Index structures, metadata, source content, replicas and backups add to those figures.
Reducing dimensions can lower raw-vector storage, but retrieval quality must be tested. The Matryoshka guide explains why supported shorter representations differ from arbitrarily truncating any embedding.
For pricing, estimate new and changed source tokens, query volume, retries and re-embedding. Add storage and reranking if they are part of the chosen system.
Plan model changes before they happen
Pin model identifiers and preprocessing where possible. Two models producing the same number of dimensions are not automatically compatible. If changing representations, evaluate a parallel index and cut over only after the new path passes the acceptance cases.
Preserve source IDs and revisions so reprocessing can target the intended documents. A rollback should restore a compatible query model and index together; changing only one side can silently damage retrieval.
Decide which layer you want to operate
Choose an embedding API when you need control over representations and already have, or want to build, the surrounding retrieval path. Choose a database integration when it provides the controls you need with less encoding orchestration. Evaluate a memory service when persistent context, ingestion and lifecycle behavior are the larger requirement.
Try Supermemory alongside your retrieval baseline if that last requirement applies. Compare the complete answers and operating work, while keeping the embedding-model experiment a separate, reproducible decision.
Frequently asked questions
Which measurements help compare embedding APIs?
Compare evidence recall, query-encoding latency, error rates and cost on the same corpus and question set. Use relevant embedding-benchmark tasks to identify candidates for that evaluation.
Does a longer embedding input limit remove the need for chunking?
No. Smaller units may still be necessary for precise evidence retrieval. Test the document types and questions you need to answer.
Does Weaviate require manually generated vectors?
No. It supports configured model-provider integrations as well as supplied vectors. Check the selected integration and its data flow.