Hybrid Search: Combining Lexical and Semantic Retrieval
Combine keyword and vector results without mixing incompatible scores. Learn rank fusion, reranking, filtering, and how to evaluate the tradeoffs.

Hybrid search combines more than one retrieval signal, commonly lexical matching and vector similarity. Lexical search is useful for exact terms such as error codes and product identifiers. Semantic search helps when a question and its answer use different wording. Combining them can broaden the candidate set, but the result still needs evaluation and permission filtering.
The word “hybrid” is overloaded. Some products use it for dense plus sparse retrieval; others use it for several content types or stores. Check the API's definition before assuming two systems perform the same operation.
Match the search method to the question
| Question | Signal likely to help | Failure to test |
|---|---|---|
| “What does ERR-482 mean?” | Exact lexical matching | A semantically similar but different code ranks first |
| “How do I stop invoices renewing?” | Semantic matching | The source says “cancel recurring billing” |
| “Does the enterprise plan support SSO?” | Both | A consumer-plan passage uses similar language |
| “Which service depends on the failed database?” | Explicit relationships or structured lookup | A passage mentions both entities without establishing dependency |
These are hypotheses for your corpus. A good keyword system can use synonyms, and a vector model may recognize an exact code. Measure the cases that matter instead of declaring a universal winner.
Combine ranks when scores are incompatible
BM25 scores and cosine similarities do not share a natural scale. Adding their raw values can let one method dominate simply because its scores are numerically larger.
Reciprocal rank fusion (RRF) combines positions instead. Each result receives a contribution based on its rank in each list. This standalone Python example treats each list as unique document IDs and uses a conventional illustrative constant of 60:
def rrf(rankings, k=60):
if k <= 0:
raise ValueError("k must be positive")
scores = {}
for ranking in rankings:
seen = set()
rank = 0
for doc_id in ranking:
if doc_id in seen:
continue
seen.add(doc_id)
rank += 1
scores[doc_id] = scores.get(doc_id, 0.0) + 1 / (k + rank)
return sorted(scores, key=lambda doc_id: (-scores[doc_id], doc_id))
assert rrf([["policy", "faq"], ["faq", "guide"]])[0] == "faq"
assert rrf([["a", "a", "b"]]) == rrf([["a", "b"]])
assert rrf([]) == []
Here, faq appears in both lists and wins the fused ranking. That does not establish that it answers the question. Fusion creates a candidate ordering; the evidence still needs to support the answer. Tune candidate counts and fusion settings on held-out questions rather than treating this constant as a product requirement.
Apply permissions to every retrieval path
If the lexical path respects the tenant filter but the vector path does not, the combined result can leak data. Apply authorized scope consistently, including reranking inputs, caches, and any parent-document expansion.
An inaccessible document must not enter a downstream model prompt just because the application removes its citation afterward. For shared applications, use the isolation checklist to test the full path.
Version filters matter too. A stale policy that matches exact words can outrank the current policy. Keep the old version only when the question and access rules allow historical evidence, and label it clearly.
Add reranking for a specific reason
A reranker evaluates candidate relevance after initial retrieval. It is useful to test when supporting evidence already appears somewhere in the candidate pool but is not selected for the final prompt.
It cannot recover a passage that neither search path returned. It also adds processing work, and any external reranking service becomes another destination for the candidate text. Measure its effect on answer support, latency, and cost under the same workload.
Weaviate's hybrid-search documentation describes a concrete implementation and fusion options. Those details are product-specific; the Python example above is a general rank-fusion illustration rather than a reproduction of Weaviate's scoring.
Keep graph traversal a separate choice
Graph traversal follows relationships such as account-to-contract or service-to-dependency. It can complement lexical and semantic search, but “hybrid search” does not automatically mean graph traversal is involved.
First identify whether the task needs an explicit relationship path. If it does, preserve the path and its supporting sources in the answer. The graph-versus-vector guide describes that decision.
Build a comparison that can change your decision
Label a sample of exact-identifier questions, paraphrases, mixed queries, and no-answer cases. Compare lexical-only, vector-only, and fused candidates before adding a reranker. Keep filters, corpus version, and answer settings fixed.
Track evidence recall, first useful result, duplicate context, unsupported answers, and time spent in each stage. Report results by query type: an average can hide the exact-ID failures that affect your support team most.
If the combined system brings no meaningful benefit for your workload, keep the simpler path. If it helps one query class, consider routing that class deliberately rather than paying for every retrieval method on every request.
Supermemory's search modes use “hybrid” to return both extracted memories and document chunks. Distinguish that content-selection meaning from the lexical-plus-vector design discussed here when configuring an integration.
If relevant passages enter the candidate pool but rank too low, use the contextual reranking guide to test that stage separately.
To evaluate Supermemory as a retrieval option, test it with your own query set using the documented search modes above. Compare returned evidence, context size, and latency against your current path; let the workload determine whether it is a useful fit.