Hybrid Search: Why BM25 Plus Vectors Beats Either Alone for RAG
Hybrid search combines two different retrieval methods, full-text keyword matching (BM25) and vector similarity search, then merges their results into a single ranked list. For professional documents full of exact codes, model numbers, acronyms and defined terms, this combination consistently outperforms either method used alone. Reciprocal rank fusion (RRF) is the simplest and most robust way to merge the two ranked lists without having to tune fragile weighting parameters.
Why one retrieval method is rarely enough
Every retrieval-augmented generation system has to answer the same question before it writes a single word of an answer: which passages, out of potentially thousands, are actually relevant to this query? Get that step wrong and no amount of clever prompting rescues the answer, because the model simply never sees the right text. This is why retrieval quality, not the choice of language model, is usually the single biggest lever on answer quality in a RAG system. If you want the full picture of how a query turns into a grounded answer, our guide on how RAG works step by step walks through the whole pipeline.
The trouble is that no single retrieval method is good at everything. Keyword search is excellent at finding the exact string a user typed, but blind to synonyms and paraphrase. Vector search is excellent at finding conceptually related passages, but can quietly drop a document that uses the precise term the user needs, in favour of something that merely sounds similar. Professional knowledge bases, contracts, technical manuals, regulatory texts, policy documents, are full of exactly the kind of content that exposes this trade-off: part numbers, clause references, statutory citations, drug names, error codes. Hybrid search exists because relying on only one of these signals means systematically missing a class of queries.
BM25 full-text search: strengths and blind spots
BM25 (Best Matching 25) is a decades-old ranking function built on term frequency and inverse document frequency. In plain terms, it scores a passage highly if it contains the query's words often, especially words that are rare across the whole collection. It is fast, cheap to run, requires no model inference and is extremely predictable: if a document contains the exact phrase the user searched for, BM25 will almost always surface it near the top.
- Exact matches on part numbers, invoice references, clause numbers and error codes ('Error 0x8007045D', 'clause 14.3', 'ISO 13485')
- Acronyms and proper nouns that a vector model may not have seen enough during training to embed distinctively
- Rare, highly specific terminology where frequency across the corpus is a genuinely strong relevance signal
- Predictable, auditable behaviour: you can explain exactly why a passage scored highly
Its blind spot is equally clear: BM25 has no concept of meaning. A query for 'can I cancel my subscription' will not match a passage that only says 'terminating your plan' unless both phrases happen to share enough overlapping terms. Synonyms, paraphrases, different languages describing the same concept, and queries phrased as questions rather than keyword strings, all expose this weakness.
Vector search: strengths and blind spots
Vector search works differently. Text is converted into embeddings, dense numerical representations that place semantically similar content close together in a high-dimensional space, and retrieval becomes a nearest-neighbour search in that space. If you have not come across embeddings before, our plain-English guide to text embeddings explains the mechanism without the maths. The practical upshot is that vector search can match a question to an answer even when they do not share a single word in common.
- Paraphrased or conversational queries ('how do I get a refund' matching a passage about 'reimbursement procedures')
- Cross-lingual or loosely worded searches where exact terms differ
- Conceptual questions that span several related ideas rather than one keyword
- Natural-language questions typed into a chat interface rather than a search box
Its blind spot is the mirror image of BM25's strength. Embedding models compress meaning, and in doing so they can blur distinctions that matter enormously in professional contexts: a model number that differs by one digit, a clause reference, a drug dosage, a regulation number. Two passages that are textually almost identical except for one critical figure can end up with very similar embeddings, which is precisely the kind of near-miss that causes confident, wrong answers. This is one of the mechanisms behind AI hallucinations that our guide on reducing AI hallucinations covers in more depth.
BM25 versus vector search at a glance
| Aspect | BM25 (full-text) | Vector search (semantic) |
|---|---|---|
| Best at | Exact terms, codes, rare words | Paraphrase, synonyms, concepts |
| Weak at | Synonyms, paraphrase | Exact codes, precise figures |
| Cost | Very low | Requires embedding inference |
| Explainability | High (term overlap is visible) | Lower (similarity score only) |
| Typical failure | Misses a differently-worded match | Confuses near-identical passages |
Reciprocal rank fusion: merging two ranked lists without guesswork
Once you accept that you need both signals, the next problem is how to combine them. A naive approach is to add the two raw scores together, but BM25 scores and cosine similarities live on completely different numeric scales, so a weighted sum requires constant re-tuning as your document collection grows or changes. Reciprocal rank fusion sidesteps this by ignoring the raw scores entirely and working only with rank positions, which makes it far more stable across collections and query types.
The formula in plain English
For each passage, you take its rank in the BM25 results list and its rank in the vector results list, then compute a score as 1 divided by (a constant k plus the rank), and sum that across both lists. A typical value for k is 60. Passages that rank highly in either list, or ideally both, end up with the highest combined score. A passage that appears near the top of only one list can still rank well overall; a passage that appears in both lists, even in the middle, often beats a passage that appears at the very top of just one.
Reciprocal rank fusion score
RRF score = 1 / (k + rank in BM25 results) + 1 / (k + rank in vector results). A passage missing from one list simply contributes nothing from that term, no normalisation of raw scores required.
A worked example
Imagine a maintenance manual knowledge base and the query 'replacement part for the E4 filter housing'. BM25 ranks a passage mentioning 'E4 filter housing, part reference FH-2210' at position 1, because of the exact code match. Vector search ranks that same passage at position 9, because the embedding leans more heavily on the surrounding descriptive text than the alphanumeric code. Meanwhile, a different passage describing 'replacing a clogged housing filter' in more general terms ranks 2nd by vector search but doesn't appear in the top 20 BM25 results at all, since it never uses the word 'E4'. With RRF (k=60), the first passage scores roughly 1/61 + 1/69 (about 0.0310), comfortably ahead of a passage that only scores well on one list. This is exactly the outcome you want: the exact-code passage wins, but the conceptually relevant one is still in the mix rather than discarded.
Where hybrid search goes wrong in practice
Hybrid search is not a magic switch you flip once and forget. A few recurring mistakes undermine it in real deployments.
- Chunking that ignores structure: if a table or clause is split mid-row, neither BM25 nor vector search has a clean passage to match against. See our guide on chunking strategies for RAG for structure-aware approaches that keep both signals strong.
- Retrieving too few candidates from each method before fusion: RRF needs a reasonably deep candidate pool (commonly 20-50 results per method) to have anything meaningful to merge.
- Treating keyword expansion as optional: a good hybrid implementation also expands query terms with related keywords before the full-text pass, catching near-synonyms that pure BM25 would miss.
- Forgetting that fusion happens before reranking, not instead of it: for high-stakes answers, a final pass that checks the fused top results against the literal query still adds value.
- Assuming one k value suits every corpus: it rarely needs much tuning, but it is worth sanity-checking on a handful of representative queries rather than trusting a default blindly.
A checklist for choosing a RAG retrieval setup
- Audit your documents for exact-match content: part numbers, clause references, statutory citations, product codes, dosages.
- If exact-match content is common, do not rely on vector search alone, however good the embedding model.
- Run both BM25 and vector search over the same candidate pool, with a reasonable depth (20+ results each).
- Merge with reciprocal rank fusion rather than a hand-tuned weighted sum, especially if your document set will keep growing.
- Keep chunk boundaries aligned with document structure (headings, clauses, table rows) so both signals have clean units to score.
- Check citations on a sample of answers before trusting the system in production; our guide on checking AI answer citations sets out a practical method.
- Re-test retrieval whenever you add a very different type of document (a new language, a much longer manual, a spreadsheet-heavy export).
Hybrid search without building the pipeline yourself
Implementing BM25 indexing, embedding generation, candidate retrieval and reciprocal rank fusion correctly, and keeping it fast as a document collection grows, is a genuine engineering project. It is one of the main reasons teams look at RAG as a service rather than assembling their own vector database, search index and fusion logic from scratch. Kopik runs hybrid search (full-text plus semantic keyword expansion) by default on every knowledge base: you upload PDFs, Word documents, text or Markdown files, Kopik extracts, chunks and indexes them, and queries are answered with cited passages drawn from both retrieval signals. You can browse existing public knowledge bases in the catalogue, create your own for free in minutes at create a knowledge base, or integrate retrieval directly into your own tools through the REST API and MCP server, usable from Claude, Cursor, ChatGPT and other MCP clients.
See hybrid search in action
Upload a set of professional documents and query them with full-text plus semantic search combined, cited passages included.
Final thought: pick the combination that matches your documents
There is no universal 'best' retrieval method, only the method best suited to what your documents actually contain. A corpus of narrative reports with few exact codes may lean more on semantic search; a corpus of technical manuals, contracts or regulatory texts almost always benefits from BM25 catching what vectors miss. Hybrid search with reciprocal rank fusion gives you both without forcing a brittle manual trade-off, which is why it has become the default rather than the exception for serious RAG retrieval.
Frequently asked questions
What is hybrid search in RAG?
Hybrid search is a retrieval approach that runs both full-text keyword search (typically BM25) and vector similarity search over the same documents, then merges the two ranked result lists into one, usually with reciprocal rank fusion, so that exact matches and conceptual matches are both captured.
Is BM25 better than vector search?
Neither is universally better. BM25 wins on exact terms, codes and rare words; vector search wins on paraphrase, synonyms and conceptual queries. Combining both with hybrid search outperforms using either one alone, especially on professional documents.
What is reciprocal rank fusion?
Reciprocal rank fusion (RRF) is a method for merging ranked result lists from different retrieval methods. Instead of combining raw scores, which sit on different scales, it scores each passage as the sum of 1 divided by (a constant plus its rank) in each list, rewarding passages that rank well in either or both lists.
Do I need to tune weights for hybrid search?
With reciprocal rank fusion you generally avoid weight tuning altogether, since it works on rank positions rather than raw scores. The main parameter, the constant k (commonly 60), rarely needs adjustment, though it is worth checking against a few representative queries.
Does hybrid search slow down retrieval?
Running two retrieval methods adds some overhead compared with one, but BM25 is very cheap computationally and the fusion step itself is trivial arithmetic on rank positions, so the added latency is usually small relative to the overall answer generation time.
Can I use hybrid search without building my own infrastructure?
Yes. Platforms offering RAG as a service, including Kopik, run hybrid search by default when you upload documents, handling indexing, embedding and reciprocal rank fusion automatically so you can query through the website, a REST API or an MCP server.
Get the Kopik newsletter
New knowledge bases, RAG guides and product news. One email every week or two, unsubscribe in one click.
By subscribing you agree to receive our newsletter. We never share your address.