Hybrid Search: Why BM25 + Vectors Beats Either Alone for RAG
Hybrid search runs a keyword search (BM25) and a vector similarity search on the same query, then merges the two ranked lists into one. For professional documents full of part numbers, acronyms, clause references and domain jargon, this consistently retrieves better passages than either method alone, because BM25 catches exact terms that embeddings blur and vectors catch paraphrases that keyword matching misses entirely.
What BM25 actually does
BM25 (Best Matching 25) is the scoring function behind most full-text search engines, including Elasticsearch, PostgreSQL's full-text search and the classic library-catalogue search box. It counts how often query terms appear in a document, weights rare terms higher than common ones, and penalizes documents that are unusually long (so a 200-page manual doesn't automatically outrank a focused paragraph just because it repeats a word more times). It has no understanding of meaning: if the document says "ECCN" and the query says "export control classification number," a pure BM25 search may not connect the two unless both exact strings appear somewhere.
What BM25 is extremely good at is exact-match precision: part numbers, error codes, statute citations, product SKUs, drug names, acronyms. These are the terms where a vector embedding often does more harm than good, because "torque value AN365-1032" and "torque value AN365-1032A" can end up with nearly identical embeddings even though they refer to different fasteners.
What vector search adds, and where it falls short alone
Vector search converts text into numerical embeddings and ranks documents by how close their embedding is to the query's embedding in that space. It is what makes a RAG system answer "why does my mower keep stalling when it's hot" even if the manual never uses the word "stalling" and instead talks about "the engine loses power under thermal load." For a deeper walkthrough of how that conversion works, see Embeddings Explained for Non-Specialists and What Is a Vector Database?.
The weakness is the mirror image of BM25's weakness. Embeddings compress meaning, and that compression can flatten the exact distinctions that matter in professional documents. A contract clause numbered "9.3" and one numbered "9.4" will have nearly identical embeddings because the surrounding language is similar. A query for "CFR 1910.147" might retrieve a passage about a completely different standard that happens to discuss similar concepts, because the embedding model never treats citation numbers as meaningfully different tokens the way a human reader does.
Why professional documents need both
Most RAG systems built for general knowledge (trivia, news summaries, casual Q&A) can get away with vector search alone, because the cost of a slightly fuzzy match is low. Professional and regulatory documents are a different case: a technician, auditor or compliance officer is often searching for one specific fact, and being off by one clause number or one part revision is a real error, not a rounding error.
Exact identifiers: part numbers, error codes, citations
Maintenance manuals, FAA advisory circulars, OSHA standards and export control lists are full of alphanumeric identifiers that behave like proper nouns. A hybrid system lets the BM25 side lock onto "AC 43.13-1B" or "error code 21" exactly while the vector side keeps searching for the conceptual neighborhood in case the user phrased things loosely.
Synonyms, paraphrasing and plain-language questions
Real users rarely type the exact vocabulary of the source document. Someone asking "can I replace the battery myself" needs to match a manual section titled "owner-performed maintenance," and someone asking "do I need a new filter" needs to match a section about "regeneration failure diagnostics." This is where vector search carries the retrieval almost entirely.
Negation and numeric thresholds
Both methods struggle here on their own, but combining them helps. A query like "quarter-inch point-of-operation guard requirement" benefits from BM25 catching the exact fraction and unit, while the vector side still retrieves related guarding passages that use different phrasing for the same threshold.
How reciprocal rank fusion (RRF) combines the two
The simplest and most widely used method to merge a BM25 ranking and a vector ranking into a single list is reciprocal rank fusion. Instead of trying to compare BM25 scores and cosine similarity scores directly, which live on incompatible scales, RRF only looks at each document's *rank position* in each list. A document's fused score is the sum, across both rankers, of 1 divided by (a constant k, typically around 60, plus its rank in that list). Documents that rank well in either list, or reasonably well in both, rise to the top.
Reciprocal rank fusion example (k = 60)
| Document | BM25 rank | Vector rank | RRF score | Final rank |
|---|---|---|---|---|
| Doc B | 4 | 1 | 0.0320 | 1 |
| Doc A | 1 | 5 | 0.0318 | 2 |
| Doc C | 2 | 8 | 0.0308 | 3 |
| Doc D | 10 | 2 | 0.0304 | 4 |
Notice what happened: Doc A was the top BM25 match, but Doc B, which ranked a respectable 4th on keywords and 1st on vector similarity, edged it out after fusion. Doc D, despite a strong vector rank of 2, stayed near the bottom because its BM25 rank was weak, meaning it probably doesn't contain the literal terms the user typed. RRF rewards documents that are good candidates from more than one angle, without requiring you to tune a weighting formula between two differently scaled scores.
Why not just average the raw scores?
BM25 scores are unbounded and depend on corpus statistics (term frequency, document length), while cosine similarity is bounded between -1 and 1. Averaging them directly means whichever score happens to have a larger numeric range silently dominates the fusion. RRF sidesteps that entirely by working with rank positions instead of raw scores, which is why it has become the default in most hybrid search implementations, including Elasticsearch's native hybrid retriever and most managed RAG platforms.
A worked example: a maintenance question
Say a technician asks a knowledge base built from a CNC machine's service manual: "what's the torque spec for the spindle housing bolts." A pure vector search might return a passage about general torque safety practices because it's semantically close to the whole topic of "torque" and "bolts," without actually containing the specific housing spec. A pure BM25 search might miss a relevant passage that uses "fastener" instead of "bolt" throughout. Hybrid search runs both: BM25 surfaces the passage that literally contains "torque," "spindle," and "housing" near each other, vector search surfaces the passage using "fastener torque values for the spindle assembly," and RRF merges them so the technician sees the exact spec passage ranked first with a related passage close behind, rather than gambling on a single retrieval method guessing correctly.
This is the same mechanism that makes chunking strategy matter so much for retrieval quality. If a chunk boundary splits a torque value from the bolt it refers to, neither BM25 nor vector search can save you. See Chunking Strategies for RAG for how fixed-size, semantic and structure-aware splitting affect what ends up searchable in the first place.
How Kopik implements hybrid search
Every knowledge base created on Kopik runs hybrid search by default: full-text indexing alongside semantic keyword expansion, so a query benefits from exact-match precision and conceptual recall without any configuration. When you upload a PDF, Word file or Markdown document to a base, Kopik extracts the text, chunks it, and indexes it for both retrieval paths automatically. You don't need to choose between a keyword engine and a vector database, or tune a fusion weight yourself; the retrieval pipeline behind each knowledge base already does that work before the answer gets generated with cited passages, or returned as raw passages if you prefer answers without generation. If you want to see how the full pipeline fits together from upload to answer, How RAG Works, Step by Step walks through one question end to end, and What Is RAG (Retrieval-Augmented Generation)? covers the broader concept for readers who are new to it.
- Upload documents once; both the keyword index and the vector index are built automatically
- Query from the website, a REST API with API keys, or an MCP server usable from Claude, Cursor or ChatGPT
- Choose grounded answers with citations, or raw passages only, depending on what you're building
- Keep a base private for internal use, or list it publicly in the catalogue and set a per-question price
Try hybrid search on your own documents
Create a knowledge base for free, upload your manuals, policies or technical files, and query them with hybrid search already built in.
Checklist: is your retrieval actually hybrid?
If you're evaluating a RAG platform, a vector database, or building your own pipeline, use this checklist to confirm you're getting real hybrid retrieval rather than vector search with a marketing label.
- Does the system index full text separately from embeddings, or does it only do approximate nearest-neighbor search?
- Does it fuse rankings with a documented method (RRF or an equivalent), or just vector results with keyword filters bolted on?
- Can you test an exact identifier (a part number, a citation, an error code) and confirm it's retrieved reliably?
- Can you test a paraphrased, plain-language version of the same question and confirm it's still retrieved?
- Is the fusion automatic per query, or does someone have to manually pick "keyword" vs "semantic" mode each time?
If you answered "no" to more than one of these, you likely have vector search with a search bar, not hybrid search. For technical teams wiring this into developer tools directly, the developer documentation covers the REST API and MCP server setup so you can query hybrid search from your own application or from an AI agent.
When hybrid search is overkill
Hybrid search isn't free: it means maintaining two indexes and running a fusion step on every query, which adds a small amount of latency and infrastructure compared to vector search alone. For small, conversational knowledge bases where exact terminology barely matters, such as an internal FAQ written in plain English with no part numbers or citations, pure vector search is often good enough and simpler to reason about. The tradeoff tips hard toward hybrid the moment your documents contain identifiers, codes, clause numbers or domain acronyms that a reader might type verbatim, which describes most regulatory, technical and legal document sets.
The practical takeaway is to default to hybrid unless you have a specific reason not to. The fusion step costs very little compared to the cost of a wrong answer in a compliance or maintenance context, and a well-built hybrid pipeline degrades gracefully: if BM25 finds nothing, vector search still carries the query, and vice versa.
Frequently asked questions
What is hybrid search in RAG?
Hybrid search runs a keyword search (typically scored with BM25) and a vector similarity search on the same query, then merges the two ranked result lists, usually with reciprocal rank fusion, so the final results benefit from both exact-term matching and conceptual similarity.
Is BM25 still relevant now that we have vector embeddings?
Yes. BM25 remains the best tool for exact identifiers like part numbers, error codes, legal citations and acronyms, which embeddings often blur together because they focus on overall meaning rather than literal string matches.
What is reciprocal rank fusion?
Reciprocal rank fusion (RRF) combines two or more ranked result lists by giving each document a score based on 1 divided by (a constant, usually around 60, plus its rank) in each list, then summing those scores. It avoids comparing incompatible score scales like BM25 scores and cosine similarity directly.
Does hybrid search slow down retrieval compared to vector search alone?
It adds a small amount of latency because two indexes are queried and then fused, but on a well-built pipeline this overhead is minor, typically tens of milliseconds, and is almost always worth it for the accuracy gain on professional documents.
Can I get hybrid search without building my own retrieval pipeline?
Yes. Managed platforms like Kopik build hybrid search (full-text plus semantic keyword expansion) into every knowledge base automatically when you upload documents, so you don't need to run separate keyword and vector infrastructure or tune a fusion method yourself.
Does hybrid search fix bad chunking?
No. If a chunk boundary splits a key fact from its context, neither keyword nor vector search can recover it reliably. Chunking strategy and hybrid search are complementary, not substitutes for each other.
Get the Kopik newsletter
New knowledge bases, RAG guides and product news. One email every week or two, unsubscribe in one click.
By subscribing you agree to receive our newsletter. We never share your address.