What Is a Vector Database? Embeddings, Similarity Search, and When You Need One
A vector database stores pieces of content as embeddings, arrays of numbers that capture meaning, so a search engine can find the passages closest to a question even if they don't share a single word with it. It's the component that lets retrieval-augmented generation (RAG) systems fetch the right context before an AI model writes an answer. Whether you need one, or whether full-text or hybrid search already covers your case, depends on how your users phrase questions and how large and varied your documents are.
What a Vector Database Actually Stores
A traditional database indexes exact values: a customer ID, a date, a string you can match with an equals sign or a keyword. A vector database indexes embeddings, dense numeric vectors (typically a few hundred to a few thousand dimensions) produced by an embedding model from a chunk of text, an image, or even audio. Two chunks that mean similar things end up with vectors that point in similar directions in that high-dimensional space, even if the wording is completely different. Instead of asking "does this row equal X," a vector database answers "which stored vectors are closest to this query vector," which is exactly the question you need answered when a user asks something in their own words and the right answer lives in a PDF written in different words entirely.
Embeddings in Plain English
An embedding model reads a sentence like "the tenant must give 30 days' notice before moving out" and outputs, say, 768 floating-point numbers. Feed it "how much notice does a renter need to provide," and despite sharing almost no vocabulary, the two sentences land close together in vector space because the model was trained to encode meaning rather than surface form. This is the mechanism that makes semantic search possible, and it's covered in more depth, including how chunks are created before they're embedded, in our guide to chunking, indexing and retrieval.
How Similarity Search Works
Once a document collection is embedded, answering a query is a three-step process: embed the query with the same model, compare that query vector against every stored vector using a distance metric, and return the top-k closest matches. The distance metric matters because it defines what "close" means mathematically.
- Cosine similarity: measures the angle between two vectors, ignoring their length. The most common choice for text embeddings because it focuses purely on direction, i.e., meaning.
- Dot product: similar to cosine but also weighted by vector magnitude; used when the embedding model was trained with magnitude carrying useful signal.
- Euclidean distance (L2): straight-line distance between two points; common in image and some legacy embedding setups.
In theory you could compute these distances against every vector in the collection (a brute-force or "flat" search) and it would be perfectly accurate. The problem is speed: comparing a query against a few hundred thousand vectors one by one gets slow, and comparing it against tens of millions becomes impractical in real time. That's where indexes come in.
Indexing Strategies: HNSW, IVF and Flat Search
An index trades a small amount of accuracy for a large amount of speed by organizing vectors so the search only has to look at a promising subset instead of the whole collection. The two names you'll see most often in vector database documentation are HNSW (Hierarchical Navigable Small World) and IVF (Inverted File Index), alongside plain flat/brute-force search for small collections.
Comparing vector index types
| Index | How it works | Best for | Trade-off |
|---|---|---|---|
| Flat / brute-force | Compares the query to every vector directly | Small collections (under ~50k vectors) | Perfectly accurate but slow at scale |
| HNSW | Builds a layered graph of neighbors for fast navigation toward close matches | Low-latency search on large collections | High recall and speed, but more memory and slower to build the index |
| IVF | Clusters vectors into buckets and only searches the most relevant buckets | Very large collections where memory is constrained | Faster to build and lighter on memory, but needs tuning of cluster count |
| IVF + PQ | Adds product quantization to compress vectors | Massive collections (tens of millions+) with tight memory budgets | Smallest footprint, but lowest accuracy of the group |
You rarely choose an index by hand
Most managed RAG and vector search platforms pick a sensible default index for you based on collection size and automatically rebuild or re-tune it as the base grows. Understanding HNSW versus IVF helps you reason about latency and recall trade-offs, but it's rarely a decision you need to make manually day to day.
Vector Database vs Full-Text vs Hybrid Search
Semantic search isn't automatically better than keyword search, it's better at a specific kind of query. Full-text search (the BM25/TF-IDF family used by classic search engines) is excellent at exact matches: product SKUs, legal citations, error codes, acronyms, names, or any query where the precise term matters more than the general idea. A vector database can actually underperform plain full-text search on these because embedding models tend to smooth over exact tokens in favor of overall meaning. Semantic search wins on paraphrased questions, conceptual queries, and cross-lingual matching, where the user's words don't overlap with the document's words at all.
Hybrid search runs both in parallel, full-text and vector similarity, then merges and re-ranks the results. It's the approach that holds up best across real-world question sets because users mix precise terms ("section 4.2", "invoice #4471") with loose, conversational phrasing in the same session.
When Full-Text or Hybrid Search Is Enough
- Your documents are short, structured, and queries tend to reuse the same vocabulary (FAQs, glossaries, reference tables). Full-text alone often suffices.
- Your users search with codes, IDs, names, or exact phrases more often than open-ended questions. Weight full-text higher in the blend.
- Your corpus mixes contract language, numbered clauses, and conversational questions from non-experts. This is the classic case for hybrid search, and it's the default most RAG-as-a-service platforms, including Kopik, ship with rather than pure vector search.
- You're searching a single, small document under a few thousand words: a dedicated vector index adds overhead without much benefit; a straightforward keyword search or even manual review may be faster.
When You Actually Need a Vector Database for RAG
If you're building or evaluating a retrieval-augmented generation system, a few signals tell you a proper vector database (rather than ad hoc keyword search) is worth the investment:
- Your knowledge base spans hundreds of documents or more, and relevant passages are scattered across files with inconsistent terminology.
- Users ask natural-language questions rather than typing search terms, so the system needs to match intent, not just words.
- You need results in multiple languages or phrasings of the same question to surface the same source passage.
- You're feeding retrieved passages to a language model and need the top few chunks to actually be relevant, since irrelevant context degrades answer quality and wastes tokens.
- The corpus changes often enough that you need incremental indexing rather than rebuilding a search index from scratch each time.
If none of those apply, you may not need a vector database at all, a well-tuned full-text index might cover your use case with far less setup. For a deeper comparison of retrieval approaches versus adjusting the model itself, see RAG vs fine-tuning.
Running Vector Search Without Managing Infrastructure Yourself
Standing up a vector database from scratch means choosing an embedding model, picking an index type, deciding on a hybrid search strategy, handling re-indexing as documents change, and exposing it all through an API your application can call. That's a reasonable amount of plumbing for a team that just wants accurate answers grounded in its own documents. Kopik handles that layer directly: you upload PDFs, Word files, Markdown or plain text, and the platform extracts, chunks, embeds and indexes them for hybrid search automatically, no index tuning required. You can read more about preparing source material so retrieval actually performs well in our guide on preparing documents for RAG, and about the end-to-end process in building a knowledge base from your documents.
Once a base is built, you can query it three ways: directly on the site, through a REST API with your own API keys, or through an MCP server usable from Claude, Cursor, ChatGPT and other MCP-compatible clients, detailed in connecting a knowledge base with MCP. Bases can stay private for internal use, as explored in a private knowledge base as your own RAG, or be published to the public catalogue where you set a price per question and keep 70% of what it earns.
Try hybrid vector search on your own documents
Upload a handful of PDFs or Word files and see embeddings, indexing and hybrid search working together without any infrastructure to configure.
A Quick Mental Model to Keep
Full-text search finds documents that contain your words. A vector database finds documents that match your meaning. Hybrid search finds documents that match either, ranked together. For most real knowledge bases, mostly ones mixing policy language, technical terms and everyday questions, hybrid is the safer default, and that's exactly why it's treated as the baseline rather than an add-on in most modern RAG-as-a-service tools, including the REST API and MCP integration that Kopik exposes for developers.
Frequently asked questions
Is a vector database the same thing as a RAG system?
No. A vector database is one component, the storage and retrieval layer, inside a RAG system. RAG also needs document chunking, an embedding model, and a language model that turns retrieved passages into an answer.
Do I need to choose between HNSW and IVF myself?
Usually not. Managed platforms pick an index automatically based on collection size and re-tune it as data grows. You only need to reason about HNSW versus IVF if you're self-hosting a vector database and tuning it by hand.
Can a vector database replace full-text search entirely?
It can for pure semantic queries, but it tends to underperform on exact matches like codes, IDs, or precise legal citations. Most production systems combine both through hybrid search rather than replacing one with the other.
What's the difference between embeddings and a vector database?
Embeddings are the numeric representations of your content, produced by an embedding model. A vector database is the system that stores those embeddings and efficiently finds the closest ones to a query vector.
How big does my document collection need to be before a vector database matters?
There's no fixed threshold, but if you're under a few thousand short documents with consistent vocabulary, full-text search alone often performs just as well with far less setup. Vector search starts paying off as phrasing diversity and collection size grow.
Does hybrid search cost more than plain vector or full-text search?
It requires running two retrieval methods and merging results, which adds some computation, but on a managed platform this is typically abstracted away and billed the same as a single query, not as two separate searches.
Get the Kopik newsletter
New knowledge bases, RAG guides and product news. One email every week or two, unsubscribe in one click.
By subscribing you agree to receive our newsletter. We never share your address.