Under the hood

What Is a Vector Database? Embeddings, Similarity Search and Indexes Explained

The Kopik team9 min read

A vector database stores documents as lists of numbers (embeddings) that capture their meaning, so a search engine can find passages that are conceptually similar to a question rather than just matching keywords. It's a core building block of retrieval-augmented generation (RAG), but it isn't always necessary, and for many knowledge bases a well-tuned full-text or hybrid search works just as well, or better.

What is a vector database, in plain English?

Traditional databases store rows and columns, or in the case of search engines, words and their positions in a document. A vector database stores something different: numerical fingerprints called embeddings, usually hundreds or thousands of numbers per piece of text. Two passages that mean similar things end up with fingerprints that sit close together in this numerical space, even if they don't share a single word.

This matters because real questions rarely use the exact vocabulary of the source document. Someone might ask "how do I cancel my subscription" while the manual says "terminating your plan". A keyword search can miss this entirely. A vector database, by comparing meaning rather than spelling, can surface the right passage anyway. That's the whole point: it's a tool for finding things that are similar in meaning, not identical in wording.

Embeddings: turning text into numbers

An embedding model reads a chunk of text and outputs a vector, a fixed-length array of numbers, that represents its meaning in a mathematical space. The same model, applied consistently, will place "invoice payment terms" and "when do I need to pay my bill" relatively close together, and place both far from "recipe for lasagne".

  1. A document is split into chunks (paragraphs or sections), because embedding an entire 50-page PDF as one vector would blur too many distinct ideas together.
  2. Each chunk is passed through an embedding model, producing a vector, typically between 384 and 1,536 numbers long depending on the model.
  3. The vectors are stored alongside the original text and some metadata (source file, page number, date).
  4. At query time, the question is embedded with the same model, and the database finds the stored vectors nearest to it.

The quality of this whole pipeline depends heavily on chunking strategy and on using the same embedding model consistently. Mixing embeddings from different models, or re-chunking inconsistently, silently degrades search quality in ways that are hard to debug later.

How similarity search actually works

Once everything is a vector, finding "similar" items becomes a geometry problem: which stored vectors are closest to the query vector? Closeness is usually measured with cosine similarity (the angle between two vectors) or Euclidean distance (straight-line distance). Both answer the same practical question: how alike are these two pieces of text, numerically speaking?

With a handful of documents, you could compare the query against every single vector, brute force, and it would be instant. The problem appears at scale: comparing a query against ten million vectors, one by one, every single time, becomes slow and expensive. That's exactly the problem that vector indexes are built to solve.

Indexes: HNSW versus IVF

An index is a data structure that organises vectors so the database can find the nearest ones without checking every single one. This is the same trade-off every index makes, whether it's a book's table of contents or a database's B-tree: a bit of upfront organisation in exchange for much faster lookups later. The two most common families for vector search are HNSW and IVF.

HNSW vs IVF at a glance

IndexHow it worksStrengthsTrade-offs
HNSW (Hierarchical Navigable Small World)Builds a multi-layer graph linking similar vectors, then "hops" through it towards the nearest matchVery fast queries, excellent accuracy, handles updates reasonably wellHigher memory use; index build time grows with dataset size
IVF (Inverted File Index)Clusters vectors into buckets (via k-means), then only searches the closest bucketsLower memory footprint, scales well to very large datasetsSlightly lower accuracy unless tuned carefully; less efficient for frequent updates

In practice, most managed vector search tools default to HNSW because it offers the best balance of speed and accuracy for typical knowledge-base sizes (thousands to low millions of chunks). IVF, and hybrid variants like IVF-PQ (which also compresses vectors), tend to matter once you're operating at genuinely huge scale, hundreds of millions of vectors, where memory cost becomes the limiting factor. For most teams building an internal knowledge base or a customer-facing assistant, this choice is an implementation detail hidden behind whichever platform or library they use, not something to lose sleep over.

You rarely need to pick an index yourself

Unless you're running your own vector search infrastructure at serious scale, the index type is usually chosen for you by the platform. What matters more day to day is chunking quality, metadata filtering, and whether the system also supports keyword matching alongside semantic search.

Do you actually need a vector database for RAG?

Retrieval-augmented generation relies on fetching relevant passages before generating an answer, so the model works from real source material instead of guessing. We cover the full pattern in RAG as a Service: What It Is and When to Use It, but the short version for this question is: a vector database solves retrieval, and retrieval quality is what determines whether your RAG answers are accurate or embarrassing.

You genuinely benefit from vector (or hybrid) search when:

  • Users ask questions in natural language, with synonyms and paraphrasing the source documents don't use.
  • Your documents are long and conceptually dense, such as policies, contracts, technical manuals or scientific reports.
  • You need to search across many documents where exact terminology varies (regional spelling, abbreviations, jargon).
  • You want the system to find a relevant passage even when the question doesn't share a single keyword with it.

You can often get away with simpler full-text search when:

  • Users search for exact terms: product codes, reference numbers, names, error messages.
  • The corpus is small enough that a human could reasonably skim it (a handful of short documents).
  • Precision on exact phrasing matters more than conceptual recall, such as legal clause lookups by citation.

In reality, the strongest results usually come from combining both approaches rather than choosing one. Full-text search (classic keyword matching, often boosted with ranking algorithms like BM25) is excellent at precision for exact terms. Vector search is excellent at recall for paraphrased or conceptual queries. Hybrid search runs both and merges the results, so a query like "VAT invoice number format" can match the exact phrase VAT invoice while also surfacing a passage that discusses "tax reference formatting" without using those precise words.

This is the approach Kopik takes with the knowledge bases on kopik.io: when you upload documents, they're chunked and indexed for hybrid search combining full-text matching with semantic keyword expansion, so a question phrased loosely still finds the right passage, and an exact reference number still gets matched precisely. You don't need to configure an index or choose between HNSW and IVF yourself; you upload documents and the retrieval layer is handled for you.

Quick comparison

ApproachBest atWeak at
Full-text searchExact terms, codes, names, phrasesSynonyms, paraphrasing, conceptual questions
Vector (semantic) searchMeaning-based matching, paraphrased questionsExact codes or rare terms the embedding model hasn't seen distinctly
Hybrid searchBoth of the above, merged and re-rankedSlightly more complex to implement from scratch

A quick checklist before you build or buy

If you're deciding how to approach search for a new project, these questions will save time:

  1. How will people phrase their questions? Natural language suggests you need semantic search; exact-code lookups suggest full-text is enough.
  2. How large and how varied is the document set? A handful of short files rarely justifies a dedicated vector index.
  3. Does accuracy matter enough to need cited sources? If answers must be traceable to a specific passage, make sure whatever you choose supports citations, not just raw similarity scores.
  4. Who will maintain it? Running your own embedding pipeline, index, and re-ranking logic is a real engineering commitment; a managed knowledge-base platform removes most of that maintenance.
  5. Do you need this via an API or an AI agent? If so, check whether the platform exposes a REST API or an MCP server, so your tools (including Claude, Cursor or ChatGPT) can query it directly.

For teams that want to skip the infrastructure decisions entirely, creating a knowledge base on Kopik takes a few minutes: upload PDFs, Word documents, text or Markdown files, and the platform handles chunking, embeddings and hybrid indexing. You can keep the base private for internal use, or list it publicly in the catalogue and set a price per question, keeping 70% of what it earns.

Querying a vector-backed knowledge base

Once a knowledge base exists, there are generally three ways to use it: through a web interface for quick manual questions, through a REST API with API keys for integration into your own applications, or through an MCP server for AI agents and assistants that need to pull in grounded answers on demand. The developer documentation covers both the API and MCP routes in detail, including how answers are returned with cited passages or as raw passages only, depending on what your use case needs.

See hybrid search in action

Browse existing knowledge bases or create your own in minutes, no index tuning, no embedding pipeline to maintain.

Common pitfalls worth knowing about

A few mistakes come up repeatedly when teams build vector search themselves. Chunking too large (whole pages) dilutes the embedding and makes retrieval vague; chunking too small (single sentences) loses context and returns fragments that don't make sense on their own. Mixing embedding models over time, after switching providers, for instance, silently corrupts similarity comparisons, since vectors from different models aren't directly comparable. And relying on vector search alone, without any keyword fallback, tends to under-perform on exact lookups like order numbers, legal citations or product SKUs, exactly the cases where hybrid search earns its keep.

None of this means vector databases are overengineered or unnecessary, quite the opposite: for any knowledge base where questions are phrased in natural language and documents are more than a page or two, semantic retrieval is usually the difference between a useful assistant and a frustrating one. The practical lesson is simply to match the tool to the query pattern, and to lean on hybrid search by default when in doubt.

Frequently asked questions

Is a vector database the same thing as RAG?

No. A vector database is the retrieval component, the part that finds relevant passages. RAG is the broader pattern that combines retrieval with a language model generating an answer from those passages. You can read more in RAG as a Service: What It Is and When to Use It.

Do I need to choose between HNSW and IVF myself?

Usually not. Most managed platforms pick a sensible default, typically HNSW, and only very large-scale deployments (hundreds of millions of vectors) need to think carefully about index trade-offs.

Can I use full-text search instead of a vector database?

Yes, if your users search with exact terms, codes or names rather than natural-language questions, and your document set is small. For most knowledge bases, hybrid search (combining both) gives the best results.

What's the difference between embeddings and keywords?

Keywords match exact words. Embeddings represent meaning as numbers, so two passages with no words in common can still be recognised as similar if they discuss the same idea.

How big does my document set need to be before I need a vector database?

There's no fixed threshold, but as a rule of thumb, once you have more documents than someone could reasonably skim by eye, or once users start phrasing questions in their own words rather than copying terminology from the source, semantic or hybrid search starts paying off.

Does Kopik use vector search or full-text search?

Both. Knowledge bases on Kopik are indexed for hybrid search, combining full-text matching with semantic keyword expansion, so exact terms and paraphrased questions are both handled without any manual configuration.

Get the Kopik newsletter

New knowledge bases, RAG guides and product news. One email every week or two, unsubscribe in one click.

By subscribing you agree to receive our newsletter. We never share your address.