Vector Database Pricing: What You Actually Pay For
Vector database pricing rarely comes down to a single number. What you actually pay for is memory and storage (driven by how many vectors you keep and how many dimensions each one has), compute for queries and indexing, and the number of replicas you run for availability. On top of that sit costs that never appear on the database invoice: generating embeddings, re-indexing when you change models, and the engineering time to keep it all running. This guide breaks down each driver, shows how to estimate memory with simple arithmetic, and explains when you don't need a vector database at all.
First, what are you storing?
A vector database stores embeddings: lists of numbers that represent the meaning of a piece of text, an image or another object. Each document you index is usually split into chunks, and each chunk gets its own vector. If you need a refresher on how embeddings and similarity search work, start with What Is a Vector Database?, then come back here for the money side.
The key point for cost is that you are not paying per document. You are paying per vector, and the number of vectors depends on how you chunk. A 50-page PDF can become 50 vectors or 500, depending on chunk size and overlap. That single design choice, which most teams make once and forget, can change your bill by an order of magnitude.
The four cost drivers of a vector database
Whether you self-host an open-source engine, add an extension to Postgres, or use a managed cloud service, the same physical resources sit underneath. Vendors package them differently (per pod, per gigabyte, per read unit, per hour), but the underlying levers are always these four.
1. Number of vectors
This is the base of every calculation. Count your chunks, not your files. Then add growth: if your corpus doubles every year, your index will too, and some index types slow down or need more memory as they grow.
2. Dimensions and precision
Each vector has a fixed number of dimensions set by the embedding model: common sizes include 384, 768, 1,024, 1,536 and 3,072. Each dimension is typically stored as a 32-bit float, which is 4 bytes. So the raw size of one vector is simply dimensions × 4 bytes. A 1,536-dimension vector weighs about 6 KB before any index overhead.
3. Query volume and latency targets
Every search compares your query vector against part of the index. Approximate nearest neighbor (ANN) indexes such as HNSW, described in the original HNSW paper, make this fast by navigating a graph instead of scanning every vector, but they need the graph in memory to be fast. Higher queries per second, stricter latency targets and metadata filters all push you toward more CPU and more RAM.
4. Replicas and environments
One copy of the index is a single point of failure. Production setups usually run at least two replicas, sometimes three, and each replica holds the full index in memory. Add a staging environment and maybe a separate index per customer for isolation, and the same data can be paid for four or five times over.
Estimating embedding storage cost with simple arithmetic
You can estimate raw vector memory without any price list. The formula is: number of vectors × dimensions × 4 bytes (for float32). Here is what that gives for a few typical sizes.
Raw float32 vector size, before index overhead and replicas
| Vectors | Dimensions | Raw size |
|---|---|---|
| 100,000 | 768 | about 0.3 GB |
| 1,000,000 | 384 | about 1.5 GB |
| 1,000,000 | 768 | about 3.1 GB |
| 1,000,000 | 1,536 | about 6.1 GB |
| 10,000,000 | 1,536 | about 61 GB |
Take the 1,000,000 × 1,536 line: 1,000,000 × 1,536 × 4 = 6,144,000,000 bytes, or roughly 6.1 GB. That is only the beginning. A graph index stores neighbor links for every vector, you usually keep the chunk text and metadata alongside, and then you multiply by the number of replicas. A realistic planning figure is the raw size, plus a healthy margin for the index and metadata, times your replica count.
Quantization is the biggest lever on memory
Storing each dimension as an 8-bit integer instead of a 32-bit float divides raw size by 4, and binary quantization (1 bit per dimension) divides it by 32. You trade some recall for that saving, so measure retrieval quality on your own questions before and after. Engines such as pgvector document the vector types and index options they support.
The hidden costs that never show up on the invoice
Teams that budget only for the database are usually surprised. The bigger line items often sit outside it.
- Embedding generation. Every chunk must pass through an embedding model at ingestion, and every user question must be embedded at query time. Whether you call an API or run a model on your own GPUs, this is a real and recurring cost.
- Re-indexing when you change models. Vectors from two different embedding models are not comparable. Switch models, or even change dimensions, and you must re-embed and re-index the entire corpus, often while keeping the old index live in parallel.
- Re-chunking. If retrieval quality is poor and you change your chunking strategy, you re-embed everything again. Expect to do this more than once in the first months.
- Updates and deletes. Documents change. Keeping vectors in sync with source files requires a pipeline, and some index types degrade with many deletions until they are rebuilt.
- Data egress and backups. Moving vectors between regions or clouds, and keeping snapshots, adds storage and transfer charges.
- Engineering and on-call time. Tuning index parameters, monitoring recall, upgrading versions and handling incidents is work that someone on your team has to own.
For regulated data, add compliance work too. If your chunks contain protected health information covered by HIPAA, or customer data in scope for a SOC 2 audit, your vector store becomes one more system to secure, log and document. Embeddings are derived from the original text and should be treated as sensitive data, not as anonymized output.
When you don't need a vector database
A dedicated vector database is the right tool for large, fast-changing collections with high query volume. Many projects are nowhere near that. Before you sign up for anything, check whether one of these situations describes you.
- Your corpus is small. A few thousand chunks fit comfortably in memory, and a brute-force similarity scan is fast enough. A vector column in the database you already run may be all you need.
- Your questions are about exact terms. Product codes, regulation numbers, error messages and names are often found better by full-text search than by pure semantic similarity. Hybrid search, combining keywords with meaning, frequently beats vectors alone.
- You already run Postgres. An extension like pgvector keeps vectors next to your relational data, with one backup policy and one access-control model.
- You only need answers, not infrastructure. If the goal is for people or AI assistants to query a set of documents, a managed retrieval service removes the database, the embedding pipeline and the re-indexing problem from your plate.
The last case is where RAG as a service fits. On Kopik, for example, you upload PDF, Word, text or Markdown files for free; Kopik extracts, chunks and indexes them for hybrid search (full-text plus semantic expansion of keywords). You don't size a cluster or choose dimensions: questions are paid for with prepaid credits, or through a chat subscription at €12 per month with a usage gauge. If you want to see what happens to your files under the hood, the guide to chunking, indexing and retrieval walks through each step.
A checklist before you compare vendors
Pricing pages are hard to compare because each vendor meters something different. Translate every offer back to the same inputs, and you will see the real differences.
- Count your chunks today and estimate them in 12 months.
- Pick your embedding dimensions and compute raw size with vectors × dimensions × 4 bytes.
- Decide whether quantization is acceptable after testing recall on 50-100 real questions.
- Set your replica count and the number of environments (production, staging, per-tenant).
- Estimate queries per day and peak queries per second, including AI agents that may call search in loops.
- Budget embedding costs for the initial load, for daily updates, and for one full re-index per year.
- List the security requirements (encryption, audit logs, data residency) and check which tier includes them.
- Price your own time: who maintains the pipeline, and what happens when they leave?
Watch usage-based pricing with AI agents
Agents connected through MCP can issue many searches for a single user request. If you pay per query, set a ceiling. Kopik's MCP tools, for instance, accept a maxPriceCents parameter that refuses a call at no charge if the base costs more than you allow.
Skip the vector database bill
Turn your PDFs and Word files into a searchable knowledge base without sizing clusters or managing embeddings. Creating a base is free.
Frequently asked questions
How much memory does 1 million embeddings need?
Multiply vectors by dimensions by 4 bytes for float32. One million 1,536-dimension vectors take about 6.1 GB raw, and one million 768-dimension vectors about 3.1 GB. Add index overhead, metadata and every replica you run.
Is a vector database expensive?
It depends far more on your volume, dimensions, query rate and replica count than on the vendor. Small corpora can run in a database you already have at almost no extra cost; large, high-traffic indexes held in RAM are where bills grow.
Do smaller embeddings reduce cost?
Yes. Halving dimensions halves raw storage and speeds up comparisons. Quantization reduces it further. The trade-off is retrieval quality, so test on your own documents and questions before switching.
Why do I have to re-index when I change embedding models?
Each model maps text into its own vector space. A query embedded with model B cannot be compared meaningfully with documents embedded with model A, so the whole corpus must be re-embedded and the index rebuilt.
Can I do RAG without a vector database?
Yes. Full-text search, hybrid search in an existing database, or a managed retrieval service can all power grounded answers. Kopik, for example, indexes uploaded documents for hybrid search and exposes them through the web, a REST API and an MCP server.
Get the Kopik newsletter
New knowledge bases, RAG guides and product news. One email every week or two, unsubscribe in one click.
By subscribing you agree to receive our newsletter. We never share your address.