Business

Vector Database Cost: What You Are Really Paying For

The Kopik team6 min read

The cost of a vector database is set by four levers: how many vectors you store, how many dimensions each one has, how many queries you serve, and how many replicas you keep running. Around those sit costs that never appear on the database invoice, chiefly generating embeddings and re-indexing your whole corpus when you change model. This guide explains each lever, shows how to size memory with simple arithmetic rather than price lists, and helps you decide whether you need a dedicated vector database in the first place.

Why vector database pricing is so hard to compare

Providers meter different things. One charges by the hour for a sized pod, another by stored gigabyte plus read and write units, a third bundles everything into a cluster tier. Self-hosting an open-source engine swaps the invoice for virtual machines, disks and staff time. The only way to compare them fairly is to work out your own physical requirements first, then translate each offer back to those numbers.

If embeddings and similarity search are still a little hazy, read What Is a Vector Database? first. The short version: each chunk of text becomes a list of numbers, and search means finding the stored lists closest to the one computed from the question.

The four levers behind the bill

Whatever the packaging, each of these levers maps to something physical: RAM, CPU, disk or network. Work them out for your own corpus before reading any pricing page, and you will be able to tell a cheap-looking entry tier from one that genuinely fits your workload.

Vector count, which depends on chunking

You pay per vector, not per document. A 40-page policy manual might become 40 vectors with large chunks or 400 with small, overlapping ones. Chunk size is a quality decision, but it is also a cost decision, and it is worth making consciously.

Dimensions and numeric precision

The embedding model fixes the number of dimensions; 384, 768, 1,024, 1,536 and 3,072 are all common. Stored as 32-bit floats, each dimension takes 4 bytes, so a single 1,536-dimension vector is about 6 KB on its own.

Query volume, latency and filters

Approximate nearest neighbour indexes, such as the HNSW graphs introduced in this paper, answer quickly because they walk a graph rather than scanning every vector. To stay quick, that graph generally needs to live in RAM. More queries per second, tighter latency targets and heavy metadata filtering all call for more memory and CPU.

Replicas, environments and tenants

A production index usually runs on two or three replicas for resilience, each holding a full copy. Add a staging environment, and perhaps one index per client to keep data separated, and the same vectors are paid for several times over.

Sizing embedding storage with back-of-envelope maths

You do not need a tariff to estimate memory. Use: vectors × dimensions × 4 bytes. The figures below are raw float32 sizes, before index overhead, metadata or replicas.

Raw vector memory in float32

VectorsDimensionsRaw size
250,000768about 0.77 GB
1,000,000384about 1.5 GB
1,000,0001,024about 4.1 GB
1,000,0001,536about 6.1 GB
5,000,0003,072about 61 GB

Worked example: 1,000,000 × 1,536 × 4 = 6,144,000,000 bytes, roughly 6.1 GB. Now add the graph's neighbour links, the stored chunk text and metadata, and multiply by three replicas. A corpus that looked like 6 GB on paper quickly needs a few tens of gigabytes of RAM across your cluster.

Quantise before you scale up

Storing each dimension as an 8-bit integer cuts raw size by a factor of 4; binary quantisation (1 bit per dimension) cuts it by 32. Recall usually drops a little, so benchmark on your own questions. Open-source engines such as pgvector document which vector types and indexes they support.

The hidden costs: embeddings, re-indexing and people

  • Embedding every chunk. Ingestion runs each chunk through an embedding model, and every question is embedded at query time too. Whether you pay an API or run your own GPUs, the meter is running.
  • Re-indexing after a model change. Vectors from different models live in different spaces and cannot be mixed. Changing model means re-embedding and rebuilding the entire index, usually with the old one still serving traffic.
  • Re-chunking after quality reviews. Poor answers often lead to a new chunking strategy, and that means another full re-embed.
  • Keeping in sync. Updated and deleted documents need a pipeline, and some indexes degrade after many deletions until rebuilt.
  • Backups and transfers. Snapshots and cross-region copies add storage and egress charges.
  • Operations. Tuning, monitoring recall, upgrades and incident response all need an owner.

There is a compliance cost as well. Under UK GDPR, embeddings derived from personal data are still personal data in practice, because they come from identifiable text and sit next to it. Your vector store needs the same access controls, retention rules and records of processing as your other systems. The ICO's guidance on AI and data protection is a sensible starting point.

When you can do without a vector database

  • Small collections. A few thousand chunks fit easily in memory and a brute-force scan is quick enough.
  • Exact-term questions. Policy numbers, part codes, legislation references and names are often better found by full-text search; hybrid search that blends keywords and meaning frequently beats vectors alone.
  • An existing Postgres estate. A vector extension keeps embeddings beside your relational data, under one backup and permissions model.
  • You want answers, not infrastructure. If the aim is to let colleagues or AI assistants query documents, a managed retrieval service removes the database, the embedding pipeline and the re-indexing headache.

That final case is what RAG as a service means. With Kopik, you upload PDF, Word, text or Markdown files and create a base for free; Kopik extracts, chunks and indexes them for hybrid search (full-text plus semantic expansion of keywords). There is no cluster to size and no dimensions to choose. Questions are paid with prepaid credits, or via a chat subscription at €12 a month with a usage gauge. To see how a question travels through such a pipeline, read How RAG Works, Step by Step.

A costing checklist before you talk to suppliers

  1. Count today's chunks and forecast the figure in 12 months.
  2. Choose your embedding dimensions and compute vectors × dimensions × 4 bytes.
  3. Test whether quantisation is acceptable on 50-100 real questions.
  4. Fix replica count and environments (production, staging, per client).
  5. Estimate daily queries and peak queries per second, allowing for AI agents that search repeatedly.
  6. Budget embeddings for the first load, for ongoing updates and for at least one full re-index a year.
  7. Check where data is hosted, what encryption and audit logging are included, and in which tier.
  8. Put a figure on your own team's time to run and maintain it.

Agents multiply queries

An assistant connected over MCP may run several searches for one request. With pay-per-query pricing, set a cap. Kopik's MCP tools accept a maxPriceCents parameter that refuses a call at no charge when a base costs more than your limit.

Get searchable documents without the infrastructure

Upload your files and query them from the web, a REST API or an MCP client. Creating a knowledge base is free.

Frequently asked questions

How do I calculate embedding storage cost?

Start with memory: number of vectors × dimensions × 4 bytes for float32. Then add index overhead and metadata, multiply by replicas and environments, and apply your provider's or your hosting's rate to the total.

How big is one million 1,536-dimension vectors?

About 6.1 GB raw (1,000,000 × 1,536 × 4 bytes). The real footprint is larger once you include the index graph, stored text, metadata and every replica.

Does changing embedding model mean starting again?

Yes. Each model produces its own vector space, so documents must be re-embedded and the index rebuilt before queries from the new model return sensible results.

Are embeddings personal data under UK GDPR?

If they are derived from text about identifiable people and stored alongside it, treat them as personal data: apply access controls, retention rules and include the vector store in your records of processing.

Can I build RAG without a vector database?

Yes. Full-text or hybrid search in a database you already run, or a managed service such as Kopik, can retrieve passages for grounded answers without a dedicated vector store.

Get the Kopik newsletter

New knowledge bases, RAG guides and product news. One email every week or two, unsubscribe in one click.

By subscribing you agree to receive our newsletter. We never share your address.