Guide

Embeddings Explained for Non-Specialists: How Text Becomes Numbers

The Kopik team9 min read

An embedding is a list of numbers that represents the meaning of a piece of text, so that texts with similar meanings get similar numbers. Once text is turned into these vectors, a computer can compare meanings with simple arithmetic, which is what powers semantic search: finding “Can I get my money back?” when the document says “refund policy”. This guide explains embeddings without jargon, walks through a cosine similarity example you can check on a calculator, and covers the limits that matter before you build anything on them.

What is an embedding, in plain English?

Computers are very good at comparing numbers and very bad at comparing meanings. Ask a classic search engine whether “car” and “automobile” are related and, unless someone typed in a synonym list, it has no idea: the letters are different, so the words are different. Embeddings solve this by translating text into numbers in a way that preserves meaning.

Concretely, an embedding model reads a word, a sentence or a paragraph and outputs a fixed-length list of numbers, for example 768 or 1,024 of them. That list is called a vector. The model is trained so that texts which mean similar things end up with vectors that point in similar directions, and texts about unrelated things end up far apart.

An analogy: a map of meaning

Picture a giant map where every sentence ever written has a pin. Sentences about returning a purchase cluster in one neighborhood, sentences about parking in another, sentences about tax deductions somewhere else. An embedding is simply the GPS coordinates of a sentence on that map. To find documents related to a question, you drop a pin for the question and look at which pins are nearby.

The only twist is that this map does not have two dimensions like a road atlas. It has hundreds or thousands, because meaning has many more directions than north and east: topic, tone, tense, domain, formality and countless subtler features. Humans can't picture that space, but the arithmetic works exactly the same way as in two or three dimensions.

How text becomes vectors

You don't need to understand neural networks to grasp the process. It happens in three stages:

  1. Tokenization. The text is cut into small units called tokens: whole words, pieces of words or punctuation. “Unbelievable” might become “un”, “believ” and “able”.
  2. Encoding. A trained neural network reads the tokens in context. Context matters: “bank” in “river bank” and “bank account” gets a different representation in modern models.
  3. Pooling. The network's internal representations are combined into one vector for the whole text, often normalized so its length is 1. That final vector is the embedding.

Where does the model's sense of meaning come from? From training on very large amounts of text, using the old linguistic insight that words appearing in similar contexts tend to mean similar things. An influential early example was word2vec, described in the 2013 paper Efficient Estimation of Word Representations in Vector Space, which produced one vector per word. Modern sentence embedding models go further and encode whole passages, so a question and the paragraph that answers it can land close together even when they share few words.

The dimensions don't have names

It is tempting to imagine that dimension 12 means “money” and dimension 40 means “legal”. In real models, individual numbers are not human-readable. Meaning is spread across all of them at once, which is one reason embeddings are hard to debug.

Cosine similarity, with a worked example

Once two texts are vectors, how do you measure how close they are? The most common measure is cosine similarity: it looks at the angle between the two vectors. If they point in the same direction, the score is 1. If they are unrelated (at right angles), the score is 0. Scores can go down to -1 for opposite directions, although with text embeddings most scores sit between 0 and 1.

The formula is short: multiply the vectors element by element and add the results (the dot product), then divide by the product of their lengths. Let's do it by hand with toy vectors of only three dimensions. To make the example readable, pretend the three numbers mean “about money”, “about returning a purchase” and “about a location”. Real embeddings don't work with labeled dimensions, but the math is identical.

Three toy embeddings

TextMoneyReturnsLocation
A: “How do I get a refund?”0.90.80.1
B: “Can I get my money back?”0.80.90.2
C: “Where can I park?”0.10.20.9
  1. Dot product A·B = 0.9×0.8 + 0.8×0.9 + 0.1×0.2 = 0.72 + 0.72 + 0.02 = 1.46.
  2. Lengths. |A| = √(0.81 + 0.64 + 0.01) = √1.46 ≈ 1.208. |B| = √(0.64 + 0.81 + 0.04) = √1.49 ≈ 1.221.
  3. Cosine A,B = 1.46 ÷ (1.208 × 1.221) ≈ 1.46 ÷ 1.475 ≈ 0.99. Very similar, even though the two questions share almost no words.
  4. Dot product A·C = 0.9×0.1 + 0.8×0.2 + 0.1×0.9 = 0.09 + 0.16 + 0.09 = 0.34. |C| = √(0.01 + 0.04 + 0.81) = √0.86 ≈ 0.927.
  5. Cosine A,C = 0.34 ÷ (1.208 × 0.927) ≈ 0.34 ÷ 1.120 ≈ 0.30. Much less similar.

That is the whole trick behind semantic search. Embed every passage of your documents once, embed the user's question at search time, compute the similarity between the question and every passage, and return the passages with the highest scores. With millions of passages, comparing against every one becomes slow, which is why specialized indexes and vector databases exist: they find the nearest vectors approximately but very quickly.

Multilingual embeddings: one map for many languages

Some embedding models are trained on many languages at once, often with pairs of translated sentences. The result is a shared map where “refund policy”, “politique de remboursement” and “Rückerstattungsrichtlinie” land near each other. That makes cross-language search possible: a user asks in Spanish and finds a passage written in English.

This is genuinely useful for US companies with multilingual customers or international documentation, but quality is uneven. Models generally perform best in the languages that dominated their training data, and less common languages, regional terms or specialist vocabulary can land in the wrong neighborhood. If cross-language search matters to you, test it with real questions in each language rather than assuming it works.

Embeddings are powerful, but they are not magic, and several failure modes surprise people the first time:

  • Exact identifiers. Part numbers, statute sections like “29 CFR 1910.134”, invoice references or error codes carry little “meaning” for a model. Semantic search can return something vaguely related instead of the exact match a keyword search would find instantly.
  • Negation and nuance. “Covered by the warranty” and “not covered by the warranty” are about the same topic, so their vectors can be close. Similarity measures topic proximity, not agreement.
  • Jargon and new terms. A product name invented last month or an industry acronym may be poorly represented if the model never saw it in training.
  • Long passages get blurry. One vector for a long chunk averages many ideas together, so a specific detail buried inside may not pull the score up enough.
  • Vectors are tied to their model. Vectors from two different embedding models are not comparable. Switching models means re-embedding every document.
  • Hard to explain. When a keyword search misses, you can see why. When a vector search returns an odd result, there is rarely a clear reason to point to.
  • Privacy. Embeddings are not anonymization. Research has shown that text can be partly reconstructed from its embedding, so vectors of sensitive documents (think HIPAA-covered records) deserve the same protection as the text itself.

None of these limits rule embeddings out. They explain why many production systems use hybrid search, combining semantic similarity with classic full-text search, so that exact terms and paraphrases are both covered.

Do you actually need embeddings?

If you are building a RAG system (an AI that answers from your documents), embeddings are one way to retrieve passages, not the only one. The right choice depends on your documents and how people ask questions.

Embeddings or full-text search?

SituationEmbeddings helpFull-text search is enough
Users phrase things very differently from the documentsYesOnly with query expansion
Documents full of codes, references, section numbersPartlyYes, and more precisely
You must explain why a result came backHardYes
Very large, loosely structured corpusYesPossible, with good ranking
You want minimal infrastructureNeeds a model and a vector indexBuilt into common databases

There is also a middle path. Instead of embedding everything, you can ask a language model to rewrite the user's question into several keyword variants and synonyms, then run a full-text search. That captures much of the “money back” versus “refund” benefit without storing a single vector. It is the approach Kopik takes: Kopik does not use embeddings or a vector database. Its search is hybrid in a different sense, combining full-text search with semantic expansion of the keywords, and it works across languages because the question is expanded into the language of the documents.

If you do go the embedding route, budget for more than the model calls: storage, indexing and re-embedding when documents or models change all add up. Our breakdown of vector database pricing covers what you actually pay for, and how a RAG pipeline works under the hood shows where retrieval fits among chunking and answer generation.

A quick checklist before you rely on embeddings

  • Write 20 to 30 real questions and check that the right passage appears in the top results.
  • Include questions with exact codes, names and numbers, not just fuzzy ones.
  • Test negations and near misses (“eligible” versus “not eligible”).
  • If you serve several languages, test each one separately.
  • Record which embedding model you used, so you know when a re-index is needed.
  • Treat stored vectors with the same security controls as the source documents.

If you would rather skip the infrastructure, Kopik lets you upload PDFs, Word files, text or Markdown and get a knowledge base that answers with cited passages, on the website, through a REST API or from MCP clients such as Claude, Cursor or ChatGPT. Creating a base is free.

Search your documents without managing vectors

Upload your files and get a knowledge base that finds the right passages and answers with citations. Creating a base is free.

Frequently asked questions

What is the difference between an embedding and a vector?

A vector is just an ordered list of numbers. An embedding is a vector produced by a model to represent the meaning of something, such as a sentence or an image. Every embedding is a vector, but not every vector is an embedding.

What is a good cosine similarity score?

There is no universal threshold. Typical score ranges differ from one embedding model to another, so a 0.8 can be excellent with one model and mediocre with another. The reliable approach is to compare scores relative to each other for the same query and to calibrate any cut-off on your own test questions.

Are embeddings the same as a language model?

No. An embedding model turns text into a vector and stops there; it does not write answers. A language model generates text. In a RAG system, an embedding model (or a full-text index) finds the passages, and a language model writes the answer from them.

Can embeddings work across languages?

Yes, with multilingual embedding models trained to place translations close together. Quality varies by language and domain, so test with real questions in each language you support before relying on it.

Do I need a vector database to use embeddings?

Not always. For a few thousand passages, comparing the question against every vector is fast enough, and extensions such as pgvector add vector search to PostgreSQL. Dedicated vector databases become useful at large scale or when you need advanced filtering and fast approximate search.

Does Kopik use embeddings?

No. Kopik indexes passages with full-text search and expands each question into keywords and synonyms in the documents' language before searching. That handles paraphrases and cross-language questions without storing vectors.

Get the Kopik newsletter

New knowledge bases, RAG guides and product news. One email every week or two, unsubscribe in one click.

By subscribing you agree to receive our newsletter. We never share your address.