Text Embeddings Explained: A Plain-English Guide for Non-Specialists
Text embeddings are lists of numbers that capture what a piece of text means, so that two passages about the same thing receive similar numbers even when they use different words. That is what lets semantic search match “My train was late, can I get money back?” with a page titled “Delay compensation”. Below you will find a jargon-free explanation, a cosine similarity calculation worked through by hand, a word on multilingual models, and the limits worth knowing before your organisation commits to the approach.
What are embeddings, without the maths?
Traditional search compares strings of letters. If your staff handbook says “annual leave” and an employee types “holiday allowance”, a basic keyword search shrugs: the words don't match, so nothing comes back. Embeddings take a different route. A trained model reads the text and produces a fixed-length list of numbers (a vector), typically several hundred to a few thousand numbers long. The model has learnt to give texts with related meanings vectors that sit close together.
So “annual leave” and “holiday allowance” end up as neighbours, while “annual accounts” lands somewhere else entirely, despite sharing the word “annual”. Search then becomes a question of geometry: which stored passages are nearest to the question?
An analogy: postcodes for meaning
Think of an embedding as a postcode for an idea. Two houses with nearly identical postcodes are on the same street; two with very different postcodes are at opposite ends of the country. An embedding model assigns every sentence a postcode on a vast map of meaning, and semantic search simply looks for the houses on the same street as your question.
The map is far stranger than the Ordnance Survey, mind you. Instead of two dimensions, it has hundreds or thousands, each capturing some faint aspect of meaning. Nobody can visualise it, but distances and angles are calculated in exactly the same way as on paper.
From words to vectors: what happens inside
- The text is split into tokens. These are words or fragments of words. A long word such as “decarbonisation” may be split into several pieces.
- A neural network reads the tokens in context. “Charge” in “electric vehicle charge point” and “criminal charge” gets treated differently, because modern models look at the surrounding words.
- The result is pooled into a single vector. The model combines its internal workings into one list of numbers for the whole passage, usually scaled to a length of 1. That list is the embedding.
The model learns all this from enormous quantities of text, relying on a simple observation: words used in similar contexts tend to have similar meanings. The idea was popularised by word2vec, introduced in Efficient Estimation of Word Representations in Vector Space in 2013, which gave each word its own vector. Today's sentence embedding models encode whole paragraphs, so a question and the paragraph that answers it can sit side by side even with little vocabulary in common.
No dimension means “money”
It is natural to imagine one number per concept. In practice the individual numbers in a real embedding cannot be read by humans; meaning is spread across all of them. That is part of why embedding search is difficult to audit.
Cosine similarity: a worked example
To compare two embeddings, most systems use cosine similarity, which measures the angle between the two vectors. Pointing the same way gives 1; at right angles gives 0; pointing in opposite directions gives -1. With text embeddings you will mostly see values between 0 and 1.
The recipe: multiply the two vectors number by number and add up the products (the dot product), then divide by the product of their lengths. Here it is with toy vectors of three dimensions. For readability we pretend the three numbers mean “money”, “travel delays” and “official documents”. Real models don't label their dimensions, but the arithmetic is identical.
Toy embeddings for three questions
| Question | Money | Delays | Documents |
|---|---|---|---|
| A: “How do I claim delay compensation?” | 0.9 | 0.8 | 0.1 |
| B: “My train was late, can I get money back?” | 0.8 | 0.9 | 0.2 |
| C: “How do I renew my passport?” | 0.1 | 0.2 | 0.9 |
- A·B = (0.9 × 0.8) + (0.8 × 0.9) + (0.1 × 0.2) = 0.72 + 0.72 + 0.02 = 1.46.
- Length of A = √(0.81 + 0.64 + 0.01) = √1.46 ≈ 1.208. Length of B = √(0.64 + 0.81 + 0.04) = √1.49 ≈ 1.221.
- Similarity of A and B = 1.46 ÷ (1.208 × 1.221) = 1.46 ÷ 1.475 ≈ 0.99: practically the same question.
- A·C = 0.09 + 0.16 + 0.09 = 0.34. Length of C = √(0.01 + 0.04 + 0.81) = √0.86 ≈ 0.927.
- Similarity of A and C = 0.34 ÷ (1.208 × 0.927) = 0.34 ÷ 1.120 ≈ 0.30: a different subject.
Scale this up and you have semantic search. Every passage in your documents is embedded once and stored; each incoming question is embedded on the fly; the system returns the passages with the highest similarity. Comparing a question against millions of vectors one by one is slow, hence specialised indexes and vector databases, which find the nearest neighbours approximately but quickly.
Multilingual embeddings
Some models are trained on many languages together, frequently using sentences paired with their translations. They place “delay compensation”, “indemnisation retard” and “Verspätungsentschädigung” close to one another, so a question in Polish or Welsh could in principle retrieve an English passage.
In practice, results vary. Languages that were well represented in training usually behave best, while smaller languages, dialect and specialist terminology are less reliable. If your organisation serves customers in several languages, test each one with genuine questions before relying on cross-language retrieval.
Where embeddings fall short
- Exact references. “VAT Notice 700”, a Companies House number or a policy clause such as “4.2(b)” carry little semantic weight. A vector search may return a loosely related passage where a keyword search would hit the exact one.
- Negation. “Eligible for Delay Repay” and “not eligible for Delay Repay” are about the same topic and can score as very similar. Similarity is about subject, not about whether two statements agree.
- Fresh or niche vocabulary. Internal project names, new product lines or sector acronyms may be poorly represented if the model never met them.
- Long chunks blur. A single vector for a long passage averages several ideas, so one specific detail can be drowned out.
- Model lock-in. Vectors from different models cannot be compared. Changing model means re-embedding the whole corpus.
- Limited explainability. It is hard to tell a colleague why a vector search returned a particular result.
- Data protection. Embeddings are not anonymised data. Studies have shown that text can be partly recovered from its vectors, so under UK GDPR it is prudent to treat embeddings of personal data as personal data, and the ICO guidance on security applies to them as much as to the source files.
These are reasons for care, not for avoidance. Many production systems therefore run hybrid search, blending semantic similarity with conventional full-text search so that both paraphrases and exact references are caught.
Does your project actually need embeddings?
When embeddings earn their keep
| Your situation | Embeddings | Full-text search |
|---|---|---|
| People describe things in their own words | Strong | Fine with query expansion |
| Documents packed with references and codes | Patchy | Strong and precise |
| You need to justify each result | Difficult | Straightforward |
| Huge, loosely organised archive | Strong | Workable with good ranking |
| You want as little infrastructure as possible | Model plus vector index | Built into common databases |
There is a third option that is often overlooked: keep full-text search, but ask a language model to expand the question into extra keywords and synonyms first. “Money back for a late train” becomes a search that also includes “compensation”, “refund” and “delay”. You gain much of the semantic benefit without storing a single vector. This is how Kopik works: it does not use embeddings or a vector database, but a hybrid of full-text search and semantic keyword expansion, with the question expanded into the documents' language so cross-language questions still land.
If you do build on embeddings, budget for the full picture: storage, indexing, re-embedding when documents change and the occasional model migration. Our guide to vector database cost sets out what you are really paying for, and how RAG works, step by step shows where retrieval sits in the wider pipeline.
A short test plan for semantic search
- Collect 20 to 30 questions your colleagues or customers really ask.
- Include precise ones (reference numbers, names, dates) alongside vague ones.
- Add pairs that differ only by a negation and check they are ranked sensibly.
- Test each language you support separately.
- Note the embedding model and version so you know when to re-index.
- Store vectors under the same access controls as the original documents.
Prefer not to run any of this yourself? Kopik turns PDFs, Word documents, text and Markdown into a knowledge base that answers with cited passages, usable on the website, via a REST API or from MCP clients such as Claude, Cursor or ChatGPT.
Search your documents, no vectors to manage
Upload your files and get a knowledge base that retrieves the right passages and cites them. Creating a base is free.
Frequently asked questions
Are embeddings and vectors the same thing?
Not quite. A vector is any ordered list of numbers. An embedding is a vector produced by a model to represent meaning. All embeddings are vectors; plenty of vectors are not embeddings.
What cosine similarity score counts as a match?
There is no fixed threshold. Score ranges differ between models, so 0.8 may be a strong match with one and an average one with another. Compare scores for the same query against each other and set any cut-off using your own test questions.
Is an embedding model the same as a chatbot model?
No. An embedding model converts text into a vector and does nothing else. A language model writes text. In a retrieval system, the search step finds passages and a language model then drafts the answer from them.
Do embeddings count as personal data under UK GDPR?
If they are derived from personal data, it is safest to assume so, because text can be partly reconstructed from embeddings. Apply the same security, retention and access rules as for the original documents, and seek specialist advice for your specific case.
Can I search in one language and find documents in another?
Yes, with multilingual embedding models, or by translating or expanding the question into the documents' language before a full-text search. Either way, test quality language by language.
Get the Kopik newsletter
New knowledge bases, RAG guides and product news. One email every week or two, unsubscribe in one click.
By subscribing you agree to receive our newsletter. We never share your address.