Computers have always been good at exact matching and bad at meaning. Search for “cheap flights” and a traditional search engine finds documents containing those exact words. A document about “affordable airfare” does not match, even though it is the same thing.
Embeddings are how that problem got solved. They are the reason modern search finds what you meant rather than what you typed, and the reason retrieval-augmented generation works at all.
The core idea #
An embedding turns a piece of text into a list of numbers, typically a few hundred to a few thousand of them. That list is called a vector.
The property that makes it useful is that texts with similar meanings produce vectors that are close together. “Cheap flights” and “affordable airfare” land near each other. “Cheap flights” and “French cheese” land far apart.
Once meaning is expressed as position in a space, comparing meanings becomes measuring distance, and measuring distance is something computers do extremely fast.
What the space looks like #
The vector has hundreds or thousands of components. You cannot visualize that, so use the two-dimensional intuition and remember it is a simplification.
Imagine a map where every word or sentence is a point. Words about food cluster in one region. Words about weather cluster in another. Within the food region, fruits sit together, and within the fruits, citrus sits together.
Nothing built this map by hand. It emerged from training.
The training objective is roughly that words appearing in similar contexts should get similar vectors. “Cat” and “dog” both show up near “pet,” “vet,” “fur,” and “adopted.” The model has no idea what an animal is. It observed that these two words occupy interchangeable positions in an enormous amount of text, and placed them near each other. Meaning, as far as the model is concerned, is entirely a function of company kept.
The famous arithmetic #
The demonstration that made this click for a lot of people:
king − man + woman ≈ queen
Take the vector for “king,” subtract the vector for “man,” add the vector for “woman,” and the result lands near “queen.”
What that shows is that directions in the space carry meaning. There is a rough gender direction, so moving along it converts king to queen, actor to actress, uncle to aunt. There are directions for plurality, for verb tense, for country-to-capital.
None of this was programmed. It fell out of training on text.
The caveat: this works cleanly on curated examples and messily in general. The arithmetic is a real property of the space, not a party trick, and it is also less reliable than the famous example suggests. Modern sentence embeddings are optimized for similarity rather than for this kind of analogy, and the arithmetic works less neatly on them.
Word embeddings versus sentence embeddings #
Two distinct generations, and confusing them causes problems.
Word embeddings (word2vec, GloVe) give each word one fixed vector. “Bank” gets a single vector that has to serve both the river bank and the financial one. It cannot distinguish them, because it does not look at context.
Contextual and sentence embeddings (BERT and its descendants, and the modern embedding models from the major providers) embed an entire passage, and each word’s representation depends on the words around it. “I sat on the bank of the river” and “I deposited it at the bank” produce different vectors for “bank,” and the two sentences as a whole land in different regions.
Essentially everything practical today uses the second kind. When someone says “embeddings” in 2026 they almost always mean a sentence or passage embedding from a model trained specifically for similarity.
How similarity is measured #
The standard measure is cosine similarity, which is the cosine of the angle between two vectors. It ranges from 1 (pointing the same direction, very similar) through 0 (perpendicular, unrelated) to −1 (opposite).
The reason to use angle rather than straight-line distance is that it ignores magnitude. A long document and a short sentence about the same topic point in the same direction even though one vector is longer. You care about direction, which encodes what the text is about, not length, which encodes something closer to how much text there is.
In practice most systems normalize all vectors to the same length, at which point cosine similarity and Euclidean distance rank results identically and the distinction stops mattering.
What this is used for #
Semantic search. Embed every document in your collection once and store the vectors. When a query arrives, embed it and find the nearest stored vectors. You get results that match meaning rather than keywords, and this is the backbone of modern search inside products.
RAG. Retrieval-augmented generation is semantic search feeding a language model. Retrieve the passages closest to the question, put them in the prompt, and let the model answer from them. Embeddings are the retrieval half, and the quality of your embeddings sets the ceiling on the quality of your RAG system.
Recommendations. Embed items and users, then recommend items whose vectors are near things the user liked. This works for products, music, articles, and anything else you can represent as text.
Clustering and deduplication. Group similar documents automatically, find near-duplicate support tickets, detect that two differently-worded bug reports describe the same problem.
Classification. Embed labeled examples, then classify new text by which labeled examples it is nearest to. Often surprisingly competitive with training a dedicated classifier, and far cheaper.
Multimodal search. Some models, CLIP being the well-known one, embed images and text into the same space, so a photo of a dog and the word “dog” land near each other. That is how you search a photo library with a text description.
Vector databases #
Once you have millions of vectors, finding the nearest ones to a query becomes a performance problem. Comparing against every stored vector is linear in the size of your collection and gets slow.
Vector databases solve this with approximate nearest neighbor search. They build index structures that let you find almost certainly the closest vectors without exhaustively checking all of them, trading a small amount of accuracy for an enormous speedup. HNSW, a graph-based method, is the most common approach.
Practical note: for under a hundred thousand vectors, you do not need a vector database. A numpy array and a brute-force comparison is fast enough and dramatically simpler. A surprising number of systems adopt heavy infrastructure for a dataset a laptop could scan in twenty milliseconds.
Where embeddings fall short #
They compress. A thousand-word document becomes maybe 1,536 numbers, so detail is lost, necessarily. Two documents that differ in one critical clause can embed almost identically.
Negation is a known weakness. “The drug is effective” and “the drug is not effective” often embed close together, because they share nearly all their words and most of their structure. This is a real and dangerous failure mode in medical and legal retrieval.
They inherit training bias. If the training text associates certain professions with certain genders, the embedding space encodes that association, and any system built on it propagates it.
Exact matching is a weak spot. Searching for a specific product code, error number, or name is a job keyword search does better. Serious systems run hybrid search, combining keyword matching with vector similarity, precisely because each covers the other’s failure mode.
And chunking decisions matter enormously. How you split documents before embedding affects retrieval quality more than the choice of embedding model does, and it gets far less attention.
The mental model #
An embedding is a coordinate. Text goes in, a position in a high-dimensional space comes out, and that position was assigned so that similar meanings end up nearby. Everything else is measuring distance.
Once you have that, meaning becomes geometry, and computers are very good at geometry.
Related: what is RAG, which is the main practical application; what is a token, for the other way text gets turned into numbers; and how LLMs work, which uses embeddings internally as its very first step.