What are embeddings?

An embedding is a list of numbers that represents a piece of text. The numbers are chosen so that texts about similar things get similar lists. Once text is in this form, "find passages related to this question" becomes arithmetic: compare the question's numbers with each passage's numbers and keep the closest.

The list is also called a vector, and searching by comparing vectors is called vector search. It is the retrieval step of RAG.

The simplest embedding: counting words

The easiest way to turn text into numbers is to count words. Imagine a very long list with one position for every possible word. For a given passage, each position holds the number of times that word appears; almost all positions are zero.

The RAG tool's word mode works like this, with a list of 2,048 positions. Very common words such as "the", "is" and "how" are ignored, and word endings are trimmed so that "returns" and "returned" land on the same position. The question "How many days do I have to return a bike?" becomes a list in which just four positions are not zero: the ones for "many", "days", "return" and "bike".

Comparing two vectors

The usual measure is cosine similarity. It asks how closely two lists point in the same direction, and gives a number between 0 and 1 for word counts: 1 means identical, 0 means nothing in common.

The sum has three parts. Multiply the two lists position by position and add the results; this is the overlap. Then work out the length of each list (the square root of the sum of its squared numbers). The similarity is the overlap divided by the two lengths multiplied together.

For the bike question and the handbook passage about returns:

  1. Shared words: "return" appears once in the question and three times in the passage, "bike" once and once, "days" once and once. Overlap = 3 + 1 + 1 = 5.
  2. Length of the question vector = 2. Length of the passage vector = 5.831.
  3. Similarity = 5 ÷ (2 × 5.831) = 0.429.

Dividing by the lengths matters. A long passage contains more words, so it would win every comparison on raw overlap alone. Dividing removes that advantage. Open steps 3 and 4 of the Step by step panel in the RAG tool to see this sum for your own document.

Embeddings that capture meaning

Word counting has an obvious weakness: "refund" and "money back" share no words, so they score zero. Modern embeddings are produced by a language model trained on large amounts of text. It reads the whole passage and outputs a shorter, dense list, typically a few hundred to a few thousand numbers, in which every position holds a value. Texts that mean similar things end up close together even when they use different words.

The individual numbers no longer stand for anything a person can name. You cannot point at position 37 and say what it means. What you can do is compare two lists, with the same cosine sum as before.

The RAG tool's "meaning search" uses a small model of this kind that runs in your browser and produces 384 numbers per passage.

What each kind is good at

Because the weaknesses are opposite, many systems use both together. That is hybrid search.

Things worth knowing

See it on your own text in the RAG tool: step 3 shows the vectors, step 4 the similarity sum and step 5 the map.

More guides

What is RAG? · How to choose a chunk size for RAG · What are tokens, and what do they cost? · Hybrid search: words plus meaning · How to test RAG retrieval · Seven common RAG mistakes