Blog
4 min read

What Are Embeddings? How AI Turns Meaning Into Numbers

An embedding is a list of numbers that captures what a piece of text means, so a computer can find similar things. How embeddings work, what they're used for (search, RAG, recommendations), how to store them, and practical tips on models, dimensions and cost.

If you've read about AI search, RAG or vector databases, you've met the word embedding. It sounds abstract, but the idea is simple and useful.

The short version

An embedding is a long list of numbers that represents the meaning of a piece of text (or an image). Texts with similar meanings get similar lists of numbers. That lets a computer answer "what's similar to this?" with arithmetic.

An analogy: a map of meaning

Imagine placing every sentence on a map so that sentences meaning similar things sit close together. "How do I get my money back?" lands right next to "What's your refund policy?" — even though they share no words. "Best pizza in Naples" lands far away.

An embedding is a sentence's coordinates on that map. Real embeddings use hundreds or thousands of dimensions instead of two, which lets them capture many kinds of similarity at once — topic, tone, intent.

What they look like

You send text to an embedding model, and get back something like:

"What's your refund policy?"  →  [0.021, -0.113, 0.087, ..., 0.004]   (1,024 numbers)

You never read these numbers yourself. You store them and compare them.

How similarity is measured

To compare two embeddings, you measure how close they are — usually with cosine similarity, which checks whether the two lists of numbers "point in the same direction". Close to 1 means very similar; close to 0 means unrelated.

You don't implement this yourself; databases with vector support do it for you, quickly, across millions of rows.

What embeddings are used for

  • Semantic search: find documents by meaning, not just matching words. A search for "laptop won't turn on" finds an article titled "Troubleshooting power issues".
  • RAG: find the relevant chunks of your content to hand to an AI model before it answers. See what is RAG?
  • Recommendations: "more like this" for articles, products or songs.
  • Duplicate detection: spot near-identical support tickets or listings.
  • Grouping: cluster feedback by topic without reading every item.

Where to get them

Embeddings come from dedicated embedding models, separate from chat models. Major AI providers and specialist companies offer them through an API; Anthropic's documentation points to Voyage AI for embeddings, and OpenAI, Google and others have their own. Open-source embedding models can also run on your own server.

Things to know when choosing:

  • Stick with one model. Embeddings from different models aren't comparable. If you switch models, you must re-embed everything.
  • Dimensions. More numbers per embedding can capture more nuance but use more storage. Many models let you choose a smaller size with little quality loss.
  • Cost is low. Embedding is far cheaper than generating text, and you embed each document once. Re-embed only when the content changes.
  • Language support matters if your content isn't in English.

Where to store them

You need a database that can store vectors and search them by similarity:

  • Postgres with pgvector — a popular choice because your embeddings sit next to your normal data. See our pgvector tutorial.
  • Dedicated vector databases — built only for this; useful at very large scale.

For most apps, start with pgvector. One database is easier to run, back up and secure than two.

Common beginner mistakes

  • Embedding whole documents. A 20-page document squashed into one embedding is a blurry average. Split into chunks first.
  • Forgetting to re-embed updates. Edited content with an old embedding gets found for the wrong things.
  • Expecting exact matches. Embeddings are fuzzy. Searching for an order number like INV-20391 works better with ordinary keyword search.
  • Mixing models after an upgrade without re-embedding.

The summary

  • An embedding is a list of numbers representing meaning; similar meanings get similar numbers.
  • They power semantic search, RAG, recommendations and duplicate detection.
  • Use one embedding model consistently, chunk long documents, and re-embed when content changes.
  • Postgres with pgvector is the simplest place to store them for most apps.

EasySpawn servers include PostgreSQL, so your embeddings can live beside the rest of your app's data — one database to query, back up and restore. See how it works or join the waitlist.

Related: What Is an LLM? · What Are Tokens in AI? · Postgres Full-Text Search · What Is a Database?

Keep reading