Word Embeddings: How AI Turns Words into Meaning

Learn how Word Embeddings turn words into coordinates of meaning—from simple personality sliders and Word2Vec to Cosine Similarity, PyTorch nn.Embedding, and modern LLM embeddings.

24 minBeginnerCode Examples

The Core Idea: In the last lesson, our tokenizer gave every word an ID number—like giving "cat" ID #412 and "kitten" ID #8951. But those ID numbers are just random barcode labels! They do not tell the computer that a cat and a kitten are almost the same thing. Word Embeddings fix this by giving every word a list of trait scores (like a character sheet in a video game), placing words with similar meanings right next to each other on a giant map.

  1. Beginner: Why Barcode IDs Are Not Enough (The Slider Bar Analogy)

Imagine walking into a supermarket where every item has a random barcode number: Apples = #105, Car Batteries = #106, and Oranges = #9402. If you only look at those numbers, Apples (#105) look like they are right next to Car Batteries (#106) and miles away from Oranges (#9402)!

That is the exact problem a neural network faces after tokenization. An ID number is just a label—it contains zero information about what the word actually means.

How do we teach a computer what a word means? Imagine creating a Video Game Character Sheet for every word using simple −1.0-1.0 to +1.0+1.0 slider bars for different traits:

WordIs it an Animal?Is it Royal?Is it Food?Final Word Embedding (Vector)
"Cat"+0.950.02-0.90[0.95, 0.02, -0.90]
"Dog"+0.920.01-0.88[0.92, 0.01, -0.88]
"King"0.10+0.98-0.95[0.10, 0.98, -0.95]
"Pizza"-0.990.00+0.99[-0.99, 0.00, 0.99]

Look at that last column: [0.95, 0.02, -0.90]. That list of slider numbers IS a Word Embedding!

Why "Cat" and "Dog" Match Instantly

Even if the computer has never seen a real cat or dog, it can compare their two lists of numbers—[0.95, 0.02, -0.90] and [0.92, 0.01, -0.88]—and immediately see that they are almost identical!

Who Sets All Those Slider Numbers?

Humans do not type these numbers in by hand! In a real AI model, each word gets 300300 to 4,0964{,}096 sliders that start as random numbers and adjust themselves automatically as the AI reads billions of sentences.

  1. Beginner: The Map of Meaning, Word Math & Cosine Similarity

Once every word has a list of numbers (a vector), you can plot every word as a dot on a map—just like GPS coordinates! On this map:
• All the fruit words (apple, banana, mango) cluster together in one neighborhood.
• All the programming words (Python, Java, code) cluster together in another neighborhood.

Doing Math on Ideas: King − Man + Woman = Queen

Because words are now just lists of numbers, we can literally add and subtract words! Look at the most famous example in AI:

Vec(King)−Vec(Man)+Vec(Woman)≈Vec(Queen)\text{Vec}(\text{King}) - \text{Vec}(\text{Man}) + \text{Vec}(\text{Woman}) \approx \text{Vec}(\text{Queen})

Why does this work? Start at "King" (Royal + Male). Subtract "Man" (removes the Male trait, leaving pure Royalty). Now add "Woman" (adds the Female trait). When you look at the map to see which word lives at those new coordinates, you land right on "Queen"!

How Does AI Measure if Two Words Are Similar? (Cosine Similarity)

Imagine drawing an arrow from the center of the map (0,0)(0, 0) to each word. To check if two words mean the same thing, we measure the angle between their two arrows using a score called Cosine Similarity:

Cosine Similarity(u,v)=u⋅v∥u∥∥v∥(Ranges from −1.0 to +1.0)\text{Cosine Similarity}(\mathbf{u}, \mathbf{v}) = \dfrac{\mathbf{u} \cdot \mathbf{v}}{\|\mathbf{u}\| \|\mathbf{v}\|} \quad \text{(Ranges from } -1.0 \text{ to } +1.0\text{)}
Score near +1.0

Arrows point the exact same way! The words have very similar meanings (like "happy" and "joyful" = 0.910.91).

Score near 0.0

Arrows are 90° apart. The words have nothing to do with each other (like "pizza" and "algebra" = 0.020.02).

Score near -1.0

Arrows point in opposite directions along the map.

⚡ Knowledge Check

If you take the embedding vector for "Paris", subtract "France", and add "Japan", which word's vector will you land closest to on the embedding map?

A) "Tokyo" (the capital of Japan)▼
✓ Correct!Paris−France\text{Paris} - \text{France} extracts the concept of "Capital City of". Adding Japan\text{Japan} moves to the capital city of Japan: Tokyo\text{Tokyo}!
B) "Croissant"▼
✕ Incorrect.Subtracting "France" removes the French country trait and replaces it with "Japan", steering the vector toward Japan's capital, Tokyo.

  1. Medium: How AI Learns Embeddings Automatically (Word2Vec & GloVe)

Here is the million-dollar question: if nobody types in the slider numbers by hand, how does the computer figure out that "coffee" and "tea" should have similar numbers?

By playing a simple "Fill-in-the-Blank" guessing game across millions of sentences! Think about these two sentences:
• "I poured a hot cup of coffee this morning."
• "I poured a hot cup of tea this morning."

Notice that the surrounding neighbor words ("poured", "hot", "cup", "morning") are identical! If a neural network is trained to guess a word from its neighbors, it is forced to give "coffee" and "tea" almost the exact same embedding vectors so both words predict the same neighbors!

In 2013, Google released Word2Vec, which uses this exact neighbor-guessing game in two ways:

1. CBOW (Guess the Center Word)

Continuous Bag of Words: Give the model the surrounding neighbor words ("The", "cat", "on", "the") and train it to guess the missing middle word ("sat").

Neighbors → Guess Middle Word

2. Skip-Gram (Guess the Neighbors)

Flips the game around: give the model just the middle word ("sat") and train it to guess the words sitting next to it ("cat", "on"). Works especially well for rare words!

Middle Word → Guess Neighbors

The Clever Speed Trick: Negative Sampling (True Pairs vs. Fake Pairs)

Predicting a neighbor out of all 100,000100{,}000 words in a dictionary is slow. So Word2Vec uses a shortcut called Negative Sampling: it shows the model 11 real word pair from the text (like ("orange", "juice") → True) and 55 random fake pairs (like ("orange", "tractor") → False). The model simply nudges real pairs closer together on the map and pushes fake pairs apart!

GloVe (Stanford, 2014)

Instead of scanning one sentence at a time, GloVe (Global Vectors) counts how often every pair of words appears near each other across the entire encyclopedia first, and trains vectors to match those global counts.

FastText (Meta AI, 2016)

Learns vectors for 3-letter chunks (like "tea", "eac", "ach" inside "teacher") and adds them up—so even if a user makes a typo, FastText can still build a good vector for it!

  1. Medium: How PyTorch nn.Embedding Works (The Giant Spreadsheet)

When you write PyTorch code for a language model, the very first layer is always nn.Embedding(num_embeddings, embedding_dim). Do not let the name intimidate you—under the hood, it is literally just a giant 2D spreadsheet (lookup table)!

1. Rows = Vocab Size
num_embeddings = 50,000

There is one row in the spreadsheet for every token ID in your tokenizer's vocabulary.

2. Columns = Sliders
embedding_dim = 768

How many numbers (trait sliders) describe each word. Small models use 128–768128\text{--}768; large LLMs use 4,096+4{,}096+.

3. Instant Row Lookup
table.weight[token_id]

When you pass in Token ID #412, PyTorch simply grabs Row #412 from the table!

Input IDs Shape: (B,  T)→nn.Embedding(V,  d)Output Tensor Shape: (B,  T,  d)\text{Input IDs Shape: } (B, \; T) \quad \xrightarrow{\text{nn.Embedding}(V, \; d)} \quad \text{Output Tensor Shape: } (B, \; T, \; d)
⚡ Knowledge Check

You create embed = nn.Embedding(10000, 256) (10,000 words in vocabulary, 256 numbers per word). If you pass in a batch of 44 sentences where each sentence has 1010 token IDs—shape (4, 10)—what shape comes out?

A) (4, 10, 256) — every token ID is replaced by its 256-number vector▼
✓ Correct!Each of the 1010 integer IDs in each of the 44 sentences is swapped for its 256256-number row from the lookup table, giving a 3D tensor of shape (4, 10, 256)!
B) (4, 10000)▼
✕ Incorrect.10,00010{,}000 is the total number of rows in the dictionary table. The output replaces each token with a vector of length embedding_dim = 256.

  1. Advanced: Static vs. Contextual Embeddings (How ChatGPT & RAG Work)

Now let us bridge from classic embeddings (Word2Vec/GloVe) to modern AI (BERT, GPT-4, Llama-3, and Vector Databases).

Classic Word2Vec had one big blind spot: it gave every word only ONE fixed vector in the table, even when the same spelling has two completely different meanings! Look at the word "apple" in these two sentences:
1. "I ate a sweet red apple for lunch." (The fruit)
2. "Apple released a new iPhone today." (The tech company)

1. Static Embeddings (Word2Vec / GloVe)

Because Word2Vec just looks up one frozen row in a table, it gives "apple" the exact same vector in both sentences—an awkward, blurry mix of half-fruit and half-computer!

1 Word = 1 Frozen Vector Forever

2. Contextual Embeddings (Transformers / LLMs)

In GPT-4 and BERT, the initial nn.Embedding table is just the starting point! Next, Self-Attention layers let surrounding words ("ate, sweet" vs. "iPhone") update "apple"'s vector on the fly!

Vector Changes Based on the Sentence!

How Sentence Embeddings Power AI Search & RAG (Retrieval-Augmented Generation)

In modern AI apps, we do not just embed single words—we pass entire paragraphs through an Embedding Model (like OpenAI's text-embedding-3 or open-source BGE) to get one vector per paragraph and store them in a Vector Database (like Pinecone or Qdrant). When a user asks a question, we embed their question and use Cosine Similarity to instantly find the most relevant paragraphs—even if the user and the document used completely different words!

⚡ Knowledge Check

Look at the word "bat" in: (1) "A bat flew out of the dark cave" and (2) "He swung the baseball bat." How does a modern Transformer (like BERT or Llama-3) handle the embedding for "bat" compared to Word2Vec?

A) Word2Vec gives the exact same vector for both; the Transformer updates the vector using surrounding words▼
✓ Correct!Word2Vec is a static lookup table, whereas a Transformer blends context from "cave" vs. "baseball" so the output vector for "bat" moves to the animal neighborhood in Sentence 1 and the sports neighborhood in Sentence 2!
B) Both models output the exact same fixed vector for "bat"▼
✕ Incorrect.Only static models output a single fixed vector regardless of context; Transformers create dynamic, context-aware vectors.

  1. Visual Explanation: How Words Become Coordinates of Meaning

Follow the exact 4-step journey of the word "cat"—from a meaningless barcode number into a rich vector of meaning that powers both AI search and next-word prediction[cite: 3]:

STAGE 1 · TOKEN ID

Input
Word Token
"cat"

ID #412

Right now, #412 is just a random barcode—the AI has zero clue what a cat is!

STAGE 2 · nn.Embedding TABLE LOOKUP

Grabs Row #412
Row #AnimalRoyalFood
#411 (bus)-0.850.01-0.92
➔ #412 (cat)+0.950.02-0.90
#413 (pie)-0.910.00+0.96

Initial Vector Pulled: [0.95, 0.02, -0.90] (Now the AI knows it's an animal!)

▼ Pass Vector Into Transformer Self-Attention Layers ▼

STAGE 3 · TRANSFORMER ATTENTION BLENDS SENTENCE CONTEXT

Surrounding words update the vector!
Full Sentence Read by AI:

"The sleepy cat purred on the rug."

  • Blends "sleepy" & "purred" ➔
Updated Contextual Vector:

Now knows this isn't a wild jungle cat—it's a cozy, indoor sleeping housepet!

↙ Now This Rich Vector Powers Two AI Superpowers ↘

SUPERPOWER A · COSINE SIMILARITY CHECK

🔍 AI Semantic Search & RAG

Test

▼

Compares our "cat" vector angle against millions of documents:
🐈 Document about "Kittens & Pets"0.94 Match ✓
🚗 Document about "Car Engines"0.03 Match ✕

SUPERPOWER B · OUTPUT LAYER

💬 Predict the Next Word (ChatGPT)

Test

▼

Uses the vector for "The sleepy cat..." to score the next word:
Next word: "purred"78% Prob ✓
Next word: "barked"0.1% Prob ✕

  1. Python Implementation: PyTorch Embeddings, Cosine Similarity & Word Math

Here is a clean, beginner-friendly PyTorch script showing: (1) how nn.Embedding looks up vectors for token IDs, (2) how to check word similarity using Cosine Similarity, and (3) how to run the famous King−Man+Woman=Queen\text{King} - \text{Man} + \text{Woman} = \text{Queen} word math in code:

word_embeddings_demo.pyPython 3.11+ · PyTorch 2.x
import torchimport torch.nn as nnimport torch.nn.functional as F # 1. PyTorch nn.Embedding: A Lookup Table of 1,000 Words x 64 Sliders Eachtorch.manual_seed(42)embeddingTable = nn.Embedding(num_embeddings=1000, embedding_dim=64) # Pass in 2 sentences of 4 Token IDs each -> Shape: (2, 4)tokenIds = torch.tensor([[10, 25, 99, 4], [8, 12, 55, 1]])wordVectors = embeddingTable(tokenIds)          # Output Shape: (2, 4, 64)print("Input IDs Shape:", tokenIds.shape, "-> Output Vectors Shape:", wordVectors.shape) # 2. Mini Map of Meaning: 4 Sliders = [Royal, Female, Human, Fruit]words = {    "man":   torch.tensor([0.0, 0.0, 1.0, 0.0]),    "woman": torch.tensor([0.0, 1.0, 1.0, 0.0]),    "king":  torch.tensor([1.0, 0.0, 1.0, 0.0]),    "queen": torch.tensor([1.0, 1.0, 1.0, 0.0]),    "apple": torch.tensor([0.0, 0.0, 0.0, 1.0])} # 3. Measure Cosine Similarity (-1 to +1)simKingQueen = F.cosine_similarity(words["king"], words["queen"], dim=0)simKingApple = F.cosine_similarity(words["king"], words["apple"], dim=0)print("Similarity (king vs queen):", round(simKingQueen.item(), 3))  # 0.816 (High!)print("Similarity (king vs apple):", round(simKingApple.item(), 3))  # 0.000 (Unrelated!) # 4. Word Math: King - Man + Woman = ?mysteryVec = words["king"] - words["man"] + words["woman"]simToQueen = F.cosine_similarity(mysteryVec, words["queen"], dim=0)print("King - Man + Woman matches Queen with score:", round(simToQueen.item(), 3))  # 1.0!

Pro Tip (Always Keep Token IDs as Integers for nn.Embedding):

Because nn.Embedding uses your token IDs as row numbers to look up in its table (like Row #10 or Row #25), PyTorch requires tokenIds to be whole integers (torch.long / int64). If you accidentally convert tokenIds to decimals (float32), PyTorch will throw an error because there is no such thing as Row #10.5!

Key Points

✓Token IDs are just arbitrary barcode numbers; Word Embeddings replace each ID with a dense list of trait numbers (a vector) so words with similar meanings have similar coordinates.
✓Cosine Similarity measures the angle between two word vectors from −1.0-1.0 to +1.0+1.0: a score near +1.0+1.0 means the words have very similar meanings, while 0.00.0 means they are unrelated.
✓Word2Vec learns embeddings automatically from raw text by playing a neighbor-guessing game: CBOW guesses the middle word from its neighbors, while Skip-Gram guesses the neighbors from the middle word.
✓In PyTorch, nn.Embedding(vocab_size, dim) is a learnable lookup table that turns a 2D batch of integer token IDs (Batch, Seq) into a 3D tensor of vectors (Batch, Seq, dim).
✓Classic embeddings (Word2Vec, GloVe) give each word one fixed static vector forever, whereas modern Transformers (BERT, GPT-4, Llama-3) update word vectors dynamically based on the surrounding sentence context.

Common Mistakes

✕ Passing float numbers into nn.Embedding instead of integer token IDs.

nn.Embedding looks up rows by whole-number index. Always keep your input token IDs as torch.long (integers); the output coming out of nn.Embedding will automatically be float32 vectors!

✕ Comparing vectors from two different AI models (like OpenAI and Llama) in the same database.

Every AI model invents its own custom map of meaning during training. Mixing vectors from Model A and Model B is like mixing latitude/longitude from Earth and Mars—always use the exact same embedding model for both your saved documents and your search queries!

✕ Using static Word2Vec embeddings when a word's meaning depends heavily on context.

Static models cannot tell the difference between "river bank" and "bank account". Use a Transformer-based contextual embedding model whenever sentence context matters.

The Big Picture

Plain Token IDs (Random Barcodes)

"Cat" = #412, "Kitten" = #8951 → Zero Shared Meaning → Computer Is Blind to Synonyms

Word Embeddings (The Map of Meaning)

Token ID → Trait Vector in nn.Embedding → Similar Ideas Live Together → Math on Meaning!

The big takeaway is simple: Word Embeddings turn human words into map coordinates. Once words are coordinates on a map, finding similar ideas is just measuring how close two points are, and understanding a sentence is just moving those points around based on context.

Remember: Tokenization chops a sentence into ID numbers, Word Embeddings turn each ID number into a meaningful vector, and in our next lesson, Sequence Models will read those vectors in order from start to finish!