The Core Idea: In the last lesson, our tokenizer gave every word an ID number—like giving "cat" ID #412 and "kitten" ID #8951. But those ID numbers are just random barcode labels! They do not tell the computer that a cat and a kitten are almost the same thing. Word Embeddings fix this by giving every word a list of trait scores (like a character sheet in a video game), placing words with similar meanings right next to each other on a giant map.
- Beginner: Why Barcode IDs Are Not Enough (The Slider Bar Analogy)
Imagine walking into a supermarket where every item has a random barcode number: Apples = #105, Car Batteries = #106, and Oranges = #9402. If you only look at those numbers, Apples (#105) look like they are right next to Car Batteries (#106) and miles away from Oranges (#9402)!
That is the exact problem a neural network faces after tokenization. An ID number is just a label—it contains zero information about what the word actually means.
How do we teach a computer what a word means? Imagine creating a Video Game Character Sheet for every word using simple to slider bars for different traits:
| Word | Is it an Animal? | Is it Royal? | Is it Food? | Final Word Embedding (Vector) |
|---|---|---|---|---|
| "Cat" | +0.95 | 0.02 | -0.90 | [0.95, 0.02, -0.90] |
| "Dog" | +0.92 | 0.01 | -0.88 | [0.92, 0.01, -0.88] |
| "King" | 0.10 | +0.98 | -0.95 | [0.10, 0.98, -0.95] |
| "Pizza" | -0.99 | 0.00 | +0.99 | [-0.99, 0.00, 0.99] |
Look at that last column: [0.95, 0.02, -0.90]. That list of slider numbers IS a Word Embedding!
Even if the computer has never seen a real cat or dog, it can compare their two lists of numbers—[0.95, 0.02, -0.90] and [0.92, 0.01, -0.88]—and immediately see that they are almost identical!
Humans do not type these numbers in by hand! In a real AI model, each word gets to sliders that start as random numbers and adjust themselves automatically as the AI reads billions of sentences.
- Beginner: The Map of Meaning, Word Math & Cosine Similarity
Once every word has a list of numbers (a vector), you can plot every word as a dot on a map—just like GPS coordinates! On this map:
• All the fruit words (apple, banana, mango) cluster together in one neighborhood.
• All the programming words (Python, Java, code) cluster together in another neighborhood.
Doing Math on Ideas: King − Man + Woman = Queen
Because words are now just lists of numbers, we can literally add and subtract words! Look at the most famous example in AI:
Why does this work? Start at "King" (Royal + Male). Subtract "Man" (removes the Male trait, leaving pure Royalty). Now add "Woman" (adds the Female trait). When you look at the map to see which word lives at those new coordinates, you land right on "Queen"!
How Does AI Measure if Two Words Are Similar? (Cosine Similarity)
Imagine drawing an arrow from the center of the map to each word. To check if two words mean the same thing, we measure the angle between their two arrows using a score called Cosine Similarity:
Arrows point the exact same way! The words have very similar meanings (like "happy" and "joyful" = ).
Arrows are 90° apart. The words have nothing to do with each other (like "pizza" and "algebra" = ).
Arrows point in opposite directions along the map.
If you take the embedding vector for "Paris", subtract "France", and add "Japan", which word's vector will you land closest to on the embedding map?
A) "Tokyo" (the capital of Japan)▼
B) "Croissant"▼
- Medium: How AI Learns Embeddings Automatically (Word2Vec & GloVe)
Here is the million-dollar question: if nobody types in the slider numbers by hand, how does the computer figure out that "coffee" and "tea" should have similar numbers?
By playing a simple "Fill-in-the-Blank" guessing game across millions of sentences! Think about these two sentences:
• "I poured a hot cup of coffee this morning."
• "I poured a hot cup of tea this morning."
Notice that the surrounding neighbor words ("poured", "hot", "cup", "morning") are identical! If a neural network is trained to guess a word from its neighbors, it is forced to give "coffee" and "tea" almost the exact same embedding vectors so both words predict the same neighbors!
In 2013, Google released Word2Vec, which uses this exact neighbor-guessing game in two ways:
Continuous Bag of Words: Give the model the surrounding neighbor words ("The", "cat", "on", "the") and train it to guess the missing middle word ("sat").
Neighbors → Guess Middle Word
Flips the game around: give the model just the middle word ("sat") and train it to guess the words sitting next to it ("cat", "on"). Works especially well for rare words!
Middle Word → Guess Neighbors
The Clever Speed Trick: Negative Sampling (True Pairs vs. Fake Pairs)
Predicting a neighbor out of all words in a dictionary is slow. So Word2Vec uses a shortcut called Negative Sampling: it shows the model real word pair from the text (like ("orange", "juice") → True) and random fake pairs (like ("orange", "tractor") → False). The model simply nudges real pairs closer together on the map and pushes fake pairs apart!
Instead of scanning one sentence at a time, GloVe (Global Vectors) counts how often every pair of words appears near each other across the entire encyclopedia first, and trains vectors to match those global counts.
Learns vectors for 3-letter chunks (like "tea", "eac", "ach" inside "teacher") and adds them up—so even if a user makes a typo, FastText can still build a good vector for it!
- Medium: How PyTorch nn.Embedding Works (The Giant Spreadsheet)
When you write PyTorch code for a language model, the very first layer is always nn.Embedding(num_embeddings, embedding_dim). Do not let the name intimidate you—under the hood, it is literally just a giant 2D spreadsheet (lookup table)!
There is one row in the spreadsheet for every token ID in your tokenizer's vocabulary.
How many numbers (trait sliders) describe each word. Small models use ; large LLMs use .
When you pass in Token ID #412, PyTorch simply grabs Row #412 from the table!
You create embed = nn.Embedding(10000, 256) (10,000 words in vocabulary, 256 numbers per word). If you pass in a batch of sentences where each sentence has token IDs—shape (4, 10)—what shape comes out?
A) (4, 10, 256) — every token ID is replaced by its 256-number vector▼
(4, 10, 256)!B) (4, 10000)▼
embedding_dim = 256.
- Advanced: Static vs. Contextual Embeddings (How ChatGPT & RAG Work)
Now let us bridge from classic embeddings (Word2Vec/GloVe) to modern AI (BERT, GPT-4, Llama-3, and Vector Databases).
Classic Word2Vec had one big blind spot: it gave every word only ONE fixed vector in the table, even when the same spelling has two completely different meanings! Look at the word "apple" in these two sentences:
1. "I ate a sweet red apple for lunch." (The fruit)
2. "Apple released a new iPhone today." (The tech company)
Because Word2Vec just looks up one frozen row in a table, it gives "apple" the exact same vector in both sentences—an awkward, blurry mix of half-fruit and half-computer!
1 Word = 1 Frozen Vector Forever
In GPT-4 and BERT, the initial nn.Embedding table is just the starting point! Next, Self-Attention layers let surrounding words ("ate, sweet" vs. "iPhone") update "apple"'s vector on the fly!
Vector Changes Based on the Sentence!
How Sentence Embeddings Power AI Search & RAG (Retrieval-Augmented Generation)
In modern AI apps, we do not just embed single words—we pass entire paragraphs through an Embedding Model (like OpenAI's text-embedding-3 or open-source BGE) to get one vector per paragraph and store them in a Vector Database (like Pinecone or Qdrant). When a user asks a question, we embed their question and use Cosine Similarity to instantly find the most relevant paragraphs—even if the user and the document used completely different words!
Look at the word "bat" in: (1) "A bat flew out of the dark cave" and (2) "He swung the baseball bat." How does a modern Transformer (like BERT or Llama-3) handle the embedding for "bat" compared to Word2Vec?
A) Word2Vec gives the exact same vector for both; the Transformer updates the vector using surrounding words▼
B) Both models output the exact same fixed vector for "bat"▼
- Visual Explanation: How Words Become Coordinates of Meaning
Follow the exact 4-step journey of the word "cat"—from a meaningless barcode number into a rich vector of meaning that powers both AI search and next-word prediction[cite: 3]:
STAGE 1 · TOKEN ID
InputID #412
Right now, #412 is just a random barcode—the AI has zero clue what a cat is!
STAGE 2 · nn.Embedding TABLE LOOKUP
Grabs Row #412Initial Vector Pulled: [0.95, 0.02, -0.90] (Now the AI knows it's an animal!)
▼ Pass Vector Into Transformer Self-Attention Layers ▼
STAGE 3 · TRANSFORMER ATTENTION BLENDS SENTENCE CONTEXT
Surrounding words update the vector!"The sleepy cat purred on the rug."
- Blends "sleepy" & "purred" ➔
Now knows this isn't a wild jungle cat—it's a cozy, indoor sleeping housepet!
↙ Now This Rich Vector Powers Two AI Superpowers ↘
SUPERPOWER A · COSINE SIMILARITY CHECK
🔍 AI Semantic Search & RAG
Test▼
SUPERPOWER A · COSINE SIMILARITY CHECK
🔍 AI Semantic Search & RAG
▼
SUPERPOWER B · OUTPUT LAYER
💬 Predict the Next Word (ChatGPT)
Test▼
SUPERPOWER B · OUTPUT LAYER
💬 Predict the Next Word (ChatGPT)
▼
- Python Implementation: PyTorch Embeddings, Cosine Similarity & Word Math
Here is a clean, beginner-friendly PyTorch script showing: (1) how nn.Embedding looks up vectors for token IDs, (2) how to check word similarity using Cosine Similarity, and (3) how to run the famous word math in code:
import torchimport torch.nn as nnimport torch.nn.functional as F # 1. PyTorch nn.Embedding: A Lookup Table of 1,000 Words x 64 Sliders Eachtorch.manual_seed(42)embeddingTable = nn.Embedding(num_embeddings=1000, embedding_dim=64) # Pass in 2 sentences of 4 Token IDs each -> Shape: (2, 4)tokenIds = torch.tensor([[10, 25, 99, 4], [8, 12, 55, 1]])wordVectors = embeddingTable(tokenIds) # Output Shape: (2, 4, 64)print("Input IDs Shape:", tokenIds.shape, "-> Output Vectors Shape:", wordVectors.shape) # 2. Mini Map of Meaning: 4 Sliders = [Royal, Female, Human, Fruit]words = { "man": torch.tensor([0.0, 0.0, 1.0, 0.0]), "woman": torch.tensor([0.0, 1.0, 1.0, 0.0]), "king": torch.tensor([1.0, 0.0, 1.0, 0.0]), "queen": torch.tensor([1.0, 1.0, 1.0, 0.0]), "apple": torch.tensor([0.0, 0.0, 0.0, 1.0])} # 3. Measure Cosine Similarity (-1 to +1)simKingQueen = F.cosine_similarity(words["king"], words["queen"], dim=0)simKingApple = F.cosine_similarity(words["king"], words["apple"], dim=0)print("Similarity (king vs queen):", round(simKingQueen.item(), 3)) # 0.816 (High!)print("Similarity (king vs apple):", round(simKingApple.item(), 3)) # 0.000 (Unrelated!) # 4. Word Math: King - Man + Woman = ?mysteryVec = words["king"] - words["man"] + words["woman"]simToQueen = F.cosine_similarity(mysteryVec, words["queen"], dim=0)print("King - Man + Woman matches Queen with score:", round(simToQueen.item(), 3)) # 1.0!Pro Tip (Always Keep Token IDs as Integers for nn.Embedding):
Because nn.Embedding uses your token IDs as row numbers to look up in its table (like Row #10 or Row #25), PyTorch requires tokenIds to be whole integers (torch.long / int64). If you accidentally convert tokenIds to decimals (float32), PyTorch will throw an error because there is no such thing as Row #10.5!
Key Points
nn.Embedding(vocab_size, dim) is a learnable lookup table that turns a 2D batch of integer token IDs (Batch, Seq) into a 3D tensor of vectors (Batch, Seq, dim).Common Mistakes
✕ Passing float numbers into nn.Embedding instead of integer token IDs.
nn.Embedding looks up rows by whole-number index. Always keep your input token IDs as torch.long (integers); the output coming out of nn.Embedding will automatically be float32 vectors!
✕ Comparing vectors from two different AI models (like OpenAI and Llama) in the same database.
Every AI model invents its own custom map of meaning during training. Mixing vectors from Model A and Model B is like mixing latitude/longitude from Earth and Mars—always use the exact same embedding model for both your saved documents and your search queries!
✕ Using static Word2Vec embeddings when a word's meaning depends heavily on context.
Static models cannot tell the difference between "river bank" and "bank account". Use a Transformer-based contextual embedding model whenever sentence context matters.
The Big Picture
Plain Token IDs (Random Barcodes)
"Cat" = #412, "Kitten" = #8951 → Zero Shared Meaning → Computer Is Blind to Synonyms
Word Embeddings (The Map of Meaning)
Token ID → Trait Vector in nn.Embedding → Similar Ideas Live Together → Math on Meaning!
The big takeaway is simple: Word Embeddings turn human words into map coordinates. Once words are coordinates on a map, finding similar ideas is just measuring how close two points are, and understanding a sentence is just moving those points around based on context.
Remember: Tokenization chops a sentence into ID numbers, Word Embeddings turn each ID number into a meaningful vector, and in our next lesson, Sequence Models will read those vectors in order from start to finish!