The Core Idea: In the last module, we saw that RNNs and LSTMs read a sentence one word at a time from left to right—so they naturally know which word came first, second, and third. A Transformer throws out that slow loop and reads all words in the sentence at the exact same time! That makes it super fast on GPUs, but creates a huge problem: without a loop, all the words get dumped into a bag at once and the model has no idea what order they were in! Positional Encodings fix this by stamping a unique "Seat Number" pattern onto every word vector before the Transformer reads it.
- Beginner: The "Bag of Scrabble Tiles" Problem
Look at these two sentences that use the exact same words:
To a human, those two sentences mean opposite things because of word order.
But if you feed both sentences into a raw Transformer without Positional Encodings, Self-Attention compares every word to every other word all at once—like shaking Scrabble tiles inside a bag! To a raw Transformer, Sentence A and Sentence B produce the exact same mathematical values!
The Fix: Give Every Word a "Seat Number" Vector!
Before the words enter the first Transformer layer, we take each word's Meaning Vector (from nn.Embedding) and add () a Position Vector that tells the model whether that word is sitting in Seat #0, Seat #1, or Seat #2:
- Beginner: Why Simple Seat Numbers (1, 2, 3...) Fail
At first, you might think: "Why not just add to the first word, to the second word, to the third word, and to the 500th word?"
Let us see why the simple ideas break down immediately:
Remember that word meaning vectors contain tiny decimals between and (like [0.42, -0.15]). If you add to the 500th word in an essay, the vector becomes [500.42, 499.85]—the giant position number completely drowns out the meaning of the word!
What if you divide by sentence length so the first word is and the last word is ? In a 4-word sentence, the step between words is . In a 100-word paragraph, the step between words shrinks to ! A 1-word step would mean something totally different in every sentence!
What a Great Positional Encoding Needs (The 3 Rules)
How do standard Positional Encodings combine with a word's Meaning Embedding vector at the input of a Transformer?
A) They have the exact same dimension d_model and are added together element-by-element (+)▼
x = word_emb + pos_emb) so the vector size stays 512!B) We replace the word embedding completely and throw the word's meaning away▼
- Medium: Sinusoidal Positional Encodings (The Hands of a Clock)
How did the creators of the original 2017 Transformer ("Attention Is All You Need") solve all three rules at once? With a brilliant analogy: The Hands of a Clock!
Think about how a clock tells time using three hands spinning at different speeds:
• The Second Hand ticks super fast (changes on every single step).
• The Minute Hand moves at medium speed.
• The Hour Hand moves super slowly (only changes over long distances).
Sinusoidal Positional Encoding does the exact same thing using smooth Sine and Cosine waves ( and ), which naturally stay bounded between and forever!
Like the Second Hand on a watch, the first few numbers in the vector wiggle up and down rapidly with every word step—helping the AI tell immediate neighbors (Position #4 vs. Position #5) apart.
Because of the denominator, the later numbers in the vector wave very slowly—like the Hour Hand—helping the AI track which paragraph or chapter a word belongs to!
- Medium: Learned Positional Embeddings (Used in BERT & GPT-2)
Instead of using fixed math waves ( and ), models like BERT and GPT-2 took an even simpler approach called Learned Positional Embeddings:
They literally created a second nn.Embedding(max_seq_len, d_model) lookup table just for seat numbers (Seat #0, Seat #1, ..., Seat #1023) and let Backpropagation learn the best numbers for each seat during training!
| Method | How It Works | Can It Read Longer Texts Than Trained On? | Famous Models |
|---|---|---|---|
| 1. Sinusoidal Waves | Fixed and frequencies added to input embeddings ( extra trainable parameters). | Somewhat (waves continue forever) | Original 2017 Transformer |
| 2. Learned Position Table | A trainable nn.Embedding(L_max, d) lookup table indexed by position . | No! Hard limit at (e.g. 512 in BERT) | BERT, RoBERTa, GPT-2, GPT-3 |
| 3. RoPE (Rotary Embeddings) | Rotates Query and Key vectors by an angle inside every Attention layer! | Yes! Scales to 128k–1M+ tokens! | Llama-3, Mistral, DeepSeek, Gemma, Qwen |
BERT was trained using a Learned Positional Embedding table of size nn.Embedding(512, 768). What happens if you try to pass a 600-token paragraph into BERT without truncating it?
A) It crashes with an IndexError because Row #512 to #599 do not exist in the position lookup table!▼
B) BERT automatically adds 88 new trained rows on the fly▼
nn.Embedding table cannot invent trained weights for unseen row indices at test time.
- Advanced: RoPE (Rotary Position Embeddings) — How Modern LLMs Work
If you open the code for Llama-3, Mistral, DeepSeek, or Qwen, you will find they do not add position vectors at the start of the model! Why?
Think about human language: when you read a sentence, does it really matter if the phrase "blue sky" is at Word #10 and #11 versus Word #510 and #511? No! What matters is the Relative Distance between the two words ( word apart)!
RoPE (Rotary Position Embedding, Su et al. 2021) encodes relative distance with a super clean geometric trick: instead of adding numbers to the word vector, RoPE rotates the word's vector like the hand of a clock by an angle proportional to its seat number !
Remember that Self-Attention compares Word and Word by measuring the angle between their vectors (dot product). If Word is rotated by and Word is rotated by , the angle between them depends ONLY on their relative gap !
Because RoPE is just clock rotation, engineers can stretch a model's context window from tokens to tokens (NTK-Aware RoPE Scaling / YaRN) simply by slowing down the rotation speed so more words fit around the clock face!
In RoPE (Rotary Position Embedding), if Word A at position is rotated by and Word B at position is rotated by , what is the angle difference between them? What if the same two words appear later at and ?
A) Both pairs have the exact same relative angle difference of 2*theta!▼
B) The second pair is 500 times farther apart in angle▼
- Visual Explanation: The Wave Fingerprint & Rotary Clock Console
Look at how a Transformer stamps seat numbers onto every word—comparing the classic Multi-Speed Wave Barcode (Sinusoidal) with modern RoPE Clock Dials (Llama-3):
PART 1 · SINUSOIDAL WAVE BARCODE (MULTI-SPEED CLOCK HANDS)
Each Seat Gets a Unique Wave Fingerprint!PART 2 · MODERN RoPE CLOCK DIALS (LLAMA-3 & DEEPSEEK)
Each Seat Rotates the Vector by +30°!🧭 Click to See How RoPE Tells "Dog Bit Man" Apart From "Man Bit Dog"!
See Proof▼
🧭 Click to See How RoPE Tells "Dog Bit Man" Apart From "Man Bit Dog"!
▼
"dog" is at 30° and "bit" is at 60° (+30° clockwise). The AI immediately sees "dog" came before "bit"!
Now "dog" is rotated to 90° (-30° behind "bit")! The relative clock angle flipped from +30° to -30°, so Attention knows "dog" is now the victim!
- Python Implementation: Sinusoidal & RoPE Positional Encodings in PyTorch
Here is a clean, runnable PyTorch script that builds both the classic Sinusoidal Positional Encoding table and a modern 2D RoPE Rotation function so you can see how both work in less than 40 lines of code:
import mathimport torch # 1. Classic Sinusoidal Positional Encoding (Attention Is All You Need, 2017)def buildSinusoidalPE(maxLen: int, dModel: int) -> torch.Tensor: pe = torch.zeros(maxLen, dModel) positions = torch.arange(0, maxLen, dtype=torch.float).unsqueeze(1) # (maxLen, 1) divTerm = torch.exp(torch.arange(0, dModel, 2).float() * (-math.log(10000.0) / dModel)) pe[:, 0::2] = torch.sin(positions * divTerm) # Even sliders: Sine wave pe[:, 1::2] = torch.cos(positions * divTerm) # Odd sliders: Cosine wave return pe posTable = buildSinusoidalPE(maxLen=10, dModel=4)print("Seat #0 Wave Vector:", posTable[0].numpy().round(3))print("Seat #1 Wave Vector:", posTable[1].numpy().round(3)) # 2. Modern RoPE (Rotary Position Embedding): Rotate a 2D Vector by m * theta!def applyRope2D(vec: torch.Tensor, seatPos: int, theta: float = 0.5) -> torch.Tensor: angle = seatPos * theta cosA, sinA = math.cos(angle), math.sin(angle) x1, x2 = vec[0], vec[1] return torch.tensor([x1 * cosA - x2 * sinA, x1 * sinA + x2 * cosA]) # Proof: Dot product between Seat #1 & #3 equals Seat #101 & #103 (Both 2 seats apart!)qWord = torch.tensor([1.0, 0.5])kWord = torch.tensor([0.8, 0.2]) scoreNear = torch.dot(applyRope2D(qWord, 1), applyRope2D(kWord, 3))scoreFar = torch.dot(applyRope2D(qWord, 101), applyRope2D(kWord, 103))print("Attention Score (Seats 1 vs 3):", round(scoreNear.item(), 5))print("Attention Score (Seats 101 vs 103):", round(scoreFar.item(), 5)) # Exact Match!Pro Tip (Look at the Two RoPE Scores at the Bottom!):
Run the code above and look at the last two lines: comparing Seat #1 vs. Seat #3 gives the exact same attention score to 5 decimal places as comparing Seat #101 vs. Seat #103! That is why RoPE allows modern LLMs like Llama-3 to understand relative word distance effortlessly across huge documents.
Key Points
nn.Embedding(L_max, d) table, but cannot extrapolate beyond .Common Mistakes
✕ Registering a fixed Sinusoidal Positional Encoding tensor as a trainable nn.Parameter.
Sinusoidal wave tables are fixed mathematical constants and should not be updated by the optimizer! In PyTorch, always register them with self.register_buffer("pe", pe) so they move to the GPU automatically without receiving gradients.
✕ Concatenating positional encodings to word embeddings instead of adding them.
Concatenating a 512-D position vector to a 512-D word vector doubles the hidden dimension to 1,024 and quadruples matrix sizes in every Transformer layer. Standard positional encodings are added (), and RoPE is applied via in-place rotation.
✕ Applying RoPE to Value (V) vectors in Self-Attention.
RoPE is applied only to Query () and Key () vectors so that the attention score captures relative distance; rotating the Value () vectors scrambles the actual semantic content being passed forward!
The Big Picture
Without Positional Encodings (Bag of Words)
Parallel Processing → Words Lose All Order → "Dog Bit Man" Equals "Man Bit Dog"
With Positional Encodings / RoPE (Parallel + Ordered!)
Stamp Every Word With Wave / Clock Angle → Process All Words at Once on GPU → Full Order Preserved!
The big takeaway is simple: Positional Encodings give us the best of both worlds. We get the blazing parallel GPU speed of processing an entire book at once, without losing track of which word came first, second, or last.
Remember: Now that our tokens have both Meaning (Word Embeddings) and Order (Positional Encodings), they are ready to enter the beating heart of the Transformer in our next lesson: Self-Attention!