Positional Encodings: How Transformers Know Word Order

Learn why Transformers are blind to word order without help—and how Sinusoidal Positional Encodings, Learned Embeddings, and RoPE (Rotary Position Embeddings) stamp every word with its place in line.

24 minBeginnerCode Examples

The Core Idea: In the last module, we saw that RNNs and LSTMs read a sentence one word at a time from left to right—so they naturally know which word came first, second, and third. A Transformer throws out that slow loop and reads all words in the sentence at the exact same time! That makes it super fast on GPUs, but creates a huge problem: without a loop, all the words get dumped into a bag at once and the model has no idea what order they were in! Positional Encodings fix this by stamping a unique "Seat Number" pattern onto every word vector before the Transformer reads it.

  1. Beginner: The "Bag of Scrabble Tiles" Problem

Look at these two sentences that use the exact same 55 words:

Sentence A
"The dog bit the man."
An everyday accident
Sentence B
"The man bit the dog."
Crazy front-page news!

To a human, those two sentences mean opposite things because of word order.

But if you feed both sentences into a raw Transformer without Positional Encodings, Self-Attention compares every word to every other word all at once—like shaking Scrabble tiles inside a bag! To a raw Transformer, Sentence A and Sentence B produce the exact same mathematical values!

The Fix: Give Every Word a "Seat Number" Vector!

Before the words enter the first Transformer layer, we take each word's Meaning Vector (from nn.Embedding) and add (++) a Position Vector that tells the model whether that word is sitting in Seat #0, Seat #1, or Seat #2:

Final Input Vector=Word Meaning Vector+Position Vector\text{Final Input Vector} = \text{Word Meaning Vector} + \text{Position Vector}

  1. Beginner: Why Simple Seat Numbers (1, 2, 3...) Fail

At first, you might think: "Why not just add +1+1 to the first word, +2+2 to the second word, +3+3 to the third word, and +500+500 to the 500th word?"

Let us see why the simple ideas break down immediately:

Bad Idea #1: Adding Raw Integers (1, 2, 3 ... 1000)

Remember that word meaning vectors contain tiny decimals between −1.0-1.0 and +1.0+1.0 (like [0.42, -0.15]). If you add +500+500 to the 500th word in an essay, the vector becomes [500.42, 499.85]—the giant position number completely drowns out the meaning of the word!

Bad Idea #2: Fractions From 0.0 to 1.0

What if you divide by sentence length so the first word is 0.00.0 and the last word is 1.01.0? In a 4-word sentence, the step between words is 0.250.25. In a 100-word paragraph, the step between words shrinks to 0.010.01! A 1-word step would mean something totally different in every sentence!

What a Great Positional Encoding Needs (The 3 Rules)

1. Small Bounded Numbers: All values must stay neatly between −1.0-1.0 and +1.0+1.0 so they never crush the word's meaning.
2. Consistent Step Size: The distance from Word #2 to Word #3 must look the same whether the sentence is 5 words long or 5,000 words long.
3. Unique Pattern per Seat: Every position needs its own distinct fingerprint vector of length dmodeld_{\text{model}}.
⚡ Knowledge Check

How do standard Positional Encodings combine with a word's Meaning Embedding vector at the input of a Transformer?

A) They have the exact same dimension d_model and are added together element-by-element (+)▼
✓ Correct!If each word embedding has 512 numbers, its positional encoding also has 512 numbers, and we simply add them together (x = word_emb + pos_emb) so the vector size stays 512!
B) We replace the word embedding completely and throw the word's meaning away▼
✕ Incorrect.The model needs both what the word means and where it sits, so we add the small wave pattern on top of the meaning vector.

  1. Medium: Sinusoidal Positional Encodings (The Hands of a Clock)

How did the creators of the original 2017 Transformer ("Attention Is All You Need") solve all three rules at once? With a brilliant analogy: The Hands of a Clock!

Think about how a clock tells time using three hands spinning at different speeds:
• The Second Hand ticks super fast (changes on every single step).
• The Minute Hand moves at medium speed.
• The Hour Hand moves super slowly (only changes over long distances).

Sinusoidal Positional Encoding does the exact same thing using smooth Sine and Cosine waves (sin⁡\sin and cos⁡\cos), which naturally stay bounded between −1.0-1.0 and +1.0+1.0 forever!

The Classic Transformer Wave Formulas (Even vs. Odd Sliders)
PE(pos,  2i)=sin⁡(pos100002i/dmodel)andPE(pos,  2i+1)=cos⁡(pos100002i/dmodel)PE_{(pos, \; 2i)} = \sin\left(\dfrac{pos}{10000^{2i / d_{\text{model}}}}\right) \qquad \text{and} \qquad PE_{(pos, \; 2i+1)} = \cos\left(\dfrac{pos}{10000^{2i / d_{\text{model}}}}\right)
Early Sliders (i=0,1,2i = 0, 1, 2): Fast Waves

Like the Second Hand on a watch, the first few numbers in the vector wiggle up and down rapidly with every word step—helping the AI tell immediate neighbors (Position #4 vs. Position #5) apart.

Late Sliders (i=250+i = 250+): Slow Waves

Because of the 100002i/d10000^{2i/d} denominator, the later numbers in the vector wave very slowly—like the Hour Hand—helping the AI track which paragraph or chapter a word belongs to!

  1. Medium: Learned Positional Embeddings (Used in BERT & GPT-2)

Instead of using fixed math waves (sin⁡\sin and cos⁡\cos), models like BERT and GPT-2 took an even simpler approach called Learned Positional Embeddings:

They literally created a second nn.Embedding(max_seq_len, d_model) lookup table just for seat numbers (Seat #0, Seat #1, ..., Seat #1023) and let Backpropagation learn the best numbers for each seat during training!

MethodHow It WorksCan It Read Longer Texts Than Trained On?Famous Models
1. Sinusoidal WavesFixed sin⁡\sin and cos⁡\cos frequencies added to input embeddings (00 extra trainable parameters).Somewhat (waves continue forever)Original 2017 Transformer
2. Learned Position TableA trainable nn.Embedding(L_max, d) lookup table indexed by position 0,1,…,L−10, 1, \dots, L-1.No! Hard limit at Lmax⁡L_{\max} (e.g. 512 in BERT)BERT, RoBERTa, GPT-2, GPT-3
3. RoPE (Rotary Embeddings)Rotates Query and Key vectors by an angle mθm\theta inside every Attention layer!Yes! Scales to 128k–1M+ tokens!Llama-3, Mistral, DeepSeek, Gemma, Qwen
⚡ Knowledge Check

BERT was trained using a Learned Positional Embedding table of size nn.Embedding(512, 768). What happens if you try to pass a 600-token paragraph into BERT without truncating it?

A) It crashes with an IndexError because Row #512 to #599 do not exist in the position lookup table!▼
✓ Correct!A learned position table only has rows for positions 00 through 511511. Asking for Position #512 triggers an out-of-bounds table lookup error!
B) BERT automatically adds 88 new trained rows on the fly▼
✕ Incorrect.A fixed nn.Embedding table cannot invent trained weights for unseen row indices at test time.

  1. Advanced: RoPE (Rotary Position Embeddings) — How Modern LLMs Work

If you open the code for Llama-3, Mistral, DeepSeek, or Qwen, you will find they do not add position vectors at the start of the model! Why?

Think about human language: when you read a sentence, does it really matter if the phrase "blue sky" is at Word #10 and #11 versus Word #510 and #511? No! What matters is the Relative Distance between the two words (11−10=111 - 10 = 1 word apart)!

RoPE (Rotary Position Embedding, Su et al. 2021) encodes relative distance with a super clean geometric trick: instead of adding numbers to the word vector, RoPE rotates the word's vector like the hand of a clock by an angle m⋅θm \cdot \theta proportional to its seat number mm!

(q1rotatedq2rotated)=(cos⁡(mθ)amp;−sin⁡(mθ)sin⁡(mθ)amp;cos⁡(mθ))(q1q2)\begin{pmatrix} q_1^{\text{rotated}} \\ q_2^{\text{rotated}} \end{pmatrix} = \begin{pmatrix} \cos(m\theta) & -\sin(m\theta) \\ \sin(m\theta) & \cos(m\theta) \end{pmatrix} \begin{pmatrix} q_1 \\ q_2 \end{pmatrix}
Why Rotating Beats Adding!

Remember that Self-Attention compares Word mm and Word nn by measuring the angle between their vectors (dot product). If Word mm is rotated by mθm\theta and Word nn is rotated by nθn\theta, the angle between them depends ONLY on their relative gap (m−n)θ(m - n)\theta!

Context Extension (8k → 128k+ Tokens)

Because RoPE is just clock rotation, engineers can stretch a model's context window from 8,1928{,}192 tokens to 128,000+128{,}000+ tokens (NTK-Aware RoPE Scaling / YaRN) simply by slowing down the rotation speed θ\theta so more words fit around the clock face!

⚡ Knowledge Check

In RoPE (Rotary Position Embedding), if Word A at position m=10m = 10 is rotated by 10θ10\theta and Word B at position n=12n = 12 is rotated by 12θ12\theta, what is the angle difference between them? What if the same two words appear later at m=510m = 510 and n=512n = 512?

A) Both pairs have the exact same relative angle difference of 2*theta!▼
✓ Correct!Because 12θ−10θ=2θ12\theta - 10\theta = 2\theta and 512θ−510θ=2θ512\theta - 510\theta = 2\theta, Attention sees the exact same 2-word relative distance regardless of where the phrase appears in a 100,000-token book!
B) The second pair is 500 times farther apart in angle▼
✕ Incorrect.Because both vectors are rotated forward together around the circle, the angle between them only depends on their difference (n−m)=2(n - m) = 2.

  1. Visual Explanation: The Wave Fingerprint & Rotary Clock Console

Look at how a Transformer stamps seat numbers onto every word—comparing the classic Multi-Speed Wave Barcode (Sinusoidal) with modern RoPE Clock Dials (Llama-3):

PART 1 · SINUSOIDAL WAVE BARCODE (MULTI-SPEED CLOCK HANDS)

Each Seat Gets a Unique Wave Fingerprint!
SEAT #0"The"
Fast Wave:0.00
Med Wave:1.00
Slow Wave:0.00
SEAT #1"dog"
Fast Wave:+0.84
Med Wave:+0.54
Slow Wave:+0.01
SEAT #2"bit"
Fast Wave:+0.91
Med Wave:-0.42
Slow Wave:+0.02
SEAT #3"man"
Fast Wave:+0.14
Med Wave:-0.99
Slow Wave:+0.03

PART 2 · MODERN RoPE CLOCK DIALS (LLAMA-3 & DEEPSEEK)

Each Seat Rotates the Vector by +30°!
Seat #0 · 0°
"The"
Seat #1 · 30°
"dog"
Seat #2 · 60°
"bit"
Seat #3 · 90°
"man"

🧭 Click to See How RoPE Tells "Dog Bit Man" Apart From "Man Bit Dog"!

See Proof

▼

Sentence A: "The dog (#1) bit (#2) man (#3)"

"dog" is at 30° and "bit" is at 60° (+30° clockwise). The AI immediately sees "dog" came before "bit"!

Sentence B: "The man (#1) bit (#2) dog (#3)"

Now "dog" is rotated to 90° (-30° behind "bit")! The relative clock angle flipped from +30° to -30°, so Attention knows "dog" is now the victim!

  1. Python Implementation: Sinusoidal & RoPE Positional Encodings in PyTorch

Here is a clean, runnable PyTorch script that builds both the classic Sinusoidal Positional Encoding table and a modern 2D RoPE Rotation function so you can see how both work in less than 40 lines of code:

positional_encodings_demo.pyPython 3.11+ · PyTorch 2.x
import mathimport torch # 1. Classic Sinusoidal Positional Encoding (Attention Is All You Need, 2017)def buildSinusoidalPE(maxLen: int, dModel: int) -> torch.Tensor:    pe = torch.zeros(maxLen, dModel)    positions = torch.arange(0, maxLen, dtype=torch.float).unsqueeze(1)  # (maxLen, 1)    divTerm = torch.exp(torch.arange(0, dModel, 2).float() * (-math.log(10000.0) / dModel))     pe[:, 0::2] = torch.sin(positions * divTerm)  # Even sliders: Sine wave    pe[:, 1::2] = torch.cos(positions * divTerm)  # Odd sliders:  Cosine wave    return pe posTable = buildSinusoidalPE(maxLen=10, dModel=4)print("Seat #0 Wave Vector:", posTable[0].numpy().round(3))print("Seat #1 Wave Vector:", posTable[1].numpy().round(3)) # 2. Modern RoPE (Rotary Position Embedding): Rotate a 2D Vector by m * theta!def applyRope2D(vec: torch.Tensor, seatPos: int, theta: float = 0.5) -> torch.Tensor:    angle = seatPos * theta    cosA, sinA = math.cos(angle), math.sin(angle)    x1, x2 = vec[0], vec[1]    return torch.tensor([x1 * cosA - x2 * sinA, x1 * sinA + x2 * cosA]) # Proof: Dot product between Seat #1 & #3 equals Seat #101 & #103 (Both 2 seats apart!)qWord = torch.tensor([1.0, 0.5])kWord = torch.tensor([0.8, 0.2]) scoreNear = torch.dot(applyRope2D(qWord, 1),   applyRope2D(kWord, 3))scoreFar  = torch.dot(applyRope2D(qWord, 101), applyRope2D(kWord, 103))print("Attention Score (Seats 1 vs 3):",     round(scoreNear.item(), 5))print("Attention Score (Seats 101 vs 103):", round(scoreFar.item(), 5))  # Exact Match!

Pro Tip (Look at the Two RoPE Scores at the Bottom!):

Run the code above and look at the last two lines: comparing Seat #1 vs. Seat #3 gives the exact same attention score to 5 decimal places as comparing Seat #101 vs. Seat #103! That is why RoPE allows modern LLMs like Llama-3 to understand relative word distance effortlessly across huge documents.

Key Points

✓Because Transformers process all tokens in parallel without a sequential loop, they are permutation-equivariant ("blind to word order") unless we inject Positional Encodings.
✓Classic Positional Encodings have the exact same dimension dmodeld_{\text{model}} as the token embeddings and are added element-wise (xt=et+pt\mathbf{x}_t = \mathbf{e}_t + \mathbf{p}_t).
✓Sinusoidal Positional Encodings use multi-frequency sin⁡\sin and cos⁡\cos waves (like the second, minute, and hour hands of a clock) to keep values bounded in [−1,+1][-1, +1] across any sequence length.
✓Learned Positional Embeddings (used in BERT and GPT-2) store position vectors in a trainable nn.Embedding(L_max, d) table, but cannot extrapolate beyond Lmax⁡L_{\max}.
✓Modern LLMs (Llama-3, Mistral, DeepSeek, Gemma) use RoPE (Rotary Position Embeddings), which rotates Query and Key vectors in 2D pairs so their dot product depends purely on relative word distance (m−n)(m - n).

Common Mistakes

✕ Registering a fixed Sinusoidal Positional Encoding tensor as a trainable nn.Parameter.

Sinusoidal wave tables are fixed mathematical constants and should not be updated by the optimizer! In PyTorch, always register them with self.register_buffer("pe", pe) so they move to the GPU automatically without receiving gradients.

✕ Concatenating positional encodings to word embeddings instead of adding them.

Concatenating a 512-D position vector to a 512-D word vector doubles the hidden dimension to 1,024 and quadruples matrix sizes in every Transformer layer. Standard positional encodings are added (++), and RoPE is applied via in-place rotation.

✕ Applying RoPE to Value (V) vectors in Self-Attention.

RoPE is applied only to Query (Q\mathbf{Q}) and Key (K\mathbf{K}) vectors so that the attention score qmTkn\mathbf{q}_m^T \mathbf{k}_n captures relative distance; rotating the Value (V\mathbf{V}) vectors scrambles the actual semantic content being passed forward!

The Big Picture

Without Positional Encodings (Bag of Words)

Parallel Processing → Words Lose All Order → "Dog Bit Man" Equals "Man Bit Dog"

With Positional Encodings / RoPE (Parallel + Ordered!)

Stamp Every Word With Wave / Clock Angle → Process All Words at Once on GPU → Full Order Preserved!

The big takeaway is simple: Positional Encodings give us the best of both worlds. We get the blazing parallel GPU speed of processing an entire book at once, without losing track of which word came first, second, or last.

Remember: Now that our tokens have both Meaning (Word Embeddings) and Order (Positional Encodings), they are ready to enter the beating heart of the Transformer in our next lesson: Self-Attention!