The Core Idea: When the original Transformer was invented in 2017 to translate English into French, it had two halves: an Encoder (the Reader that understands the full input sentence at once) and a Decoder (the Writer that generates the output one word at a time). Soon after, AI researchers realized you could split the machine in half—creating Encoder-Only models (like BERT) for super-smart reading and search, and Decoder-Only models (like ChatGPT and Llama-3) for writing and conversation!
- Beginner: The Three Transformer Families (Reader vs. Writer vs. Translator)
Imagine three different language jobs in the real world. Each job requires a different way of looking at a page of text:
Can look both left and right across the entire sentence at once! Great for reading comprehension, spam detection, and AI search embeddings—but cannot write essays.
Wears a Causal Blinder so it can only look backward at past words while predicting the very next word. Powers almost all modern chatbots and coding AIs!
Uses an Encoder to read the whole input (like Spanish text or audio) and a Decoder to write the English translation word-by-word!
| Architecture Family | Where Can Attention Look? | How Is It Trained? | Best Real-World Use Cases |
|---|---|---|---|
| Encoder-Only (BERT) | Bidirectional (Past + Future) | Fill in the blank ([MASK]) | RAG Vector Embeddings, Sentiment, Named Entity Tagging |
| Decoder-Only (GPT / Llama) | Causal (Past & Present Only) | Predict the very next token | Chatbots, Code Generation, Reasoning Agents |
| Encoder-Decoder (T5 / Whisper) | Encoder: Both · Decoder: Past + Encoder | Sequence-to-Sequence mapping | Machine Translation, Speech-to-Text (Whisper) |
- Beginner: How an Encoder Works (Bidirectional Fill-in-the-Blank)
Why is an Encoder (like BERT) so good at understanding what a sentence means? Because it has zero blinders on its eyes!
Suppose you see this sentence with a missing word in the middle:
"The surgeon walked into the [MASK] room to perform the heart operation."
To guess that the missing word is "operating", your brain looks at "surgeon" on the left AND "heart operation" on the right!
In an Encoder block, every token is allowed to attend to all tokens before it AND all tokens after it. There is no Causal Mask blocking the upper-right triangle!
Because BERT was trained with all words visible at once (except hidden behind [MASK]), it does not naturally generate text word-by-word from left to right. Instead, it outputs super-rich vectors for search engines and classifiers!
If you are building a RAG search system and need an AI model to convert 10,000 help-center articles into dense meaning vectors for Cosine Similarity search, which Transformer family is designed specifically for that job?
A) An Encoder-Only model (like BERT, RoBERTa, or BGE) with bidirectional attention▼
B) A model that blocks words from seeing the rest of the paragraph▼
- Medium: Causal Masking & Why Decoders Train Fast in Parallel!
Now let us look at Decoder-Only models (GPT-4, Claude, Llama-3). Here is the single coolest trick in modern AI—how a Causal Mask lets a Decoder learn prediction tasks in single GPU step!
Suppose our training sentence is "AI is super cool". Because of the Upper-Triangular Causal Mask (), every row of the matrix sees a different prefix of the sentence and predicts the next word simultaneously:
4 Training Examples Inside 1 Sentence (Teacher Forcing!)
Look at why that is magical: During Training, the GPU processes all rows in parallel matrix multiplication without any future cheating! Then during live Chat Generation (Inference), the model generates one new token at a time, appends it to the end, and repeats!
- Medium: Cross-Attention (The Bridge Between Encoder and Decoder)
What if we are using a full Encoder-Decoder model—like OpenAI Whisper (which transcribes spoken audio into English text) or the original 2017 Translation Transformer?
How does the Decoder (writing the English words) look back at what the Encoder read (the French sentence or audio clip)?
Through a special middle layer inside every Decoder block called Cross-Attention (Encoder-Decoder Attention)! Look at where , , and come from in Cross-Attention:
The Decoder asks: "I have written 'The cat sat' in English so far—what French word should I translate next?"
The Encoder holds up the nametags () and meanings () of the original French words so the Decoder can pull the exact translation across the bridge!
In an Encoder-Decoder Transformer's Cross-Attention layer, which side provides the Query () and which side provides the Keys () and Values ()?
A) Queries (Q) come from the Decoder; Keys (K) and Values (V) come from the Encoder's final output▼
B) All three (Q, K, V) come only from the Encoder▼
- Advanced: Why Decoder-Only Models Won the LLM Race & The 2-Mask Rule
At the Advanced / LLM Engineering level, two crucial insights explain how modern production Transformers are built:
Why are GPT-4, Claude, Llama-3, and DeepSeek all Decoder-Only instead of Encoder-Decoder? Because in a Decoder-Only model, of the parameters are used on every token, and every task—translation, summarization, coding, Q&A—can simply be formatted as next-token prediction after a prompt! Even simpler: there is only one stack of layers and one KV Cache to manage!
In real training batches, sentences have different lengths (so shorter sentences have <PAD> tokens), AND we need the Causal Mask! Therefore, a production Decoder combines TWO masks using logical AND: block position with if (Future Token) OR token is a <PAD> token!
Why do we use Left-Padding (putting <PAD> tokens on the left: [PAD, PAD, "Hi", "there"]) instead of Right-Padding when running batched generation with a Decoder-Only LLM?
A) Because a Decoder generates the next word immediately after the rightmost token in the row!▼
["Hi", "there", PAD, PAD]), the last token in the row is PAD instead of "there"! Left-padding aligns every prompt's final real word along the rightmost column so the whole batch can generate the next word together!B) Because GPUs read matrices from right to left▼
-1 where autoregressive sampling happens.
- Visual Explanation: The 3-Way Vision Goggles & Cross-Attention Bridge
Put on the AI Vision Goggles below to see the exact difference between how an Encoder (BERT), a Causal Decoder (GPT/Llama), and an Encoder-Decoder Bridge (Whisper/T5) see the exact same sentence when sitting at Word #3 ("sat"):
🕶️ GOGGLES #1 · ENCODER-ONLY (BERT) — 360° BIDIRECTIONAL VISION
All 5 Words Visible at Once!🕶️ GOGGLES #2 · DECODER-ONLY (GPT-4 / LLAMA-3) — CAUSAL MASK BLINDER
Future Words Blocked With -∞!🌉 GOGGLES #3 · ENCODER-DECODER CROSS-ATTENTION BRIDGE (T5 / WHISPER)
Q from Decoder ➔ K, V from Encoder⇄
🎯 Click for the Architecture Decision Guide: Which One Should You Choose?
Decision Guide▼
🎯 Click for the Architecture Decision Guide: Which One Should You Choose?
▼
Pick Encoder-Only when you want fast, cheap vectors for RAG search, document clustering, or sentiment classification.
Pick Decoder-Only when building a general AI assistant, code generator, or reasoning agent.
Pick Encoder-Decoder when converting one modality/format into another (like Audio → Text in Whisper).
- Python Implementation: Encoder vs. Causal Decoder & Combined Masking in PyTorch
Here is a clean, runnable PyTorch script showing how to build a Combined Causal + Padding Mask and how to prove that changing the last word of a sentence changes every word's vector in an Encoder, while leaving earlier words untouched in a Causal Decoder:
import torchimport torch.nn.functional as F # 1. Build a Combined Causal + Padding Mask for a Batchdef createDecoderMask(padMask: torch.Tensor) -> torch.Tensor: # padMask shape: (Batch, SeqLen) where True = Real Token, False = <PAD> B, T = padMask.shape causalAllowed = torch.tril(torch.ones(T, T, dtype=torch.bool)) # Lower triangle = True keyIsReal = padMask.unsqueeze(1).expand(B, T, T) # (B, T, T) return causalAllowed & keyIsReal # Both must be True! samplePadMask = torch.tensor([[True, True, True, False]]) # 3 real words + 1 PADcombinedMask = createDecoderMask(samplePadMask)print("Allowed Attention Grid (1=Allowed, 0=Blocked):\n", combinedMask[0].int().numpy()) # 2. Prove Encoder (Bidirectional) vs. Decoder (Causal) Behavior!torch.manual_seed(7)seqA = torch.randn(1, 4, 8)seqB = seqA.clone()seqB[0, 3, :] = torch.randn(8) # Change ONLY the 4th (final) word! # Run Bidirectional Encoder Attention (is_causal=False)encOutA = F.scaled_dot_product_attention(seqA, seqA, seqA, is_causal=False)encOutB = F.scaled_dot_product_attention(seqB, seqB, seqB, is_causal=False) # Run Causal Decoder Attention (is_causal=True)decOutA = F.scaled_dot_product_attention(seqA, seqA, seqA, is_causal=True)decOutB = F.scaled_dot_product_attention(seqB, seqB, seqB, is_causal=True) print("Did Word #1 change in Encoder when Word #4 changed?", not torch.allclose(encOutA[0, 0], encOutB[0, 0]))print("Did Word #1 change in Decoder when Word #4 changed?", not torch.allclose(decOutA[0, 0], decOutB[0, 0]))Pro Tip (Look at the Two True/False Prints at the Bottom!):
When we change Word #4 at the end of the sentence, Word #1 in the Encoder changes immediately (True) because it has 360° vision. But Word #1 in the Causal Decoder does not change at all (False)—proving that the Causal Mask completely blocks information from flowing backward from the future!
Key Points
Common Mistakes
✕ Training a next-word Decoder without enabling the Causal Mask.
If you forget is_causal=True during training, your loss drops to in two minutes because every token simply looks one seat to the right and copies the ground-truth answer—and then the model outputs total gibberish at test time!
✕ Using a Causal Mask inside Cross-Attention.
In an Encoder-Decoder model, the Causal Mask is only used inside the Decoder's Self-Attention layer. In Cross-Attention, the Decoder is allowed to look at the entire finished Encoder input (all French words or all audio frames) freely!
✕ Confusing PyTorch's boolean mask convention (True = Keep vs. True = Mask Out).
In PyTorch's F.scaled_dot_product_attention, a boolean mask means True = Allowed to Attend, whereas in nn.MultiheadAttention, True = Blocked/Ignored! Always verify with a 2-line test or pass a float mask with 0.0 and -inf.
The Big Picture
Encoder-Only (BERT · The 360° Reader)
No Causal Mask → Sees Past & Future Simultaneously → Best for Search Embeddings & Classification
Decoder-Only (GPT-4 & Llama-3 · The Causal Writer)
Causal Mask (-∞ on Future) → Trains in Parallel on GPUs → Generates Answers Token-by-Token!
The big takeaway is simple: the only structural difference between a Reader (Encoder) and a Writer (Decoder) is a triangle of numbers! Remove the mask and the Transformer reads the whole page at once; put the Causal Mask on and it becomes an autoregressive storyteller.
Remember: Attention lets words gather context from each other, but what else sits inside a Transformer layer to actually process and store facts? In our next lesson, we will open up the full Transformer Block: LayerNorm, Residual Connections, and Feed-Forward Networks (FFNs)!