The Core Idea (NeurIPS 2017 · Vaswani et al.): Before June 2017, every top AI system for language translation and sequence modeling was built on Recurrent Neural Networks (RNNs, LSTMs, and GRUs). Because recurrent models process text word-by-word from left to right ( must wait for ), they could not be parallelized across GPUs during training, and their memory of early words faded over long paragraphs. The "Attention Is All You Need" paper introduced the Transformer—the first sequence transduction architecture to eliminate recurrence and convolutions completely, relying entirely on Self-Attention and Position-wise Feed-Forward Networks to connect any two words in a sentence in a single direct step!
- High-Level Architecture: How the Encoder and Decoder Work Together
The original 2017 Transformer was designed for Sequence-to-Sequence (Seq2Seq) tasks like translating an English sentence into French. To do this, the architecture is divided into two main macro-components that work as a team:
Takes the entire input sequence of token symbols all at once and transforms them through stacked layers into a sequence of rich, context-aware continuous representations . Because the entire input sentence is already known, the Encoder processes all words simultaneously in parallel.
Receives the Encoder's contextual memory matrix and generates the output sequence one token at a time. The model is auto-regressive: at each step, it consumes the tokens it has already generated so far as additional input to predict the very next token.
The Input Pipeline Before Layer 1 (Embeddings + Scaling + Positional Waves)
- Scaled Token Embeddings: Input token IDs are converted into vectors of dimension using a learned embedding table. Notably, the paper multiplies the embedding weights by () before adding positional encodings so that the word's semantic meaning is not drowned out by the positional wave values!
- Sinusoidal Positional Encoding Addition: Because the Transformer contains zero recurrence and zero convolution, a -dimensional vector of multi-frequency and waves is added element-wise to inject exact token order.
- Inside the Encoder: The 2-Sublayer Architecture
The Encoder is composed of a stack of identical layers piled on top of one another. Every single Encoder layer maintains the exact same vector width () from bottom to top, and contains two distinct sub-layers:
In this sub-layer, all three inputs—Queries (), Keys (), and Values ()—come from the output of the previous Encoder layer. Every position in the Encoder can attend to every position in the entire input sentence (both left and right, with no causal mask). Across the parallel attention heads, each word gathers grammatical, coreference, and contextual clues from the whole sentence at once.
Once Self-Attention has mixed context across words, each word vector passes individually and identically through a 2-layer fully connected neural network. It expands the -dimensional vector to an inner dimension of ( wider), applies a non-linear activation, and projects it back down to .
Why did the authors force every single embedding layer, attention layer, and feed-forward output in the Transformer Base model to have the exact same dimension ()?
A) To facilitate the Residual Skip Connections (x + Sublayer(x)), which require identical vector dimensions for element-wise addition▼
B) Because GPUs cannot multiply matrices that are larger than 512▼
- Inside the Decoder: The 3-Sublayer Architecture & Output Generator
The Decoder is also composed of a stack of identical layers. However, unlike the Encoder (which has 2 sub-layers per block), each Decoder block contains THREE sub-layers, all wrapped with Residual Connections and Layer Normalization:
Allows each position in the Decoder to attend to all earlier generated positions up to and including itself. Future positions are masked out (set to before Softmax) combined with shifting target outputs right by one position, ensuring predictions for position depend only on known outputs at positions less than .
This is the bridge between the two towers! Here, the Queries () come from the previous Decoder sub-layer, while the Keys () and Values () come from the final output () of the Encoder stack. This lets every position in the Decoder attend over all positions in the input sequence!
Identical in structure to the Encoder's FFN ( with ). Processes the combined target context + source context at each position independently to prepare for predicting the next word.
The Final Top Cap: Linear Projection + Softmax Generator
- The Linear Layer (Unembedding Projection): Projects the -dimensional vector from the top of Decoder Layer #6 up to the size of the entire subword vocabulary ( logits).
- The Softmax Layer: Converts those raw logit scores into a valid probability distribution summing to , and selects the next output token!
- Residual Connections, Layer Normalization & Training Tricks
Around every single sub-layer in both the Encoder and Decoder, the authors wrapped a Residual Skip Connection () followed by Layer Normalization:
Without the skip connection, multiplying matrices across Encoder and Decoder layers ( sub-layers total!) causes gradients to vanish and destroys the token's original identity & position signal. With the wire, each layer only needs to learn a small refinement () added onto the main highway!
Repeatedly adding vectors across dozens of layers causes their magnitudes to grow uncontrollably. LayerNorm normalizes the features of each individual token to mean and variance after every step.
Ramps the learning rate up linearly over the first training steps before decaying by so early attention gradients never explode.
Softens target one-hot labels to during training so the model learns not to be overconfident, improving translation accuracy and BLEU scores.
Shares the exact same weight matrix between the Encoder embedding, the Decoder embedding, and the pre-softmax Linear layer!
- Complexity Comparison: Why Self-Attention Beat RNNs & CNNs
In Section 4 of the paper, the authors proved why Self-Attention is superior to Recurrent and Convolutional layers across three fundamental mathematical metrics:
| Layer Type (Paper Section 4) | Complexity per Layer | Sequential Waiting Operations | Maximum Path Length Between Tokens |
|---|---|---|---|
| Self-Attention | (All at once!) | (1 direct hop!) | |
| Recurrent (RNN / LSTM) | (Waits steps) | ( steps apart) | |
| Convolutional (Kernel ) |
According to Section 4 of the paper, why is it much easier for a Transformer to learn long-range dependencies between distant words in a sentence compared to an RNN or CNN?
A) Because the maximum path length between any two input/output positions in Self-Attention is a constant O(1)▼
B) Because Self-Attention deletes all words that are more than 5 positions away▼
- Visual Explanation: The Complete "Figure 1" Architecture Blueprint
Below is a complete, color-coded architectural map of Figure 1 from the 2017 paper. Follow the signal from the bottom of both towers all the way to the top Linear + Softmax output generator:
FINAL OUTPUT GENERATOR (TOP OF DECODER)
Linear Projection (512 → 37,000) ➔ Softmax ➔ Next Token Probabilities
🗼 ENCODER STACK (N = 6×)
d_model = 512
Sub-Layer 2: Position-Wise FFN
2-Layer MLP () applied to each token position independently.
Sub-Layer 1: Multi-Head Self-Attention
8 Parallel Heads (). Unmasked 360° vision across all input words ( all from Encoder).
Encoder Memory Z
Keys (K) & Values (V)
Top of Encoder #6 feeds and into Sub-Layer 2 of all 6 Decoder blocks!
🗼 DECODER STACK (N = 6×)
Auto-Regressive
Sub-Layer 3: Position-Wise FFN
Expands per token position.
Sub-Layer 2: Encoder-Decoder Attention
Queries () from Decoder Sub-Layer 1; Keys () & Values () from Encoder Tower!
Sub-Layer 1: Masked Multi-Head Attention
Blocks future positions () so token can only attend to already-generated words!
Key Points
The Big Picture
The Encoder Stack (Left Tower · 2 Sub-Layers × 6)
Input + Sine Waves → Bidirectional Self-Attention → 4x FFN → Exports Context Memory (K, V)
The Decoder Stack (Right Tower · 3 Sub-Layers × 6)
Shifted Outputs → Masked Self-Attention → Cross-Attention (Reads K, V) → 4x FFN → Linear + Softmax
The big takeaway is simple: the 2017 Transformer proved that Attention really is all you need. By separating cross-token communication (Multi-Head Attention) from per-token feature processing (Position-Wise FFN) and wiring them together along a clean Residual + LayerNorm highway, Vaswani et al. created the universal engine of modern Artificial Intelligence.