The Transformer Paper: Attention Is All You Need (Architecture Deep Dive)

A complete theoretical breakdown of the landmark 2017 'Attention Is All You Need' architecture—exploring the 6-layer Encoder stack, the 3-stage Decoder stack, Cross-Attention bridges, Residual Highways, LayerNorm, and Feed-Forward Networks.

24 minBeginner

The Core Idea (NeurIPS 2017 · Vaswani et al.): Before June 2017, every top AI system for language translation and sequence modeling was built on Recurrent Neural Networks (RNNs, LSTMs, and GRUs). Because recurrent models process text word-by-word from left to right (hth_t must wait for ht−1h_{t-1}), they could not be parallelized across GPUs during training, and their memory of early words faded over long paragraphs. The "Attention Is All You Need" paper introduced the Transformer—the first sequence transduction architecture to eliminate recurrence and convolutions completely, relying entirely on Self-Attention and Position-wise Feed-Forward Networks to connect any two words in a sentence in a single O(1)\mathcal{O}(1) direct step!

  1. High-Level Architecture: How the Encoder and Decoder Work Together

The original 2017 Transformer was designed for Sequence-to-Sequence (Seq2Seq) tasks like translating an English sentence into French. To do this, the architecture is divided into two main macro-components that work as a team:

1. The Encoder Stack (Left Tower)The Reader

Takes the entire input sequence of token symbols (x1,x2,…,xn)(x_1, x_2, \dots, x_n) all at once and transforms them through N=6N = 6 stacked layers into a sequence of rich, context-aware continuous representations Z=(z1,z2,…,zn)\mathbf{Z} = (\mathbf{z}_1, \mathbf{z}_2, \dots, \mathbf{z}_n). Because the entire input sentence is already known, the Encoder processes all words simultaneously in parallel.

2. The Decoder Stack (Right Tower)The Writer

Receives the Encoder's contextual memory matrix Z\mathbf{Z} and generates the output sequence (y1,y2,…,ym)(y_1, y_2, \dots, y_m) one token at a time. The model is auto-regressive: at each step, it consumes the tokens it has already generated so far as additional input to predict the very next token.

The Input Pipeline Before Layer 1 (Embeddings + dmodel\sqrt{d_{\text{model}}} Scaling + Positional Waves)

  1. Scaled Token Embeddings: Input token IDs are converted into vectors of dimension dmodel=512d_{\text{model}} = 512 using a learned embedding table. Notably, the paper multiplies the embedding weights by dmodel\sqrt{d_{\text{model}}} (512≈22.6\sqrt{512} \approx 22.6) before adding positional encodings so that the word's semantic meaning is not drowned out by the [−1,+1][-1, +1] positional wave values!
  1. Sinusoidal Positional Encoding Addition: Because the Transformer contains zero recurrence and zero convolution, a 512512-dimensional vector of multi-frequency sin⁡\sin and cos⁡\cos waves is added element-wise to inject exact token order.

  1. Inside the Encoder: The 2-Sublayer Architecture

The Encoder is composed of a stack of N=6N = 6 identical layers piled on top of one another. Every single Encoder layer maintains the exact same vector width (dmodel=512d_{\text{model}} = 512) from bottom to top, and contains two distinct sub-layers:

Encoder Sub-Layer 1
Bidirectional Multi-Head Self-Attention

In this sub-layer, all three inputs—Queries (Q\mathbf{Q}), Keys (K\mathbf{K}), and Values (V\mathbf{V})—come from the output of the previous Encoder layer. Every position in the Encoder can attend to every position in the entire input sentence (both left and right, with no causal mask). Across the 88 parallel attention heads, each word gathers grammatical, coreference, and contextual clues from the whole sentence at once.

Encoder Sub-Layer 2
Position-Wise Feed-Forward Network (FFN)

Once Self-Attention has mixed context across words, each word vector passes individually and identically through a 2-layer fully connected neural network. It expands the 512512-dimensional vector to an inner dimension of dff=2,048d_{\text{ff}} = 2{,}048 (4×4\times wider), applies a non-linear ReLU\text{ReLU} activation, and projects it back down to 512512.

Mathematical Flow Through One Encoder Block
H=LayerNorm(X+MultiHeadAttention(X,X,X))\mathbf{H} = \text{LayerNorm}\left(\mathbf{X} + \text{MultiHeadAttention}(\mathbf{X}, \mathbf{X}, \mathbf{X})\right)
Z=LayerNorm(H+FFN(H))whereFFN(h)=max⁡(0,  hW1+b1)W2+b2\mathbf{Z} = \text{LayerNorm}\left(\mathbf{H} + \text{FFN}(\mathbf{H})\right) \quad \text{where} \quad \text{FFN}(\mathbf{h}) = \max(0, \; \mathbf{h}\mathbf{W}_1 + \mathbf{b}_1)\mathbf{W}_2 + \mathbf{b}_2
⚡ Knowledge Check

Why did the authors force every single embedding layer, attention layer, and feed-forward output in the Transformer Base model to have the exact same dimension (dmodel=512d_{\text{model}} = 512)?

A) To facilitate the Residual Skip Connections (x + Sublayer(x)), which require identical vector dimensions for element-wise addition▼
✓ Correct!As stated directly in Section 3.1 of the paper: "To facilitate these residual connections, all sub-layers in the model, as well as the embedding layers, produce outputs of dimension dmodel=512d_{\text{model}} = 512."
B) Because GPUs cannot multiply matrices that are larger than 512▼
✕ Incorrect.Inside the FFN sub-layer, the vector actually expands to 2,0482{,}048 before projecting back down to 512512 so it can be added back to the 512512-D residual stream.

  1. Inside the Decoder: The 3-Sublayer Architecture & Output Generator

The Decoder is also composed of a stack of N=6N = 6 identical layers. However, unlike the Encoder (which has 2 sub-layers per block), each Decoder block contains THREE sub-layers, all wrapped with Residual Connections and Layer Normalization:

Decoder Sub-Layer 1
Masked Multi-Head Self-Attention

Allows each position in the Decoder to attend to all earlier generated positions up to and including itself. Future positions are masked out (set to −∞-\infty before Softmax) combined with shifting target outputs right by one position, ensuring predictions for position ii depend only on known outputs at positions less than ii.

Decoder Sub-Layer 2
Encoder-Decoder Cross-Attention

This is the bridge between the two towers! Here, the Queries (Q\mathbf{Q}) come from the previous Decoder sub-layer, while the Keys (K\mathbf{K}) and Values (V\mathbf{V}) come from the final output (Z\mathbf{Z}) of the Encoder stack. This lets every position in the Decoder attend over all positions in the input sequence!

Decoder Sub-Layer 3
Position-Wise Feed-Forward (FFN)

Identical in structure to the Encoder's FFN (512→2048→512512 \to 2048 \to 512 with ReLU\text{ReLU}). Processes the combined target context + source context at each position independently to prepare for predicting the next word.

The Final Top Cap: Linear Projection + Softmax Generator

  1. The Linear Layer (Unembedding Projection): Projects the 512512-dimensional vector from the top of Decoder Layer #6 up to the size of the entire subword vocabulary (∣V∣=37,000|\mathcal{V}| = 37{,}000 logits).
  1. The Softmax Layer: Converts those 37,00037{,}000 raw logit scores into a valid probability distribution summing to 100%100\%, and selects the next output token!

  1. Residual Connections, Layer Normalization & Training Tricks

Around every single sub-layer in both the Encoder and Decoder, the authors wrapped a Residual Skip Connection (x+Sublayer(x)\mathbf{x} + \text{Sublayer}(\mathbf{x})) followed by Layer Normalization:

1. Why the Residual (++) Wire?

Without the skip connection, multiplying matrices across 66 Encoder and 66 Decoder layers (30+30+ sub-layers total!) causes gradients to vanish and destroys the token's original identity & position signal. With the ++ wire, each layer only needs to learn a small refinement (Δx\Delta \mathbf{x}) added onto the main highway!

2. Why LayerNorm?

Repeatedly adding vectors across dozens of layers causes their magnitudes to grow uncontrollably. LayerNorm normalizes the 512512 features of each individual token to mean 00 and variance 11 after every step.

1. 4,000-Step LR Warmup

Ramps the learning rate up linearly over the first 4,0004{,}000 training steps before decaying by step−0.5\text{step}^{-0.5} so early attention gradients never explode.

2. Label Smoothing (ϵls=0.1\epsilon_{ls} = 0.1)

Softens target one-hot labels to 0.90.9 during training so the model learns not to be overconfident, improving translation accuracy and BLEU scores.

3. 3-Way Weight Sharing

Shares the exact same 37,000×51237{,}000 \times 512 weight matrix between the Encoder embedding, the Decoder embedding, and the pre-softmax Linear layer!

  1. Complexity Comparison: Why Self-Attention Beat RNNs & CNNs

In Section 4 of the paper, the authors proved why Self-Attention is superior to Recurrent and Convolutional layers across three fundamental mathematical metrics:

Layer Type (Paper Section 4)Complexity per LayerSequential Waiting OperationsMaximum Path Length Between Tokens
Self-AttentionO(n2⋅d)\mathcal{O}(n^2 \cdot d)O(1)\mathcal{O}(1) (All at once!)O(1)\mathcal{O}(1) (1 direct hop!)
Recurrent (RNN / LSTM)O(n⋅d2)\mathcal{O}(n \cdot d^2)O(n)\mathcal{O}(n) (Waits nn steps)O(n)\mathcal{O}(n) (nn steps apart)
Convolutional (Kernel kk)O(k⋅n⋅d2)\mathcal{O}(k \cdot n \cdot d^2)O(1)\mathcal{O}(1)O(log⁡k(n))\mathcal{O}(\log_k(n))
⚡ Knowledge Check

According to Section 4 of the paper, why is it much easier for a Transformer to learn long-range dependencies between distant words in a sentence compared to an RNN or CNN?

A) Because the maximum path length between any two input/output positions in Self-Attention is a constant O(1)▼
✓ Correct!The shorter the path forward and backward signals must travel between any two positions in the network, the easier it is to learn long-range dependencies. In Self-Attention, any two tokens interact directly in 11 step (O(1)\mathcal{O}(1))!
B) Because Self-Attention deletes all words that are more than 5 positions away▼
✕ Incorrect.Self-Attention compares every token across the entire sequence without distance limits.

  1. Visual Explanation: The Complete "Figure 1" Architecture Blueprint

Below is a complete, color-coded architectural map of Figure 1 from the 2017 paper. Follow the signal from the bottom of both towers all the way to the top Linear + Softmax output generator:

FINAL OUTPUT GENERATOR (TOP OF DECODER)

Linear Projection (512 → 37,000) ➔ Softmax ➔ Next Token Probabilities

🗼 ENCODER STACK (N = 6×)

2 Sub-Layers per Block

d_model = 512

Add & LayerNormLayerNorm(x + FFN(x))

Sub-Layer 2: Position-Wise FFN

2-Layer MLP (512→ReLU2048→512512 \xrightarrow{\text{ReLU}} 2048 \to 512) applied to each token position independently.

Add & LayerNormLayerNorm(x + MHA(x))

Sub-Layer 1: Multi-Head Self-Attention

8 Parallel Heads (dk=64d_k = 64). Unmasked 360° vision across all input words (Q,K,VQ, K, V all from Encoder).

⊕ Sinusoidal Positional Encoding
↑ Input Embeddings (×512\times \sqrt{512})
Source Input Tokens (x1,…,xn)(x_1, \dots, x_n)

Encoder Memory Z

➔

Keys (K) & Values (V)

Top of Encoder #6 feeds KK and VV into Sub-Layer 2 of all 6 Decoder blocks!

🗼 DECODER STACK (N = 6×)

3 Sub-Layers per Block

Auto-Regressive

Add & LayerNormLayerNorm(x + FFN(x))

Sub-Layer 3: Position-Wise FFN

Expands 512→2048→512512 \to 2048 \to 512 per token position.

Add & LayerNormCross-Attention Bridge

Sub-Layer 2: Encoder-Decoder Attention

Queries (QQ) from Decoder Sub-Layer 1; Keys (KK) & Values (VV) from Encoder Tower!

Add & LayerNorm🔒 Causal Mask (-∞)

Sub-Layer 1: Masked Multi-Head Attention

Blocks future positions (j>ij \gt i) so token ii can only attend to already-generated words!

⊕ Sinusoidal Positional Encoding
↑ Output Embeddings (Shifted Right)
Starts with <BOS> token

Key Points

✓The 2017 Transformer is an Encoder-Decoder architecture (N=6N = 6 layers each, dmodel=512d_{\text{model}} = 512, H=8H = 8 heads, dff=2,048d_{\text{ff}} = 2{,}048) that relies entirely on Attention and Feed-Forward layers with zero recurrence.
✓Each Encoder block has 2 sub-layers: (1) Bidirectional Multi-Head Self-Attention and (2) a Position-Wise Feed-Forward Network (512→2048→512512 \to 2048 \to 512), each wrapped with a Residual Connection and LayerNorm.
✓Each Decoder block has 3 sub-layers: (1) Masked Multi-Head Self-Attention (which blocks future positions with −∞-\infty), (2) Encoder-Decoder Cross-Attention (QQ from Decoder, KK & VV from Encoder), and (3) a Position-Wise FFN.
✓Self-Attention connects any two tokens in O(1)\mathcal{O}(1) sequential operations and O(1)\mathcal{O}(1) path length, enabling massive GPU parallelization during training.
✓This single paper gave birth to all three modern Transformer families: BERT (the Encoder stack), GPT / Llama / Claude (the Masked Decoder stack), and T5 / Whisper (the full Encoder-Decoder stack).

The Big Picture

The Encoder Stack (Left Tower · 2 Sub-Layers × 6)

Input + Sine Waves → Bidirectional Self-Attention → 4x FFN → Exports Context Memory (K, V)

The Decoder Stack (Right Tower · 3 Sub-Layers × 6)

Shifted Outputs → Masked Self-Attention → Cross-Attention (Reads K, V) → 4x FFN → Linear + Softmax

The big takeaway is simple: the 2017 Transformer proved that Attention really is all you need. By separating cross-token communication (Multi-Head Attention) from per-token feature processing (Position-Wise FFN) and wiring them together along a clean Residual + LayerNorm highway, Vaswani et al. created the universal engine of modern Artificial Intelligence.