Self-Attention and Query, Key, Value (Q, K, V) Explained Simply

Learn how Self-Attention lets words talk directly to each other across a sentence—from the YouTube Search analogy for Query, Key, and Value (Q, K, V) to Scaled Dot-Product Attention and Causal Masks.

25 minBeginnerCode Examples

The Core Idea: When you read the sentence "The dog chased the ball because it was fast," your brain instantly connects the word "it" back to "the dog". Old sequence models (RNNs) struggled with this because they had to pass a single sticky-note message through every word in between. Self-Attention throws out the middlemen: it lets every word in a sentence shine a direct spotlight on every other word at the exact same time to gather the exact context it needs!

  1. Beginner: The Mystery Pronoun Problem (Why Words Need to Talk)

Look at these two nearly identical sentences where only the very last word changes:

Sentence A

"The animal didn't cross the street because it was too tired."

Who was tired? The animal! (Streets don't get sleepy.)

Sentence B

"The animal didn't cross the street because it was too wide."

What was wide? The street!

Think about the tiny word "it". By itself in a dictionary, "it" is a blank placeholder—it has no meaning until it looks around the sentence to see what else is there!

Self-Attention is the mechanism that lets the word "it" look at every word in the sentence, score how relevant each word is (giving 85%85\% attention to "animal" and 2%2\% to "street" in Sentence A), and absorb their meanings directly into its own vector!

  1. Beginner: What Are Query, Key, and Value (Q, K, V)? (The YouTube Search Analogy)

Every textbook says: "Self-Attention projects each token into a Query, a Key, and a Value." Why those three weird names?

Because Self-Attention works exactly like searching for a video on YouTube! Imagine how a search engine works in 3 simple steps:

1. Query (Q)
"What I Am Looking For"

Like typing "fluffy pet that meows" into the YouTube search bar. The word "it" sends out a Query asking: "Hey, is there a tired noun earlier in this sentence?"

2. Key (K)
"My Name Tag / Title"

Like the title tags on every YouTube video. Every word in the sentence holds up a Key badge announcing what it is: "animal" holds up a tag saying "I am a living noun!"

3. Value (V)
"The Actual Video Content"

Once your Query matches a video's Key title, what do you actually watch? The Value! It is the actual rich meaning payload that gets passed over to update the word.

How One Word Creates All Three Vectors (Q, K, and V)

Every word starts with its single input vector x\mathbf{x} (Meaning + Position). The Transformer simply multiplies x\mathbf{x} by three learnable weight matrices (WQ,WK,WV\mathbf{W}_Q, \mathbf{W}_K, \mathbf{W}_V) to create that word's personal Query (q\mathbf{q}), Key (k\mathbf{k}), and Value (v\mathbf{v})!

⚡ Knowledge Check

In the YouTube search analogy for Self-Attention, which two vectors are multiplied together first to calculate the "relevance match score" between Word A and Word B?

A) Word A's Query (Q) and Word B's Key (K)▼
✓ Correct!Just like matching your search Query against a video's Key title, we take the dot product of Q\mathbf{Q} and KT\mathbf{K}^T to see how strongly two words match!
B) Word A's Value (V) and Word B's Value (V)▼
✕ Incorrect.The Value (V\mathbf{V}) vectors are only blended together after Query and Key have already calculated the percentage match scores!

  1. Medium: The 4-Step Recipe of Scaled Dot-Product Attention

Now that you know what Q\mathbf{Q}, K\mathbf{K}, and V\mathbf{V} are, the famous equation from the 2017 "Attention Is All You Need" paper is super easy to read. It is just a 4-step kitchen recipe:

Attention(Q,K,V)=softmax(QKTdk)V\text{Attention}(\mathbf{Q}, \mathbf{K}, \mathbf{V}) = \text{softmax}\left(\dfrac{\mathbf{Q}\mathbf{K}^T}{\sqrt{d_k}}\right)\mathbf{V}
Step 1: Score Every Pair (QKT\mathbf{Q}\mathbf{K}^T)

Multiply every word's Query by every word's Key. If a sentence has 55 words, this creates a 5×55 \times 5 grid of raw match scores.

Step 2: Divide by dk\sqrt{d_k} (Cool Down!)

If vectors have dk=64d_k = 64 numbers, their dot products can get huge, which pushes Softmax into flat regions where gradients vanish! Dividing by 64=8\sqrt{64} = 8 keeps scores calm and balanced.

Step 3: Softmax (Turn into 100% Pie)

Applies Softmax across each row so all attention weights are positive percentages that add up to 1.01.0 (100%100\%)!

Step 4: Blend the Values (×V\times \mathbf{V})

Multiply each word's Value vector by its percentage weight (e.g., 0.80×vanimal+0.20×vtired0.80 \times \mathbf{v}_{\text{animal}} + 0.20 \times \mathbf{v}_{\text{tired}}) and sum them up!

  1. Medium: Causal Masking (How GPT Stops Cheating From the Future!)

When we train a model like GPT-4 or Llama-3 to predict the next word in "The cat sat on the mat", we feed all 66 words into the GPU at the same time for speed.

Wait a second! If all 66 words are in the matrix at the same time, when Word #2 ("cat") tries to guess Word #3 ("sat"), what stops "cat" from just looking ahead at Seat #3 and copying the answer?

The Fix: The Upper-Triangular Causal Mask (−∞-\infty)

Right before Step 3 (Softmax), we apply a Causal Mask (Look-Ahead Blinder): for every word at seat ii, we overwrite the attention scores of all future words (j>ij \gt i) with −∞-\infty (negative infinity)!

Why −∞-\infty? Because in Softmax, e−∞=0.0e^{-\infty} = 0.0! Every future word gets exact 0%0\% attention, so Word #2 can only see Word #1 and Word #2—never the future!

⚡ Knowledge Check

Why do we fill masked future positions with −∞-\infty (negative infinity) BEFORE running Softmax, instead of filling them with 00?

A) Because in Softmax, exp(-infinity) = 0.0 (zero attention), whereas exp(0) = 1.0 (which would give positive attention to future words!)▼
✓ Correct!Softmax exponentiates every score (exe^x). Setting a score to 00 before Softmax makes e0=1e^0 = 1, leaking information from the future! Setting it to −∞-\infty guarantees e−∞=0.0e^{-\infty} = 0.0 after Softmax.
B) Because PyTorch tensors cannot store the number 0▼
✕ Incorrect.Tensors store zeros easily; the reason is purely mathematical because e0=1e^0 = 1 inside the Softmax numerator.

  1. Advanced: The O(T2)\mathcal{O}(T^2) Memory Bottleneck & FlashAttention

If Self-Attention is so great, what is its biggest weakness in modern Large Language Models?

1. The O(T2)\mathcal{O}(T^2) Quadratic Explosion

Because every word compares itself with every other word, a sequence of length TT creates a T×TT \times T Attention Matrix:
• 1,0001{,}000 tokens = 1 Million1\text{ Million} pairs
• 100,000100{,}000 tokens = 10 Billion10\text{ Billion} pairs per head!
Writing that massive T×TT \times T grid out to slow GPU RAM (HBM) is what makes long-context LLMs run out of memory!

2. FlashAttention (Dao et al., 2022)

How do modern models read 128k+128\text{k}+ tokens without crashing? FlashAttention splits Q,K,V\mathbf{Q}, \mathbf{K}, \mathbf{V} into tiny blocks inside ultra-fast GPU chip cache (SRAM) and computes the exact same Softmax result without ever saving the giant T×TT \times T matrix to GPU memory—giving a 3–5×3\text{--}5\times speedup!

  1. Visual Explanation: The Self-Attention Spotlight & Causal Mask Matrix

Explore our two-part visual studio below: first, watch the Attention Spotlight resolve what the word "it" means in real time; second, inspect the Causal Mask Grid that stops GPT from peeking into the future:

🔦 LIVE ATTENTION SPOTLIGHT · QUERY WORD = "it"

Thickness = Softmax Attention %
"The" (2%)"animal" (76%)"street" (4%)"because" (3%)"tired" (15%)🔍 Query: "it"

New Vector for "it" = 0.76 × Vec("animal") + 0.15 × Vec("tired") + 0.09 × (others)

🔒 4×4 GPT CAUSAL ATTENTION HEATMAP

🚫 = Masked (-∞ → 0%)
Query ↓ \ Key →
"The"
"cat"
"sat"
"down"
"The"
100%
🚫 0%
🚫 0%
🚫 0%
"cat"
30%
70%
🚫 0%
🚫 0%
"sat"
10%
65%
25%
🚫 0%
"down"
5%
35%
45%
15%
How to Read This Grid:
  1. Pick any word on the left (Row = Query).
  2. Read across to see how much percentage attention it pays to earlier words (Columns = Keys).
  3. Notice that every single row adds up to 100%, and the upper-right triangle is blocked (🚫) so words can never peek at future words!

🔓 What if we are using BERT instead of GPT?

Reveal

▼

BERT (Encoder) removes the red 🚫 blinder completely! Because BERT is used for reading comprehension and search embeddings (not guessing the next word), every word is allowed to look both backward and forward across the whole sentence!

  1. Python Implementation: Scaled Dot-Product Self-Attention in PyTorch

Here is a clean, runnable PyTorch script that implements the 4-step Scaled Dot-Product Self-Attention formula (with an optional Causal Mask) from scratch in 25 lines of code:

self_attention_qkv.pyPython 3.11+ · PyTorch 2.x
import mathimport torchimport torch.nn as nnimport torch.nn.functional as F class SingleHeadSelfAttention(nn.Module):    def init(self, dModel: int):        super().init()        self.wQ = nn.Linear(dModel, dModel, bias=False)  # Creates Query (What I look for)        self.wK = nn.Linear(dModel, dModel, bias=False)  # Creates Key   (My label tag)        self.wV = nn.Linear(dModel, dModel, bias=False)  # Creates Value (My content)     def forward(self, x: torch.Tensor, causalMask: bool = True):        B, T, dK = x.shape        Q, K, V = self.wQ(x), self.wK(x), self.wV(x)         # Steps 1 & 2: Compute Match Scores (Q @ K^T) and Scale by sqrt(d_k)        scores = (Q @ K.transpose(-2, -1)) / math.sqrt(dK)   # Shape: (B, T, T)         # Optional Causal Mask: Block future words with -infinity!        if causalMask:            upperTri = torch.triu(torch.ones(T, T, device=x.device), diagonal=1).bool()            scores = scores.masked_fill(upperTri, float("-inf"))         # Step 3: Turn scores into 100% row probabilities via Softmax        attnWeights = F.softmax(scores, dim=-1)              # Shape: (B, T, T)         # Step 4: Blend the Value (V) vectors using the attention weights!        output = attnWeights @ V                             # Shape: (B, T, dModel)        return output, attnWeights torch.manual_seed(42)attn = SingleHeadSelfAttention(dModel=16)sampleSentence = torch.randn(1, 4, 16)                       # 1 sentence, 4 words, 16 slidersoutVectors, weights = attn(sampleSentence, causalMask=True)print("Output Shape:", outVectors.shape)print("4x4 Causal Attention Matrix:\n", weights[0].detach().numpy().round(2))

Pro Tip (Use F.scaled_dot_product_attention in Production PyTorch 2.x!):

Writing out the 4 steps above is the best way to understand how Self-Attention works. In real production training, PyTorch 2.0+ includes F.scaled_dot_product_attention(Q, K, V, is_causal=True), which automatically calls FlashAttention on your GPU for 3×3\times faster speed and massive memory savings!

Key Points

✓Self-Attention allows every token in a sequence to look directly at every other token in a single step (11-hop connection), eliminating the sequential bottleneck of RNNs.
✓Every token vector x\mathbf{x} is projected into three roles: Query (Q\mathbf{Q}) (what it is looking for), Key (K\mathbf{K}) (its label tag for others to match against), and Value (V\mathbf{V}) (the semantic content it passes along).
✓Scaled Dot-Product Attention computes softmax(QKTdk)V\text{softmax}\left(\frac{\mathbf{Q}\mathbf{K}^T}{\sqrt{d_k}}\right)\mathbf{V}, where dividing by dk\sqrt{d_k} prevents large dot products from pushing Softmax into zero-gradient saturation.
✓Autoregressive Decoder models (like GPT and Llama) apply an upper-triangular Causal Mask of −∞-\infty before Softmax so tokens cannot peek at future tokens during parallel training.
✓Because the T×TT \times T attention matrix grows quadratically (O(T2)\mathcal{O}(T^2)), modern LLMs use FlashAttention to compute exact attention in fast GPU SRAM tiles without materializing the full T×TT \times T matrix in memory.

Common Mistakes

✕ Applying the Causal Mask AFTER Softmax instead of BEFORE Softmax.

If you zero out future tokens after Softmax, the remaining probabilities in each row no longer sum to 1.01.0 (100%100\%)! Always mask future positions with -inf before calling F.softmax(scores, dim=-1).

✕ Forgetting to divide by sqrt(d_k) before Softmax.

Without the 1dk\frac{1}{\sqrt{d_k}} scaling factor, large embedding dimensions produce extreme dot products (like +80+80 and −80-80), turning Softmax into a hard one-hot step function whose backward gradients vanish to zero.

✕ Transposing the batch dimension instead of the last two dimensions in K.transpose(-2, -1).

In batched PyTorch tensors of shape (Batch, Seq, Dim), never write K.T! Always use K.transpose(-2, -1) so only the sequence and feature dimensions swap: (B, T, d) @ (B, d, T) -> (B, T, T).

The Big Picture

Old Sequence Models (RNNs / LSTMs)

Word #1 → Word #2 → ... → Word #50 (Slow Chain · Early Context Fades Away)

Self-Attention (Direct All-to-All Spotlight)

Match Query (Q) to Keys (K) → Softmax % Weights → Blend Values (V) in 1 Parallel Step!

The big takeaway is simple: Self-Attention turns a static dictionary lookup into a dynamic conversation between words. Every word asks a question (Q\mathbf{Q}), checks every other word's nametag (K\mathbf{K}), and absorbs a custom blend of their meanings (V\mathbf{V}).

Remember: What if a single word needs to look at two different things at once—like checking who did the action AND when it happened? One single attention spotlight isn't enough! In our next lesson, we will give the AI 8 to 64 spotlights at the same time with Multi-Head Attention and Transformer Blocks!