Tokenization and Subword Vocabularies in Transformers

How modern Transformers chop human language into LEGO bricks—from Byte-Pair Encoding (BPE) and WordPiece to Byte-Level Tokenization in GPT-4 and Llama-3.

24 minBeginnerCode Examples

The Core Idea: Before a Transformer can run Self-Attention or predict the next word, it needs a way to chop messy human sentences into neat building blocks. If you use whole words, your dictionary is too huge; if you use single letters, sentences are too long. Subword Tokenizers (like BPE and WordPiece) find the sweet spot by breaking text into reusable LEGO bricks—keeping common words whole while slicing rare words, code, and emojis into familiar pieces.

  1. Beginner: The Goldilocks Dilemma (Words vs. Letters vs. LEGO Bricks)

Imagine you are building a toy city. You have three choices for what building blocks to use:

1. Whole-Word Blocks
Pre-built entire houses

Every single word has its own block. If someone invents a new word like "ChatGPT-ish" or makes a typo like "teh", you have no block for it! The AI crashes with an "Unknown Word" error.

2. Single-Letter Blocks
Tiny grains of sand

Every letter is a separate block ('c', 'a', 't'). You can spell anything, but a short paragraph turns into thousands of tiny letters, eating up the Transformer's memory!

3. Subwords (The LEGO Bricks)
The Goldilocks Solution

Common words stay whole ("apple"), while rare or long words get broken into meaningful chunks ("un" + "break" + "able"). Compact, flexible, and never crashes!

Why Transformers Care So Much About Sequence Length

In a Transformer, the cost of Self-Attention grows with the square of the sequence length (T2T^2). If an article takes 500500 subword tokens, Attention computes 500×500=250,000500 \times 500 = 250{,}000 comparisons. If you tokenized by individual characters (3,0003{,}000 letters), Attention would have to compute 3000×3000=9,000,0003000 \times 3000 = 9{,}000{,}000 comparisons—36 times slower! Subwords keep sequences short and manageable.

  1. Beginner: How Byte-Pair Encoding (BPE) Learns Subwords

How does a tokenizer decide which chunks of letters deserve to become official tokens in its dictionary?

Most modern models (like GPT-2, GPT-4, and Llama-3) use an algorithm called Byte-Pair Encoding (BPE). Think of it as a game of "Spot the Pattern":

Step 1 · Start with Characters
Every letter is a separate token

[l, o, w], [l, o, w, e, r], [n, e, w, e, s, t]

Step 2 · Count Neighbor Pairs
Find the most frequent pair

"e" followed by "s" appears everywhere!

Step 3 · Glue Them Together!
Merge into a new token

New Token added to dictionary: "es"

How BPE Builds Complex Words in Practice

  1. First merge: ('e', 's') becomes "es".
  1. Next merge: ('es', 't') becomes "est" (the common superlative ending!).
  1. Repeat this 50,00050{,}000 times on billions of web pages, and the tokenizer naturally discovers prefixes like "un-", stems like "play", suffixes like "-ing", and entire common words like "the" and "banana"!
⚡ Knowledge Check

Why does a BPE subword tokenizer never crash with an "Unknown Word" error when a user types a brand new word like "neurohacker"?

A) Because it can break the word into familiar pieces like ["neuro", "hack", "er"] or individual characters▼
✓ Correct!Even if the whole word was never seen during training, its subword roots and base letters exist in the vocabulary table, so it can always be reconstructed.
B) Because it downloads the definition from Google in real time▼
✕ Incorrect.Tokenizers run locally with a fixed offline vocabulary; they handle new words strictly through subword decomposition.

  1. Medium: Byte-Level BPE (How GPT-4 and Llama-3 Handle Any Language & Emoji)

Early tokenizers had a problem: what happens with emojis (🚀, 🤖) or non-English alphabets (Chinese, Hindi, Arabic)? If your alphabet only includes English letters, where do emojis go?

In 2019, OpenAI introduced Byte-Level BPE (BBPE) in GPT-2, which is now the industry standard used in GPT-4 and Llama-3:

Start from the 256 Raw UTF-8 Bytes

Instead of starting with text letters (A-Z), the tokenizer starts with the 256 possible bytes (0 to 255) that make up computer files. Because every character, letter, accent, code symbol, and emoji on Earth is just a combination of UTF-8 bytes, the model can literally read ANY digital text without failing!

Merge Bytes into Multi-Byte Chunks

Just like standard BPE, frequently occurring byte sequences merge together. Common English words become single tokens, while emojis like 🤖 (which is 4 UTF-8 bytes: 0xF0 0x9F 0xA4 0x96) become a single token ID in the dictionary once merged!

Model FamilyTokenizer AlgorithmVocab Size (∣V∣|\mathcal{V}|)Notable Design Feature
BERTWordPiece30,522Uses ## for subword continuations (e.g. play, ##ing)
GPT-2 / GPT-3.5Byte-Level BPE50,257Zero unknown tokens; treats spaces as part of the token (e.g. Ġapple)
Llama-3Byte-Level BPE (tiktoken)128,256Massive vocab for 15% better compression on code and multilingual text
GPT-4oByte-Level BPE (o200k_base)200,000Huge vocabulary to slash token counts across Asian, Indian, and Arabic languages

  1. Medium: Leading Spaces & Special Tokens (Why " hello" ≠ "hello")

There are two critical details that trip up almost every beginner working with Transformer tokenizers:

1. Spaces Are Attached to the Word!

In modern tokenizers, spaces are not thrown away. Instead, a leading space is glued to the front of the word (represented as Ġ or ).

"hello" (at start of text) = ID #15339
" hello" (with space before it) = ID #24741

These are two completely different token IDs in the model's embedding table!

2. Special Control Tokens

Tokenizers reserve special markers that guide the Transformer:

  • <|begin_of_text|>: Signals start of a prompt
  • <|eot_id|> / <EOS>: Tells the model to stop generating!
  • <PAD>: Fills blank spots in rectangular mini-batches
⚡ Knowledge Check

Why did models like Llama-3 and GPT-4o expand their vocabulary size to over 128,000–200,000 tokens instead of staying at GPT-2's 50,000 tokens?

A) Larger vocabularies compress text and code into fewer total tokens, making processing faster and cheaper▼
✓ Correct!With more pre-merged tokens, the same paragraph in Python, Hindi, or Japanese takes ~20–40% fewer tokens, speeding up generation and fitting more content into the context window.
B) Because smaller vocabularies run out of RAM during inference▼
✕ Incorrect.A larger vocabulary actually uses more embedding parameter RAM (∣V∣×d|\mathcal{V}| \times d), but saves computation time by shortening sequence lengths.

  1. Advanced: Why LLMs Struggle to Spell & The Concept of "Token Fertility"

Understanding tokenization gives you the answer to two of the most famous quirks in modern AI:

1. Why ChatGPT Struggles With "How Many R's in Strawberry?"

An LLM does not see letters! To its tokenizer, "strawberry" is swallowed whole as two chunks: ["straw", "berry"] (Token IDs #496 and #675). The model sees only two numbers—it cannot peek inside the token to count individual letters unless it spells the word out letter-by-letter first!

2. Token Fertility (The Multilingual Tax)

Token Fertility measures how many tokens it takes to express one word. In early tokenizers, English had a fertility of ≈1.3\approx 1.3 tokens/word, while Hindi or Arabic had a fertility of 3–53\text{--}5 tokens/word because the vocabulary lacked non-Latin merges—making AI up to 3×3\times more expensive for non-English speakers until modern 128k+ vocabularies balanced it out!

  1. Visual Explanation: The Subword Slicing Machine

Look at how different types of input text pass through the tokenizer's slicing blades. Notice how common words stay intact, while compound words and code get sliced into recognizable subwords:

RAW INPUT STRING

Length: 42 Characters

"The unbreakable code ran successfully! 🚀"

▼ Tokenizer Applies Learned Merge Rules ▼

Resulting Token Stream (Hover or Tap Each Brick to See Its Token ID):

"The" ID: 464

" un" ID: 734

"break" ID: 5932

"able" ID: 497

" code" ID: 2438

" ran" ID: 5542

" success" ID: 2261

"fully" ID: 3598

"!" ID: 0

" 🚀" ID: 94812

Final Output Sent to Transformer Embedding Layer:

[464, 734, 5932, 497, 2438, 5542, 2261, 3598, 0, 94812]

🔍 Notice What Happened in the Example Above? Click to Inspect!

Inspect

▼

• "The", "code", and "ran" are high-frequency words, so they remain single whole tokens.
• "unbreakable" was split into 3 parts: " un" + "break" + "able".
• "successfully" was split into 2 parts: " success" + "fully".
• "🚀" didn't crash—Byte-Level BPE mapped its 4 UTF-8 bytes cleanly to a single token ID!

  1. Python Implementation: Exploring Real Tokenizers with tiktoken & HuggingFace

Here is a clean, runnable Python script showing how to inspect token IDs, spaces, and subwords using modern production libraries:

transformer_tokenizers_demo.pyPython 3.10+ · tiktoken & transformers
# 1. Simple demonstration of Byte-Pair Encoding logic in Pythonfrom collections import Counter # Start with space-separated character tokenswordCounts = {"l o w </w>": 5, "l o w e r </w>": 2, "n e w e s t </w>": 6} def findBestPair(vocab):    pairs = Counter()    for word, freq in vocab.items():        symbols = word.split()        for i in range(len(symbols) - 1):            pairs[(symbols[i], symbols[i + 1])] += freq    return pairs.most_common(1)[0][0] best = findBestPair(wordCounts)print("Top adjacent pair to merge next:", best)  # ('e', 's') # 2. In Production: Use OpenAI's tiktoken or HuggingFace Tokenizers# (Runs in compiled Rust at millions of tokens per second!)## import tiktoken# enc = tiktoken.get_encoding("cl100k_base")  # GPT-4 tokenizer# tokens = enc.encode("Hello, world! 🚀")# print("Token IDs:", tokens)# print("Decoded back:", [enc.decode([t]) for t in tokens])

Pro Tip (Tokenizer Training is Separate from Model Training!):

A tokenizer is trained once, beforehand on a massive text corpus using pure statistics (counting pairs), and then its vocabulary table is permanently frozen. The Transformer model then learns its weights using that fixed vocabulary!

Key Points

✓Subword tokenization strikes the ideal balance between vocabulary size and sequence length, avoiding out-of-vocabulary errors without ballooning sequence lengths.
✓Byte-Pair Encoding (BPE) starts with base characters and iteratively merges the most frequent adjacent symbol pairs into new subword tokens.
✓Byte-Level BPE (BBPE) starts from the 256 fundamental UTF-8 byte values, allowing modern models (GPT-4, Llama-3) to process emojis, code, and any language without unknown tokens.
✓Spaces are preserved by attaching them directly to subwords, meaning words with and without a leading space have different token IDs.
✓Special tokens control sequence boundaries (BOS, EOS), padding (PAD), and conversation role switches in chat-tuned models.

Common Mistakes

✕ Stripping whitespace or lowercasing text before passing it to a modern Transformer.

Modern subword tokenizers are trained on exact casing and indentation. Forcing text to lowercase ruins code indentation, proper nouns, and changes the token IDs completely.

✕ Assuming an LLM sees individual characters.

An LLM receives token IDs, not character strings. Asking an LLM to reverse words or count letters often fails because it cannot see inside the multi-character token bricks.

✕ Swapping tokenizers between different models.

Token ID #412 might mean "apple" in GPT-4 and "the" in Llama-3. A model's embedding weights are permanently bound to its specific tokenizer.

The Big Picture

Raw Text Stream

Letters, Punctuation, Whitespace, Code & Emojis

Subword Tokenizer (BPE / WordPiece)

Slices into Reusable Subword Bricks → Maps to Fixed Integer IDs → Feeds Embedding Layer

Tokenization is the doorway into the Transformer. Once words are converted into an efficient sequence of token IDs, the model is ready to map those IDs into vectors and start the core magic of Transformers: Self-Attention!