The Core Thesis: Neural networks cannot read letters, words, or sentences—they only understand matrices of numbers. Tokenization is the critical translation bridge at the very front of every NLP model and Large Language Model that chops raw human text into discrete pieces (Tokens) and maps each piece to a unique integer ID from a fixed Vocabulary.
- Beginner Foundations: Why Computers Cannot Read Raw Text
Every neural network layer we have studied so far computes weighted sums and matrix multiplications (). You cannot multiply the English word "Apple" by a weight matrix!
Before text ever touches a neural network, it must pass through a two-stage Text-to-Numbers Pipeline:
Cleans up messy raw strings—standardizing Unicode characters (like smart quotes and accents), stripping broken HTML tags, and normalizing whitespace so the model sees consistent text.
" AI is AMAZING!! " → "AI is AMAZING!!"
Splits the string into atomic units called Tokens and looks up each token's unique integer ID in a fixed dictionary called the Vocabulary ().
["AI", " is", " amazing"] → [15836, 374, 8056]
Classical NLP Preprocessing vs. Modern LLM Preprocessing
In Classical NLP (like spam filters or search engines), engineers aggressively stripped punctuation, forced all text to lowercase, removed common Stop Words ("the", "is", "not"), and chopped word endings via Stemming / Lemmatization (turning "running" and "ran" into "run"). In Modern LLMs (like GPT-4), we DO NOT remove stop words or lowercase text—because capital letters, punctuation, and words like "not" carry crucial meaning and code syntax!
- Beginner to Intermediate: The Three Levels of Tokenization Granularity
How should we chop a sentence into tokens? Should every whole word be a token, or every single letter, or something in between? Let us compare the three fundamental approaches:
Splits on spaces. Fatal Flaw: Huge vocabulary ( words), treats "run" and "running" as unrelated, and crashes on any unseen typo with an Out-Of-Vocabulary (<UNK>) token!
Every single letter is a token. Fatal Flaw: Tiny vocabulary ( chars) and zero <UNK> errors, but sequences become longer and individual letters carry no semantic meaning!
The Goldilocks Solution! Keeps common words whole ("apple") while splitting rare or complex words into meaningful prefixes, roots, and suffixes!
| Tokenization Strategy | Typical Vocab Size | Sequence Length | Handles Typos & New Words? |
|---|---|---|---|
| Word-Level | Short ( token/word) | No — replaces unseen words with <UNK> | |
| Character-Level | Very Long ( tokens/word) | Yes — can spell any word | |
| Subword (BPE / WordPiece) | Compact ( tokens/word) | Yes — decomposes rare words into subwords! |
Suppose a user types the newly invented medical term "neuroplasticity" into a chatbot. How does a Subword Tokenizer handle this word if it was never seen as a single whole word before?
A) It splits it into known subword pieces like ["neuro", "plastic", "ity"]▼
B) It replaces the entire word with <UNK> and loses all meaning▼
<UNK> is the fatal flaw of old Word-Level tokenizers. Modern Byte-Level BPE tokenizers never produce <UNK>!
- Intermediate: How Byte-Pair Encoding (BPE) Learns Subwords
How does an AI decide which letter combinations should become subword tokens? It uses an elegant data-compression algorithm called Byte-Pair Encoding (BPE) (used by GPT-2, GPT-3.5, GPT-4, and Llama-3)!
BPE starts with a tiny vocabulary of individual characters (or 256 raw UTF-8 bytes) and repeatedly merges the most frequent adjacent pair until the vocabulary reaches the target size :
Step-by-Step Worked Example: Training BPE on a Mini Corpus
Suppose our training text contains the words "low" ( times), "lower" ( times), and "newest" ( times). Watch how BPE builds subwords in 3 iterations:
In n e w e s t (), the pair ('e', 's') appears times. BPE merges them into the new token "es"!
Now ('es', 't') appears times in n e w es t. BPE merges them into the suffix token "est"!
Across low () and lower (), the pair ('l', 'o') appears times, merging into "lo", and next into "low"!
Comparing Modern Subword Algorithms: BPE vs. WordPiece vs. Unigram
Starts from raw UTF-8 bytes and merges the highest-frequency pairs bottom-up. Can encode any emoji or language with zero <UNK> tokens!
Marks subword continuations with ## (e.g., ["play", "##ing"]) and merges pairs that maximize corpus likelihood .
Works top-down: starts with a huge vocabulary and iteratively prunes subwords that least reduce probabilistic model likelihood, treating spaces as _ characters.
- Intermediate: Special Control Tokens, Padding & Attention Masks
Beyond regular words, every tokenizer reserves a set of Special Control Tokens that tell the neural network where documents begin, where a user's chat turn ends, and which positions are empty padding:
| Special Token | Full Name | Exact Role Inside Transformers & LLMs |
|---|---|---|
| <BOS> / <|begin_of_text|> | Beginning of Sequence | Prepended to the start of a prompt to signal the start of a new document and anchor attention. |
| <EOS> / <|eot_id|> | End of Sequence / Turn | When an LLM samples this token ID during generation, the inference loop stops generating immediately! |
| <PAD> | Padding Token | Fills shorter sentences in a mini-batch so all rows share the same rectangular matrix length . |
| [CLS] and [SEP] | Classification & Separator | Used in BERT encoders to pool sentence-level meaning ([CLS]) and separate sentence pairs ([SEP]). |
Why Padding Requires a Binary Attention Mask (1 for Real Tokens, 0 for Padding)
Suppose Sentence 1 has tokens and Sentence 2 has tokens. To batch them into a tensor, we pad Sentence 2 with two <PAD> tokens: [104, 892, 41, 0, 0]. To prevent Self-Attention from paying attention to fake <PAD> tokens, the tokenizer outputs a matching attention_mask tensor: [1, 1, 1, 0, 0]!
When you ask ChatGPT a question, how does the server know the exact moment the model has finished its answer and should stop generating more words?
A) The model predicts the special <EOS> (End-of-Sequence) token ID▼
<EOS> token. As soon as the model samples that token ID, the generation loop breaks and returns the response!B) It always stops as soon as it outputs a period (.)▼
.) is just a normal punctuation token used at the end of every sentence inside a multi-paragraph essay. Stopping on a period would limit the AI to 1-sentence answers!
- Advanced: Pre-Tokenization Regex, Fertility, and Why LLMs Struggle to Count Letters
At the Advanced / LLM Architecture level, tokenization explains many of the strangest behaviors, bugs, and cost differences in modern AI systems:
An LLM never sees individual letters! To the tokenizer, "strawberry" is immediately converted into integer IDs like ["str", "aw", "berry"] = [496, 675, 15717]. Because the model only sees those 3 numbers, it cannot see the individual 'r' characters inside token 15717 unless it spells the word out letter-by-letter first!
Tokenization Hides Character-Level Spelling!
Before BPE merges run, a Pre-Tokenizer Regex prevents merges across punctuation and attaches leading spaces to words: "apple" at the start of a line is a different token ID than " apple" after a space! Furthermore, Llama-3 splits numbers into digit chunks so math is consistent.
"apple" (ID 18040) != " apple" (ID 24149)
- Vocabulary Size Trade-offs ( vs. ) and Multilingual Token Fertility
Token Fertility measures the average number of tokens required to encode one word. Llama-2 used a -token vocabulary, whereas GPT-4o ( vocab) and Llama-3 ( vocab) expanded their vocabularies by . Why? A larger vocabulary compresses English, Python code whitespace, and non-Latin languages (like Hindi, Arabic, and Japanese) into fewer tokens—making inference faster and fitting more text into the context window, at the cost of a larger Embedding and Output Softmax matrix ()!
Why did Llama-3 increase its BPE vocabulary size from tokens (in Llama-2) to tokens, and what is the primary memory trade-off of making too large?
A) Compresses text & code into ~15% fewer tokens; increases Embedding/LM-Head matrix VRAM▼
B) It makes sequences 4x longer so the model thinks harder▼
- Visual Explanation: The End-to-End Tokenization Pipeline
Look at how a raw user string passes through Unicode normalization, regex pre-tokenization, BPE subword merging, and special-token framing to produce the exact integer tensors fed into a Transformer:
- Python Implementation: Training Byte-Pair Encoding (BPE) From Scratch
Here is a complete, runnable Python implementation of the Byte-Pair Encoding (BPE) algorithm from scratch—showing how a tokenizer automatically discovers subword units like "ing" and "train" by iteratively merging the most frequent adjacent pairs:
from collections import Counter # 1. Define a Training Corpus (Word -> Frequency Count)rawCorpus = {"train": 5, "training": 6, "trained": 3, "learning": 4} # Split each word into characters + end-of-word symbol "</w>"splits = {word: list(word) + ["</w>"] for word in rawCorpus} def getPairCounts(wordSplits, corpusFreqs): pairCounts = Counter() for word, freq in corpusFreqs.items(): symbols = wordSplits[word] for i in range(len(symbols) - 1): pairCounts[(symbols[i], symbols[i + 1])] += freq return pairCounts def mergePair(bestPair, wordSplits): a, b = bestPair for word, symbols in wordSplits.items(): newSymbols = [] i = 0 while i < len(symbols): if i < len(symbols) - 1 and symbols[i] == a and symbols[i + 1] == b: newSymbols.append(a + b) i += 2 else: newSymbols.append(symbols[i]) i += 1 wordSplits[word] = newSymbols # 2. Run 6 BPE Merge Iterations to Discover Subwords Automatically!for step in range(1, 7): counts = getPairCounts(splits, rawCorpus) bestPair, maxFreq = counts.most_common(1)[0] mergePair(bestPair, splits) print("Merge", step, ":", bestPair, "->", bestPair[0] + bestPair[1], "(freq:", maxFreq, ")") print("Final Subword Splits:", splits)Pro Tip (Production Libraries: tiktoken & HuggingFace Tokenizers):
In production, you never run BPE loops in pure Python—you use OpenAI's Rust-powered tiktoken library (enc = tiktoken.get_encoding("o200k_base")) or HuggingFace's Rust-backed AutoTokenizer.from_pretrained(), which tokenize millions of words per second on CPU!
Key Points
<BOS>, <EOS>, <PAD>) control document framing, stop generation, and pad variable-length sequences inside mini-batches alongside a binary attention_mask.Common Mistakes
✕ Using Tokenizer A (e.g., BERT) with Model Weights B (e.g., Llama-3).
Every pretrained model is permanently married to the exact vocabulary table it was trained with! In Tokenizer A, ID 2054 might mean "apple", while in Model B, row 2054 of the embedding matrix means "database". Always load the matching tokenizer via AutoTokenizer.from_pretrained(model_id).
✕ Removing stop words ("not", "never", "no") before feeding text into a Transformer or sentiment model.
Stripping stop words turns "This movie is not good" into "movie good", completely inverting the true meaning! Never strip stop words when using contextual embeddings or Transformers.
✕ Padding sequences in a batch without passing the attention_mask into the Transformer.
If you pad shorter sentences with <PAD> tokens but forget to pass attention_mask into model(input_ids, attention_mask=mask), Self-Attention mixes the fake padding tokens into your real word representations!
✕ Right-padding instead of Left-padding when running batched autoregressive LLM generation.
Decoder-only LLMs generate the next token immediately after the very last token on the right. If you pad on the right during batched generation, the LLM tries to continue generating after the <PAD> tokens! Always set padding_side="left" for batched LLM inference.
The Big Picture
Rigid Word Dictionaries (Classical NLP)
Split on Spaces → Huge 500k Vocabulary → Breaks on Typos, Code & New Words (<UNK>)
Byte-Level Subword Tokenization (Modern LLMs)
Raw Text / Code / Emoji → BPE Subword Merges → Compact Integer IDs → Ready for Vector Embeddings!
The important conceptual shift is realizing that a tokenizer is the eyes of a Language Model. Before a Transformer can apply Self-Attention or predict the next word, Byte-Pair Encoding translates messy human text, source code, and emojis into a clean, lossless stream of integer IDs with zero unknown-word failures.
Remember: Tokenization turns words into arbitrary integer IDs (like ID 15836), which have no mathematical meaning on their own—in the very next lesson, Word Embeddings will turn each integer ID into a rich, multi-dimensional vector of meaning!