Text Preprocessing and Tokenization

How machines translate human language into numbers—from classical text cleaning to subword Byte-Pair Encoding (BPE) that powers GPT-4, Claude, and Llama-3.

27 minBeginnerCode Examples

The Core Thesis: Neural networks cannot read letters, words, or sentences—they only understand matrices of numbers. Tokenization is the critical translation bridge at the very front of every NLP model and Large Language Model that chops raw human text into discrete pieces (Tokens) and maps each piece to a unique integer ID from a fixed Vocabulary.

  1. Beginner Foundations: Why Computers Cannot Read Raw Text

Every neural network layer we have studied so far computes weighted sums and matrix multiplications (y=Wx+b\mathbf{y} = \mathbf{W}\mathbf{x} + \mathbf{b}). You cannot multiply the English word "Apple" by a weight matrix!

Before text ever touches a neural network, it must pass through a two-stage Text-to-Numbers Pipeline:

Stage 1: Text Normalization & Preprocessing

Cleans up messy raw strings—standardizing Unicode characters (like smart quotes and accents), stripping broken HTML tags, and normalizing whitespace so the model sees consistent text.

" AI is AMAZING!! " → "AI is AMAZING!!"

Stage 2: Tokenization & Vocabulary Lookup

Splits the string into atomic units called Tokens and looks up each token's unique integer ID in a fixed dictionary called the Vocabulary (V\mathcal{V}).

["AI", " is", " amazing"] → [15836, 374, 8056]

Classical NLP Preprocessing vs. Modern LLM Preprocessing

In Classical NLP (like spam filters or search engines), engineers aggressively stripped punctuation, forced all text to lowercase, removed common Stop Words ("the", "is", "not"), and chopped word endings via Stemming / Lemmatization (turning "running" and "ran" into "run"). In Modern LLMs (like GPT-4), we DO NOT remove stop words or lowercase text—because capital letters, punctuation, and words like "not" carry crucial meaning and code syntax!

  1. Beginner to Intermediate: The Three Levels of Tokenization Granularity

How should we chop a sentence into tokens? Should every whole word be a token, or every single letter, or something in between? Let us compare the three fundamental approaches:

1. Word-Level
["unhappiness", "is", "bad"]

Splits on spaces. Fatal Flaw: Huge vocabulary (500k+500\text{k}+ words), treats "run" and "running" as unrelated, and crashes on any unseen typo with an Out-Of-Vocabulary (<UNK>) token!

2. Character-Level
["u", "n", "h", "a", "p", ...]

Every single letter is a token. Fatal Flaw: Tiny vocabulary (≈100\approx 100 chars) and zero <UNK> errors, but sequences become 5×5\times longer and individual letters carry no semantic meaning!

3. Subword-Level (Modern AI)
["un", "happi", "ness"]

The Goldilocks Solution! Keeps common words whole ("apple") while splitting rare or complex words into meaningful prefixes, roots, and suffixes!

Tokenization StrategyTypical Vocab Size ∣V∣|\mathcal{V}|Sequence Length TTHandles Typos & New Words?
Word-Level100,000 – 1,000,000+100{,}000\text{ -- }1{,}000{,}000+Short (11 token/word)No — replaces unseen words with <UNK>
Character-Level100 – 500100\text{ -- }500Very Long (5–75\text{--}7 tokens/word)Yes — can spell any word
Subword (BPE / WordPiece)32,000 – 128,00032{,}000\text{ -- }128{,}000Compact (≈1.3\approx 1.3 tokens/word)Yes — decomposes rare words into subwords!
⚡ Knowledge Check

Suppose a user types the newly invented medical term "neuroplasticity" into a chatbot. How does a Subword Tokenizer handle this word if it was never seen as a single whole word before?

A) It splits it into known subword pieces like ["neuro", "plastic", "ity"]▼
✓ Correct!Because the subword tokenizer knows common roots like "neuro", "plastic", and "ity", the LLM immediately understands both the spelling and the combined meaning of the word without throwing an unknown-word error!
B) It replaces the entire word with <UNK> and loses all meaning▼
✕ Incorrect.Replacing unseen words with <UNK> is the fatal flaw of old Word-Level tokenizers. Modern Byte-Level BPE tokenizers never produce <UNK>!

  1. Intermediate: How Byte-Pair Encoding (BPE) Learns Subwords

How does an AI decide which letter combinations should become subword tokens? It uses an elegant data-compression algorithm called Byte-Pair Encoding (BPE) (used by GPT-2, GPT-3.5, GPT-4, and Llama-3)!

BPE starts with a tiny vocabulary of individual characters (or 256 raw UTF-8 bytes) and repeatedly merges the most frequent adjacent pair until the vocabulary reaches the target size ∣V∣|\mathcal{V}|:

Best Merge Pair (ta,tb)=arg⁡max⁡(ti,tj)∈V×VCount(ti,tj)\text{Best Merge Pair } (t_a, t_b) = \arg\max_{(t_i, t_j) \in \mathcal{V} \times \mathcal{V}} \text{Count}(t_i, t_j)

Step-by-Step Worked Example: Training BPE on a Mini Corpus

Suppose our training text contains the words "low" (55 times), "lower" (22 times), and "newest" (66 times). Watch how BPE builds subwords in 3 iterations:

Iteration 1: Merge ('e', 's')

In n e w e s t (6×6\times), the pair ('e', 's') appears 66 times. BPE merges them into the new token "es"!

Iteration 2: Merge ('es', 't')

Now ('es', 't') appears 66 times in n e w es t. BPE merges them into the suffix token "est"!

Iteration 3: Merge ('l', 'o')

Across low (55) and lower (22), the pair ('l', 'o') appears 5+2=75 + 2 = 7 times, merging into "lo", and next into "low"!

Comparing Modern Subword Algorithms: BPE vs. WordPiece vs. Unigram

1. Byte-Level BPE
GPT-4, Llama-3, Claude

Starts from 256256 raw UTF-8 bytes and merges the highest-frequency pairs bottom-up. Can encode any emoji or language with zero <UNK> tokens!

2. WordPiece
BERT, DistilBERT

Marks subword continuations with ## (e.g., ["play", "##ing"]) and merges pairs that maximize corpus likelihood Count(ab)Count(a)Count(b)\frac{\text{Count}(ab)}{\text{Count}(a)\text{Count}(b)}.

3. Unigram (SentencePiece)
T5, Gemma, Multilingual

Works top-down: starts with a huge vocabulary and iteratively prunes subwords that least reduce probabilistic model likelihood, treating spaces as _ characters.

  1. Intermediate: Special Control Tokens, Padding & Attention Masks

Beyond regular words, every tokenizer reserves a set of Special Control Tokens that tell the neural network where documents begin, where a user's chat turn ends, and which positions are empty padding:

Special TokenFull NameExact Role Inside Transformers & LLMs
<BOS> / <|begin_of_text|>Beginning of SequencePrepended to the start of a prompt to signal the start of a new document and anchor attention.
<EOS> / <|eot_id|>End of Sequence / TurnWhen an LLM samples this token ID during generation, the inference loop stops generating immediately!
<PAD>Padding TokenFills shorter sentences in a mini-batch so all rows share the same rectangular matrix length TT.
[CLS] and [SEP]Classification & SeparatorUsed in BERT encoders to pool sentence-level meaning ([CLS]) and separate sentence pairs ([SEP]).

Why Padding Requires a Binary Attention Mask (1 for Real Tokens, 0 for Padding)

Suppose Sentence 1 has 55 tokens and Sentence 2 has 33 tokens. To batch them into a (2,5)(2, 5) tensor, we pad Sentence 2 with two <PAD> tokens: [104, 892, 41, 0, 0]. To prevent Self-Attention from paying attention to fake <PAD> tokens, the tokenizer outputs a matching attention_mask tensor: [1, 1, 1, 0, 0]!

⚡ Knowledge Check

When you ask ChatGPT a question, how does the server know the exact moment the model has finished its answer and should stop generating more words?

A) The model predicts the special <EOS> (End-of-Sequence) token ID▼
✓ Correct!During Supervised Fine-Tuning (SFT), every assistant response ends with a special End-of-Turn / <EOS> token. As soon as the model samples that token ID, the generation loop breaks and returns the response!
B) It always stops as soon as it outputs a period (.)▼
✕ Incorrect.A period (.) is just a normal punctuation token used at the end of every sentence inside a multi-paragraph essay. Stopping on a period would limit the AI to 1-sentence answers!

  1. Advanced: Pre-Tokenization Regex, Fertility, and Why LLMs Struggle to Count Letters

At the Advanced / LLM Architecture level, tokenization explains many of the strangest behaviors, bugs, and cost differences in modern AI systems:

1. Why LLMs Struggle with "How Many R's in Strawberry?"

An LLM never sees individual letters! To the tokenizer, "strawberry" is immediately converted into integer IDs like ["str", "aw", "berry"] = [496, 675, 15717]. Because the model only sees those 3 numbers, it cannot see the individual 'r' characters inside token 15717 unless it spells the word out letter-by-letter first!

Tokenization Hides Character-Level Spelling!

2. Leading Spaces, Pre-Tokenization Regex & Numbers

Before BPE merges run, a Pre-Tokenizer Regex prevents merges across punctuation and attaches leading spaces to words: "apple" at the start of a line is a different token ID than " apple" after a space! Furthermore, Llama-3 splits numbers into 1–31\text{--}3 digit chunks so math is consistent.

"apple" (ID 18040) != " apple" (ID 24149)

  1. Vocabulary Size Trade-offs (32k32\text{k} vs. 128k128\text{k}) and Multilingual Token Fertility

Token Fertility measures the average number of tokens required to encode one word. Llama-2 used a 32,00032{,}000-token vocabulary, whereas GPT-4o (200k200\text{k} vocab) and Llama-3 (128,256128{,}256 vocab) expanded their vocabularies by 4×4\times. Why? A larger vocabulary compresses English, Python code whitespace, and non-Latin languages (like Hindi, Arabic, and Japanese) into fewer tokens—making inference faster and fitting more text into the context window, at the cost of a larger Embedding and Output Softmax matrix (Wvocab∈Rd×∣V∣\mathbf{W}_{\text{vocab}} \in \mathbb{R}^{d \times |\mathcal{V}|})!

⚡ Knowledge Check

Why did Llama-3 increase its BPE vocabulary size from 32,00032{,}000 tokens (in Llama-2) to 128,256128{,}256 tokens, and what is the primary memory trade-off of making ∣V∣|\mathcal{V}| too large?

A) Compresses text & code into ~15% fewer tokens; increases Embedding/LM-Head matrix VRAM▼
✓ Correct!A 128k128\text{k} vocabulary encodes the same paragraph in fewer tokens (speeding up autoregressive generation and multilingual tasks), while increasing the parameter count of the input embedding table and final Softmax projection (∣V∣×dmodel)(|\mathcal{V}| \times d_{\text{model}}).
B) It makes sequences 4x longer so the model thinks harder▼
✕ Incorrect.A larger vocabulary merges more frequent multi-character chunks and whitespace patterns, which shortens sequence length TT rather than lengthening it.

  1. Visual Explanation: The End-to-End Tokenization Pipeline

Look at how a raw user string passes through Unicode normalization, regex pre-tokenization, BPE subword merging, and special-token framing to produce the exact integer tensors fed into a Transformer:

  1. Python Implementation: Training Byte-Pair Encoding (BPE) From Scratch

Here is a complete, runnable Python implementation of the Byte-Pair Encoding (BPE) algorithm from scratch—showing how a tokenizer automatically discovers subword units like "ing" and "train" by iteratively merging the most frequent adjacent pairs:

bpe_tokenizer_from_scratch.pyPython 3.11+ · Subword BPE Engine
from collections import Counter # 1. Define a Training Corpus (Word -> Frequency Count)rawCorpus = {"train": 5, "training": 6, "trained": 3, "learning": 4} # Split each word into characters + end-of-word symbol "</w>"splits = {word: list(word) + ["</w>"] for word in rawCorpus} def getPairCounts(wordSplits, corpusFreqs):    pairCounts = Counter()    for word, freq in corpusFreqs.items():        symbols = wordSplits[word]        for i in range(len(symbols) - 1):            pairCounts[(symbols[i], symbols[i + 1])] += freq    return pairCounts def mergePair(bestPair, wordSplits):    a, b = bestPair    for word, symbols in wordSplits.items():        newSymbols = []        i = 0        while i < len(symbols):            if i < len(symbols) - 1 and symbols[i] == a and symbols[i + 1] == b:                newSymbols.append(a + b)                i += 2            else:                newSymbols.append(symbols[i])                i += 1        wordSplits[word] = newSymbols # 2. Run 6 BPE Merge Iterations to Discover Subwords Automatically!for step in range(1, 7):    counts = getPairCounts(splits, rawCorpus)    bestPair, maxFreq = counts.most_common(1)[0]    mergePair(bestPair, splits)    print("Merge", step, ":", bestPair, "->", bestPair[0] + bestPair[1], "(freq:", maxFreq, ")") print("Final Subword Splits:", splits)

Pro Tip (Production Libraries: tiktoken & HuggingFace Tokenizers):

In production, you never run BPE loops in pure Python—you use OpenAI's Rust-powered tiktoken library (enc = tiktoken.get_encoding("o200k_base")) or HuggingFace's Rust-backed AutoTokenizer.from_pretrained(), which tokenize millions of words per second on CPU!

Key Points

✓Tokenization converts raw text strings into a sequence of discrete integer IDs from a fixed vocabulary V\mathcal{V} so neural networks can process language mathematically.
✓Classical NLP pipelines used lowercasing, stop-word removal, and stemming/lemmatization, whereas modern LLMs preserve exact casing, punctuation, and whitespace via lossless subword tokenization.
✓Subword tokenization (BPE, WordPiece, Unigram) combines the best of word-level and character-level approaches: keeping frequent words whole while splitting rare or unseen words into meaningful subwords.
✓Byte-Pair Encoding (BPE) builds its vocabulary bottom-up by starting with individual characters/bytes and iteratively merging the most frequent adjacent token pairs.
✓Special tokens (<BOS>, <EOS>, <PAD>) control document framing, stop generation, and pad variable-length sequences inside mini-batches alongside a binary attention_mask.
✓Many LLM quirks—such as difficulty counting characters inside a word or sensitivity to leading spaces—are direct consequences of subword tokenization rather than neural reasoning failures.

Common Mistakes

✕ Using Tokenizer A (e.g., BERT) with Model Weights B (e.g., Llama-3).

Every pretrained model is permanently married to the exact vocabulary table it was trained with! In Tokenizer A, ID 2054 might mean "apple", while in Model B, row 2054 of the embedding matrix means "database". Always load the matching tokenizer via AutoTokenizer.from_pretrained(model_id).

✕ Removing stop words ("not", "never", "no") before feeding text into a Transformer or sentiment model.

Stripping stop words turns "This movie is not good" into "movie good", completely inverting the true meaning! Never strip stop words when using contextual embeddings or Transformers.

✕ Padding sequences in a batch without passing the attention_mask into the Transformer.

If you pad shorter sentences with <PAD> tokens but forget to pass attention_mask into model(input_ids, attention_mask=mask), Self-Attention mixes the fake padding tokens into your real word representations!

✕ Right-padding instead of Left-padding when running batched autoregressive LLM generation.

Decoder-only LLMs generate the next token immediately after the very last token on the right. If you pad on the right during batched generation, the LLM tries to continue generating after the <PAD> tokens! Always set padding_side="left" for batched LLM inference.

The Big Picture

Rigid Word Dictionaries (Classical NLP)

Split on Spaces → Huge 500k Vocabulary → Breaks on Typos, Code & New Words (<UNK>)

Byte-Level Subword Tokenization (Modern LLMs)

Raw Text / Code / Emoji → BPE Subword Merges → Compact Integer IDs → Ready for Vector Embeddings!

The important conceptual shift is realizing that a tokenizer is the eyes of a Language Model. Before a Transformer can apply Self-Attention or predict the next word, Byte-Pair Encoding translates messy human text, source code, and emojis into a clean, lossless stream of integer IDs with zero unknown-word failures.

Remember: Tokenization turns words into arbitrary integer IDs (like ID 15836), which have no mathematical meaning on their own—in the very next lesson, Word Embeddings will turn each integer ID into a rich, multi-dimensional vector of meaning!