Sequence Models: RNNs, LSTMs, and GRUs Explained Simply

Learn how AI reads sentences word-by-word with memory—from simple Recurrent Neural Networks (RNNs) and the game-of-telephone problem to LSTMs, GRUs, and PyTorch code.

25 minBeginnerCode Examples

The Core Idea: Standard neural networks have total amnesia—every time you show them a new word, they instantly forget the word that came right before it. But in human language, word order changes everything ("The dog bit the man" is normal; "The man bit the dog" is front-page news!). Sequence Models (RNNs, LSTMs, and GRUs) give AI a sticky-note memory that passes notes from one word to the next as it reads a sentence from left to right.

  1. Beginner: Why Word Order Matters & How an RNN Works (The Sticky-Note Loop)

Imagine trying to read a mystery book through a tiny keyhole where you can only see one word at a time—and the moment a new word appears, your brain completely wipes the previous word from memory! You would have no idea what the story is about.

That is how a normal Multilayer Perceptron (MLP) works: it takes a fixed input, spits out an answer, and immediately forgets everything.

A Recurrent Neural Network (RNN) fixes this with a dead-simple trick: as it reads each word, it keeps a Running Summary Note (called the Hidden State hth_t) and passes that note forward to the next step!

1. New Word (xtx_t)
Current word embedding

The word the AI is looking at right now (for example, Word #3: "not").

2. Old Memory (ht−1h_{t-1})
Summary from step t−1t-1

The sticky-note summary passed over from all the words read so far ("The movie is...").

3. Updated Memory (hth_t)
Blend old note + new word

Combines the old note and the new word into a fresh summary note to pass to Word #4!

ht=tanh⁡(Wxhxt+Whhht−1+b)h_t = \tanh\left( W_{xh} x_t + W_{hh} h_{t-1} + b \right)

Plain-English Translation of That Formula

Do not let the letters scare you! It literally says: New Memory (hth_t) = squash together (New Word xtx_t + Old Memory ht−1h_{t-1}). Even better, the RNN uses the exact same weight matrices (Wxh,WhhW_{xh}, W_{hh}) at every single word step, whether the sentence has 55 words or 500500 words!

  1. Beginner: Why Simple RNNs Forget Long Sentences (The Game of Telephone)

Simple RNNs work great on short 4-word phrases. But what happens when an RNN reads a long 40-word paragraph? Look at this sentence:

"I grew up in France, where I loved eating fresh croissants by the river every morning, so I speak fluent ______."

To guess the final word ("French"), the AI has to remember the word "France" from 1818 words ago! Why does a standard RNN fail at this?

1. The "Game of Telephone" (Memory Overwriting)

Because a simple RNN crams the new word into the exact same sticky note at every step, each new word waters down the old memory by half. After 2020 steps, the word "France" has been completely overwritten by words like "river" and "morning"!

2. Vanishing Gradients Through Time

During training, blame signals have to walk backward across all 2020 word steps (Backpropagation Through Time). Multiplying a small fraction like 0.5×0.5×…0.5 \times 0.5 \times \dots twenty times shrinks the learning signal to 0.0000010.000001 (Vanishing Gradient), so the network never learns long-distance links!

⚡ Knowledge Check

Why does a standard RNN struggle to connect the first word of a 50-word paragraph to the very last word of that paragraph?

A) Its single memory note gets overwritten at every step, and backward gradients shrink to near zero▼
✓ Correct!Repeatedly multiplying and squashing the hidden state across 50 steps dilutes early memories forward and causes gradients to vanish backward.
B) Because an RNN is only allowed to read 5 words before crashing▼
✕ Incorrect.An RNN can loop over sequences of any length; the problem is not a hard length limit, but rather that early information fades out over many steps.

  1. Medium: LSTMs (Long Short-Term Memory) — The Express Conveyor Belt

In 1997, Sepp Hochreiter and Jürgen Schmidhuber asked a brilliant question: "Why force the AI to overwrite its only memory note on every single word? What if we give it a separate Long-Term Vault that glides straight through the sentence untouched unless a valve opens?"

That invention is the LSTM (Long Short-Term Memory)! Instead of one memory note, an LSTM keeps two memories and controls them using three smart valves (Gates) that open (1.01.0) or close (0.00.0) using Sigmoid switches:

The Cell State (CtC_t)Long-Term Conveyor Belt

Think of CtC_t as an express luggage conveyor belt running straight across the top of the network. Important facts (like "grew up in France") ride this belt for 50+50+ words without getting squashed or overwritten!

1. The Forget Gate (ftf_t)The Eraser Valve (0 to 1)

Decides what old info on the conveyor belt is no longer needed. When a sentence ends and a new topic starts, ft→0f_t \to 0 wipes the old subject clean!

2. The Input Gate (iti_t)The "Save This!" Valve

Decides whether the current word is important enough to write onto the long-term conveyor belt (saves "France"; ignores filler words like "the").

3. The Output Gate (oto_t)The Working Memory Valve

Decides what part of the long-term conveyor belt needs to be pulled out right now into the short-term working memory (hth_t) to predict the very next word!

The Secret of the LSTM Conveyor Belt (Addition Instead of Multiplication!)
Ct=ft⊙Ct−1⏟Keep Old Memory+it⊙C~t⏟Add New Fact⟹ht=ot⊙tanh⁡(Ct)C_t = \underbrace{f_t \odot C_{t-1}}_{\text{Keep Old Memory}} + \underbrace{i_t \odot \tilde{C}_t}_{\text{Add New Fact}} \quad \Longrightarrow \quad h_t = o_t \odot \tanh(C_t)

Why That Plus Sign (++) Solved Vanishing Gradients

Look at the formula for CtC_t above: it updates the long-term memory using addition (++) instead of matrix multiplication! When the Forget Gate stays open (ft≈1f_t \approx 1), memories and backward learning signals glide straight down the highway untouched—just like a ResNet skip connection!

  1. Medium: GRUs (Gated Recurrent Units) — The Fast, Streamlined Cousin

In 2014, Kyunghyun Cho and his team asked: "LSTMs work great, but do we really need TWO separate memory states (CtC_t and hth_t) and THREE separate gates? Can we simplify it?"

They created the GRU (Gated Recurrent Unit), which merges the long-term and short-term memory back into a single vector hth_t and uses only two gates:

1. The Update Gate (ztz_t) — Combined Forget + Input!

One single balance knob (1−zt)⊙ht−1+zt⊙h~t(1 - z_t) \odot h_{t-1} + z_t \odot \tilde{h}_t: if ztz_t is 00, keep 100%100\% of the old memory; if ztz_t is 11, replace it with 100%100\% of the new word!

2. The Reset Gate (rtr_t)

Controls how much of the past memory to consult when proposing a new candidate meaning for the current word.

FeatureSimple RNNLSTMGRU
Memory Vectors Passed Forward1 (hth_t)2 (hth_t and Cell State CtC_t)1 (hth_t)
Number of Control Gates0 Gates3 Gates (Forget, Input, Output)2 Gates (Update, Reset)
Remembers Long Sentences?No (forgets after ~10 words)Yes (100+ words)Yes (similar to LSTM)
Speed & Parameter CountFastest (1×1\times weights)Heaviest (4×4\times weights)25% faster than LSTM (3×3\times weights)
⚡ Knowledge Check

What is the main difference between an LSTM and a GRU when you use them in PyTorch?

A) LSTM uses 3 gates and returns (h_n, c_n); GRU uses 2 gates, returns only h_n, and runs 25% faster▼
✓ Correct!GRUs merge the cell state and hidden state into one vector hth_t controlled by 2 gates instead of 3, giving you 25% fewer weights and faster training with nearly identical accuracy!
B) GRUs have zero gates and cannot remember past 3 words▼
✕ Incorrect.That describes a plain vanilla RNN. A GRU has two learnable gates (Update and Reset) specifically designed to remember long sequences.

  1. Advanced: Bidirectional LSTMs & Why Transformers Eventually Replaced Them

At the Advanced / Architecture level, two big questions remain: how do we let an LSTM see the end of a sentence before judging the beginning, and why did ChatGPT switch from LSTMs to Transformers?

1. Bidirectional LSTMs (Reading Both Ways!)

Consider: "Teddy bears are on sale" vs. "Teddy Roosevelt was president." You cannot know if "Teddy" is a toy or a person until you read the next word! A Bidirectional LSTM runs two LSTMs at once—one reading left-to-right (h→t\overrightarrow{h}_t) and one reading right-to-left (h←t\overleftarrow{h}_t)—and glues their two memories together!

nn.LSTM(..., bidirectional=True) → Doubles Output Size (2 * hidden_size)

2. The Fatal Bottleneck of LSTMs: Sequential Waiting!

Why did Transformers replace LSTMs for Large Language Models? Because an LSTM cannot compute Word #50 until it has finished waiting for Word #1, #2, ..., #49 in a slow loop! GPUs have 10,000+10{,}000+ parallel cores that sit idle waiting for the loop. Furthermore, cramming a 1,0001{,}000-word essay into a single fixed-size vector hth_t is like trying to pack an entire library into one small suitcase!

LSTMs Must Read Step-by-Step → Cannot Parallelize Across GPUs!

Where Are LSTMs, GRUs, and Modern Recurrent Models (Mamba / RWKV) Used Today?

Because an LSTM/GRU only needs to keep a tiny, constant-size memory vector hth_t in RAM (O(1)\mathcal{O}(1) memory per step), they are still widely used in low-power IoT sensors, smartwatch heart-rate monitors, real-time audio noise cancellation, and robotics! Even more exciting: a new generation of Linear Recurrent State-Space Models (like Mamba and RWKV) removes the non-linear squashing inside the recurrent loop so they can train in parallel like a Transformer while running as fast and cheap as an RNN!

⚡ Knowledge Check

Even though LSTMs solved vanishing gradients over medium-length sentences, what is the #1 reason AI labs switched to Transformers for training massive models on thousands of GPUs?

A) LSTMs must process words one-by-one in sequence (hth_t waits for ht−1h_{t-1}), preventing parallel GPU training▼
✓ Correct!Because step tt depends on step t−1t-1, an LSTM cannot process all words of a training document simultaneously on a GPU, and it forces all past context through a single fixed-size vector bottleneck.
B) Because LSTMs cannot use Word Embeddings▼
✕ Incorrect.LSTMs use nn.Embedding as their first layer just like Transformers do.

  1. Visual Explanation: Inside the LSTM Memory Conveyor Belt

Watch how an LSTM reads a tricky movie review word-by-word—using its Conveyor Belt (CtC_t) and 3 Control Valves to remember a Plot Twist ("not") that a plain RNN would mishandle:

🛤️ LONG-TERM MEMORY CONVEYOR BELT (CELL STATE C_t)

Sentence: "This movie is NOT bad!"
STEP 1–3Filler Words
"This movie is..."
Input Valve (i_t):0.10 (Ignore filler)
Belt Memory:[Topic: Movie]
STEP 4 · PLOT TWIST!Valve Opens!
"...NOT..."
Input Valve (i_t):0.99 (SAVE THIS!)
Belt Memory:[FLIP NEXT WORD!]
STEP 5 · FINAL WORDCombine!
"...bad!"
Output Valve (o_t):0.95 (Read Belt!)
Final Meaning:NOT + bad = GOOD! ✓

↙ Click Below to Inspect What Each Gate Valve Does Inside Step 4 & Step 5 ↘

VALVE INSPECTOR 1 · FORGET & INPUT GATES

🔓 How the Conveyor Belt Updates (CtC_t)

Inspect

▼

When the LSTM sees a new period . at the end of a sentence:
🗑️ Forget Gate (ftf_t) drops to 0.02Erases old sentence ✓
💾 Input Gate (iti_t) opens to 0.98 on keywordsSaves new subject ✓

VALVE INSPECTOR 2 · RNN VS LSTM COMPARISON

⚡ Why Plain RNN Gets This Review Wrong

Compare

▼

If 10 filler words sit between "not" and "bad":
❌ Plain RNN: Overwrites "not", sees "bad"Predicts: 1-Star ✕
✅ LSTM Belt: Carries "not" to flip "bad"Predicts: 5-Stars ✓

  1. Python Implementation: Building an LSTM Sentiment Classifier in PyTorch

Here is a clean, runnable PyTorch script showing how nn.Embedding and nn.LSTM work together as a team: first turning token IDs into word vectors, then reading the sequence step-by-step and using the final summary memory h_n[-1] to classify a sentence:

lstm_sequence_classifier.pyPython 3.11+ · PyTorch nn.LSTM & nn.GRU
import torchimport torch.nn as nn # 1. Define a Clean LSTM Sequence Classifierclass SentimentLSTM(nn.Module):    def init(self, vocabSize: int, embedDim: int, hiddenDim: int):        super().init()        self.embedding = nn.Embedding(vocabSize, embedDim, padding_idx=0)        # ALWAYS set batch_first=True so input shape is (Batch, SeqLen, EmbedDim)!        self.lstm = nn.LSTM(embedDim, hiddenDim, batch_first=True)        self.classifier = nn.Linear(hiddenDim, 1)     def forward(self, tokenIds: torch.Tensor) -> torch.Tensor:        # Step 1: Turn Token IDs (B, T) into Word Vectors (B, T, embedDim)        x = self.embedding(tokenIds)         # Step 2: Read Sequence Left-to-Right!        # allSteps has every word's output; (h_n, c_n) has the FINAL summary notes!        allSteps, (hFinal, cFinal) = self.lstm(x)         # Step 3: Pass the final summary note hFinal[-1] into the Linear classifier        finalSummary = hFinal[-1]               # Shape: (Batch, hiddenDim)        return self.classifier(finalSummary)    # Shape: (Batch, 1) torch.manual_seed(42)model = SentimentLSTM(vocabSize=1000, embedDim=32, hiddenDim=64) # 2. Test on a Batch of 3 Sentences (Each 6 Tokens Long)sampleBatch = torch.randint(1, 1000, (3, 6))    # Shape: (3, 6)logits = model(sampleBatch)probs = torch.sigmoid(logits)print("Input Shape:", sampleBatch.shape, "-> Output Probabilities:", probs.detach().numpy().round(3).flatten())

Pro Tip (Always Pass batch_first=True to PyTorch RNN, LSTM, and GRU!):

By a strange historical quirk, PyTorch's nn.RNN, nn.LSTM, and nn.GRU expect inputs in shape (SeqLen, Batch, Dim) by default! Always pass batch_first=True when creating the layer so it accepts standard (Batch, SeqLen, Dim) tensors just like every other layer in PyTorch!

Key Points

✓Recurrent Neural Networks (RNNs) process sequential data step-by-step, passing a running summary vector (the Hidden State hth_t) from each word to the next.
✓Simple RNNs suffer from the "game of telephone" (Vanishing Gradients Through Time)—overwriting early words and failing to learn connections across more than ≈10\approx 10 words.
✓LSTMs solve long-term memory by introducing an additive Cell State conveyor belt (CtC_t) controlled by three Sigmoid valves: the Forget Gate (ftf_t), Input Gate (iti_t), and Output Gate (oto_t).
✓GRUs streamline the LSTM architecture by merging the cell state and hidden state into a single vector hth_t with just two gates (Update and Reset), running ≈25%\approx 25\% faster with fewer parameters.
✓While LSTMs and GRUs remain popular for streaming audio, time-series sensors, and edge devices, Large Language Models moved to Transformers because recurrent loops cannot process all tokens in parallel on GPUs.

Common Mistakes

✕ Forgetting batch_first=True in PyTorch's nn.LSTM or nn.GRU.

Without batch_first=True, PyTorch treats axis 00 as sequence steps and axis 11 as batch size, silently scrambling your sentences across the batch! Always set batch_first=True.

✕ Unpacking nn.LSTM output as out, h_n = lstm(x) instead of out, (h_n, c_n) = lstm(x).

Unlike nn.RNN and nn.GRU (which return one state tensor h_n), nn.LSTM returns a tuple of two tensors (h_n, c_n) as its second output!

✕ Using the final step output out[:, -1, :] when sequences are right-padded with zeros.

If a short sentence is padded with ten <PAD> tokens on the right, the final step has just read ten fake padding tokens! Either grab the index of the last real token, use pack_padded_sequence, or average across real tokens using a mask.

The Big Picture

Simple RNN (Single Sticky-Note Memory)

Overwrites Same Note Every Step → Forgets Early Words → Fails on Long Sentences

LSTM & GRU (Gated Memory Conveyor Belt)

Smart Valves Decide What to Forget & Save → Carries Key Facts Across 100+ Steps!

The big takeaway is simple: Sequence Models taught AI how to read across time. By adding learnable Forget and Input valves onto a memory conveyor belt, LSTMs and GRUs solved the short-term amnesia of early neural networks and paved the way for modern language AI.

Remember: Even with an LSTM conveyor belt, forcing an AI to read one word at a time and cram an entire book into a single summary vector is a bottleneck—in our next lesson, Attention Mechanisms and Transformers will smash that bottleneck by letting every word look directly at every other word at the exact same time!