The Core Thesis: Backpropagation is only the compass—it tells your neural network which way is downhill by computing the gradient . The Optimizer is the engine that actually moves the weights, and the Learning-Rate Schedule is the throttle that controls how big your steps are at every stage of the journey so your model reaches the lowest valley quickly without crashing.
- Beginner Foundations: What Does an Optimizer Do? (Batch vs. Mini-Batch SGD)
Once Backpropagation finishes computing the gradient , we need a rule to update our weights . The simplest rule is Gradient Descent, which steps in the opposite direction of the gradient scaled by a step size (the Learning Rate)[cite: 10]:
How many training examples should we look at before taking one step? There are three ways, and modern Deep Learning overwhelmingly uses the third:
Computes the exact average gradient over all samples before taking step. Way too slow and cannot fit in GPU memory!
Updates weights after every single data point[cite: 10]. Fast, but the gradient jumps around wildly and wastes GPU parallel cores.
The industry standard! Computes gradients on a small batch of samples—fast on GPUs, smooth enough to steer downhill, and noisy enough to bounce out of bad traps.
Why Plain Mini-Batch SGD Struggles on Deep Neural Networks
Plain SGD has two major flaws: (1) inside a narrow, steep canyon (Ravine), it violently zig-zags back and forth across the walls while barely moving forward along the valley floor, and (2) it uses the exact same learning rate for every weight, even when some weights receive huge gradients and rare word embeddings receive tiny gradients[cite: 10]!
- Beginner to Intermediate: Momentum (The Heavy Bowling Ball)
How do we stop SGD from zig-zagging across narrow canyons and getting stuck in flat plateaus? Imagine replacing a lightweight ping-pong ball with a heavy bowling ball rolling downhill!
SGD with Momentum keeps a moving average of past gradients called Velocity (). Instead of stepping strictly along the current step's noisy gradient , the ball blends (typically , or ) of its previous velocity with the new gradient:
When gradients flip-flop between and across steep canyon walls, the moving average cancels those opposite signs out to , smoothing the trajectory!
Along directions where gradients consistently point the same way, velocity builds up speed and rolls right through flat saddle points and tiny bumps.
Suppose a weight's gradient alternates between and on every step as it bounces between two steep walls, while another weight's gradient is a steady pointing down the valley floor. What does Momentum () do to these two weights?
A) It dampens the +5/-5 oscillation and accelerates the steady +1 direction▼
B) It amplifies the +5/-5 oscillation to +50/-50▼
- Intermediate: Adaptive Optimizers (RMSProp, Adam, and AdamW)
Momentum fixes the direction of the step, but what about the size of the step for each individual parameter? In a Transformer, an embedding weight for a rare word might only get a non-zero gradient once every batches, whereas a final layer weight gets huge gradients on every batch.
Adaptive Optimizers give every single weight its own custom learning rate by dividing the step by the square root of past squared gradients ():
Tracks an exponentially decaying average of squared gradients () and divides the learning rate by . Weights with huge gradients get smaller steps; weights with tiny gradients get larger steps!
Adaptive Moment Estimation (2014): combines 1st Moment (Momentum) with 2nd Moment (RMSProp), plus a bias correction for early steps ():
Why Modern AI Uses AdamW Instead of Standard Adam (Decoupled Weight Decay )
In standard Adam, regularization (weight decay) was added directly into the gradient , which meant Adam's denominator accidentally scaled down the regularization penalty for weights with large gradients! AdamW (Loshchilov & Hutter, 2019) fixes this by decoupling weight decay—shrinking the weights directly via outside the adaptive gradient step. Today, AdamW is the universal default for training Transformers and LLMs.
| Optimizer | Tracks Direction ()? | Adapts Step Size ()? | Extra Memory per Weight | Best Use Case in AI |
|---|---|---|---|---|
| Plain SGD | No | No | states | Simple convex models & textbook demos |
| SGD + Momentum | Yes () | No | ( bytes/param) | ResNets & classic CNN image classification |
| RMSProp | No | Yes () | ( bytes/param) | Recurrent nets & some Reinforcement Learning (DQN) |
| AdamW | Yes () | Yes () | ( bytes/param) | Default for Transformers, LLMs, ViTs, and Diffusion! |
- Intermediate to Advanced: Learning-Rate Schedules (Warmup & Cosine Decay)
Should you keep the learning rate fixed at the exact same number from step to step ? Never! Early in training, you want a larger learning rate to cross the loss landscape quickly, but near the end of training, a large learning rate will bounce around the rim of the valley without ever settling into the minimum[cite: 10].
A Learning-Rate Schedule dynamically adjusts after every step or epoch. Modern LLM training combines two stages:
At step , weights are random and AdamW's variance estimate has seen almost no data—so a full learning rate causes violent gradient spikes! Linear Warmup starts near and ramps up linearly to over steps:
After reaching , Cosine Annealing smoothly glides the learning rate down to along a gentle cosine curve, allowing the weights to settle into a wide, flat minimum:
Why do Transformer models almost always require a Linear Warmup phase during the first few hundred or thousand steps when trained with AdamW?
A) Early gradients are noisy and AdamW's second-moment v_t needs time to calibrate▼
B) Because GPUs need a few minutes to warm up their physical temperature▼
- Advanced Production Optimization: Optimizer VRAM Math, 8-Bit Adam & Linear Scaling
At the Advanced / LLM Systems level, the optimizer is often the largest consumer of GPU memory in your entire cluster! Look at why:
For every single weight parameter , AdamW must store two full 32-bit state tensors in GPU VRAM: the 1st moment ( bytes) and the 2nd moment ( bytes).
7B Model = 56 GB of VRAM just for FP32 AdamW Optimizer States!
To fit large models into memory, engineers quantize and to 8-bit integers (bitsandbytes 8-bit AdamW, saving optimizer RAM) or shard states across GPUs (DeepSpeed ZeRO / FSDP).
Modern LLM Schedule: WSD (Warmup - Stable - Decay)
- The Batch-Size Scaling Rule (When You Double Batch Size, Adjust Learning Rate!)
When you multiply your mini-batch size by (for example, scaling from to across GPUs), each step averages more samples and you take fewer steps per epoch. To converge at the same speed, use the Linear Scaling Rule () for SGD, or Square-Root Scaling () for AdamW!
A neural network has ( Billion) parameters. If AdamW stores its momentum () and variance () states in 32-bit floating point (float32 = bytes per number), how much GPU memory is required strictly for the two optimizer state tensors?
A) 8 GB (4 GB for m_t + 4 GB for v_t)▼
B) 0 GB, because AdamW does not store historical states▼
- Visual Explanation: The Modern AI Optimization & Schedule Loop
Look at how gradients from Backpropagation are clipped, fed into AdamW to update 1st and 2nd moments with decoupled weight decay, and stepped using a Warmup + Cosine Annealing Schedule:
- Python Implementation: AdamW + Linear Warmup & Cosine Decay in PyTorch
Here is a complete, runnable PyTorch implementation showing how to configure torch.optim.AdamW (separating weight-decayed matrices from non-decayed biases/LayerNorms) paired with a Linear Warmup + Cosine Annealing Schedule:
import torchimport torch.nn as nnfrom torch.optim.lr_scheduler import LinearLR, CosineAnnealingLR, SequentialLR # 1. Build a Sample MLP Modeltorch.manual_seed(42)model = nn.Sequential(nn.Linear(16, 64), nn.GELU(), nn.Linear(64, 2)) # 2. Pro Practice: Apply Weight Decay ONLY to 2D Weight Matrices (NOT 1D Biases!)decayParams = [p for p in model.parameters() if p.dim() >= 2]noDecayParams = [p for p in model.parameters() if p.dim() < 2] optimizer = torch.optim.AdamW( [ {"params": decayParams, "weight_decay": 0.01}, {"params": noDecayParams, "weight_decay": 0.0} ], lr=1e-3, betas=(0.9, 0.999)) # 3. Build Warmup (5 steps) -> Cosine Decay (15 steps) SchedulewarmupSched = LinearLR(optimizer, start_factor=0.2, end_factor=1.0, total_iters=5)cosineSched = CosineAnnealingLR(optimizer, T_max=15, eta_min=1e-5)scheduler = SequentialLR(optimizer, schedulers=[warmupSched, cosineSched], milestones=[5]) # 4. Run 20 Optimization Steps and Track Learning RateX = torch.randn(8, 16)Y = torch.randint(0, 2, (8,))criterion = nn.CrossEntropyLoss() for step in range(1, 21): optimizer.zero_grad() loss = criterion(model(X), Y) loss.backward() torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0) optimizer.step() scheduler.step() if step in (1, 5, 12, 20): currentLr = optimizer.param_groups[0]["lr"] print("Step:", step, "| LR:", round(currentLr, 6), "| Loss:", round(loss.item(), 4))Pro Tip (Why We Exclude 1D Biases and LayerNorm from Weight Decay):
Look at Step 2 in the code above (p.dim() >= 2)! In every production LLM codebase (like nanoGPT, Llama, and HuggingFace), weight decay is applied only to 2D weight matrices and set to for 1D biases and LayerNorm scales. Decaying biases or LayerNorm scales hurts model accuracy without preventing overfitting!
Key Points
Common Mistakes
✕ Calling scheduler.step() BEFORE optimizer.step() in PyTorch.
In PyTorch, calling scheduler.step() before optimizer.step() skips the first value of your learning-rate schedule and triggers a user warning. Always run optimizer.step() first, then scheduler.step()!
✕ Using the same learning rate (e.g., lr = 0.1) when switching from SGD to AdamW.
Because AdamW normalizes gradients by , its effective step size is much larger than SGD. While SGD often uses , AdamW typically requires a smaller learning rate around .
✕ Applying weight decay to 1D bias vectors and LayerNorm / BatchNorm parameters.
Shrinking biases and normalization scales toward zero restricts the network's ability to shift and scale activations. Separate parameters into decay (2D+ matrices) and no_decay (1D biases/norms) groups.
✕ Stepping a batch-level warmup scheduler only once per epoch instead of once per batch.
If your warmup is configured for steps, you must call scheduler.step() inside the inner mini-batch loop (not the outer epoch loop) so the learning rate updates on every step!
The Big Picture
Plain Fixed-Step SGD (Bumpy & Unstable)
Same Step Size for Every Weight → Zig-Zags in Canyons → Overshoots Minimum at the End
AdamW + Warmup & Cosine Decay (Modern AI Standard)
Smooth Momentum Direction + Per-Weight Adaptive Step + Scheduled Throttle → Fast, Stable Convergence
The important conceptual shift is separating computing the gradient (Backpropagation) from using the gradient (Optimization)[cite: 10]. By smoothing gradient directions with Momentum, scaling each weight's step size with AdamW, and shaping the global step size with Warmup and Cosine Annealing, modern deep networks can train billions of parameters smoothly without diverging.
Remember: When starting almost any modern Deep Learning or Transformer project, your golden baseline recipe is AdamW(lr=3e-4, weight_decay=0.01) paired with Gradient Clipping (1.0) and a Linear Warmup + Cosine Decay schedule.