Derivatives and Partial Derivatives

How calculus measures rates of change across millions of parameters—forming the mathematical compass behind Gradient Descent and Backpropagation.

25 minBeginnerCode Examples

The Core Thesis: A machine learning model learns by making mistakes, measuring its error (Loss), and nudging its internal weights to reduce that error. Derivatives and Partial Derivatives are the mathematical compass that tell the model the exact direction and sensitivity of every single weight so it knows whether to nudge each parameter up or down.

  1. What is a Derivative? (The Sensitivity Dial)

Forget memorizing dusty textbook tables for a moment. In Artificial Intelligence, a Derivative answers one simple, practical question: "If I nudge this input knob xx by a tiny amount, how much does the output f(x)f(x) change?"

Geometrically, the derivative of a function f(x)f(x) at any point is the instantaneous slope of the tangent line touching the curve at that exact point:

f′(x)=dfdx=lim⁡h→0f(x+h)−f(x)hf^\prime(x) = \dfrac{df}{dx} = \lim_{h \to 0} \dfrac{f(x + h) - f(x)}{h}
Positive Slope (dfdx>0\dfrac{df}{dx} \gt 0)
Curve goes UPHILL

Increasing xx increases the output f(x)f(x). To reduce error, we must decrease xx (step left).

Negative Slope (dfdx<0\dfrac{df}{dx} \lt 0)
Curve goes DOWNHILL

Increasing xx decreases the output f(x)f(x). To reduce error, we must increase xx (step right).

Zero Slope (dfdx=0\dfrac{df}{dx} = 0)
Flat Valley / Peak

The curve is completely flat. We have reached a minimum, maximum, or saddle point!

Intuitive Numerical Example (f(x)=x2f(x) = x^2):

Suppose our error function is f(x)=x2f(x) = x^2 and our current weight is x=3x = 3, giving an error of f(3)=9f(3) = 9. What happens if we nudge xx up by a tiny step h=0.001h = 0.001 to x=3.001x = 3.001?
f(3.001)−f(3)0.001=9.006001−90.001=6.001≈6\dfrac{f(3.001) - f(3)}{0.001} = \dfrac{9.006001 - 9}{0.001} = 6.001 \approx 6
At x=3x = 3, the output grows 66 times faster than our tiny nudge! That agrees with the calculus rule ddx(x2)=2x=2(3)=6\dfrac{d}{dx}(x^2) = 2x = 2(3) = 6.

  1. Essential Derivative Rules Used in Machine Learning

You only need a small toolkit of basic derivative rules to understand every loss function and activation function in Deep Learning:

Rule NameFunction f(x)f(x)Derivative f′(x)f^\prime(x)Where It Appears in AI
Constant Rulef(x)=cf(x) = c00Fixed labels yy or unlearnable constants have zero gradient.
Linear Rulef(x)=wx+bf(x) = wx + bwwDense neural layers and Linear Regression slopes.
Power Rulef(x)=xnf(x) = x^nnxn−1n x^{n-1}Mean Squared Error (MSE): derivative of e2e^2 is 2e2e.
Exponential Rulef(x)=exf(x) = e^xexe^xPowers Softmax, Sigmoid, and Tanh activations.
Natural Log Rulef(x)=ln⁡(x)f(x) = \ln(x)1x\dfrac{1}{x}Cross-Entropy Loss (L=−ln⁡(y^)\mathcal{L} = -\ln(\hat{y})) in classifiers & LLMs.

Derivatives of Neural Network Activation Functions

1. Sigmoid Activation (σ(z)\sigma(z))

Squashes any real number into (0,1)(0, 1). Its derivative has a remarkably clean self-referential form:

σ(z)=11+e−z\sigma(z) = \dfrac{1}{1 + e^{-z}}
σ′(z)=σ(z)(1−σ(z))\sigma^\prime(z) = \sigma(z)\big(1 - \sigma(z)\big)

Notice: Its maximum slope is only 0.250.25 (at z=0z=0). Multiplying many 0.250.25s across deep layers causes the Vanishing Gradient Problem!

2. ReLU Activation (max⁡(0,z)\max(0, z))

The modern default in Deep Learning. Its derivative is simply a binary switch (11 if positive, 00 if negative):

ReLU(z)=max⁡(0,z)\text{ReLU}(z) = \max(0, z)
ReLU′(z)=1 (for z>0)or0 (for z<0)\text{ReLU}^\prime(z) = 1 \text{ (for } z \gt 0\text{)} \quad \text{or} \quad 0 \text{ (for } z \lt 0\text{)}

Because the slope is a constant 11 for all positive inputs, gradients flow backward through deep networks without shrinking!

⚡ Knowledge Check

You are training a single weight ww. At its current value w=5w = 5, you compute the derivative of the Loss with respect to ww and find dLdw=−8\dfrac{dL}{dw} = -8. To DECREASE the Loss, how should you adjust ww?

A) Increase w (move right, e.g. to 5.1)▼
✓ Correct!A negative derivative (dLdw=−8\dfrac{dL}{dw} = -8) means the error curve slopes downhill as ww increases. Mathematically, wnew=w−α(−8)=w+8αw_{\text{new}} = w - \alpha(-8) = w + 8\alpha, which increases ww to reduce the Loss.
B) Decrease w (move left, e.g. to 4.9)▼
✕ Incorrect.Because the slope is negative (−8-8), moving left (decreasing ww) walks uphill on the error curve and makes the Loss worse.

  1. Partial Derivatives: Isolating One Knob Among Millions

Real AI models never depend on just one variable xx. Even a simple Linear Regression model y^=wx+b\hat{y} = wx + b has two parameters (ww and bb), and modern LLMs have billions. How do we measure the slope when there are multiple variables?

We use a Partial Derivative, written with the curly symbol ∂\partial instead of dd. To take the partial derivative ∂f∂w\dfrac{\partial f}{\partial w}, you pretend every other variable is a frozen constant number and differentiate normally with respect to ww alone!

f(w,b)=w2b+3b2f(w, b) = w^2 b + 3b^2
Partial w.r.t. ww (Freeze bb as constant)
∂f∂w=2wb+0=2wb\dfrac{\partial f}{\partial w} = 2wb + 0 = 2wb
Partial w.r.t. bb (Freeze ww as constant)
∂f∂b=w2(1)+6b=w2+6b\dfrac{\partial f}{\partial b} = w^2(1) + 6b = w^2 + 6b

Geometric Intuition (3D Mountain Surface): Imagine standing on a 3D mountain where East-West is ww, North-South is bb, and altitude is Loss L(w,b)L(w, b).

• ∂L∂w\dfrac{\partial L}{\partial w} is the slope of the mountain if you walk strictly East-West (keeping North-South fixed).
• ∂L∂b\dfrac{\partial L}{\partial b} is the slope of the mountain if you walk strictly North-South (keeping East-West fixed).

  1. The Gradient Vector (∇L\nabla L): Packing All Partial Derivatives Together

Instead of handling thousands of partial derivatives separately, we pack every parameter's partial derivative into a single column vector called the Gradient Vector, denoted by the nabla symbol ∇L\nabla L:

∇L(w)=[∂L∂w1∂L∂w2⋮∂L∂wn]\nabla L(\mathbf{w}) = \begin{bmatrix} \dfrac{\partial L}{\partial w_1} \\[10pt] \dfrac{\partial L}{\partial w_2} \\[6pt] \vdots \\[6pt] \dfrac{\partial L}{\partial w_n} \end{bmatrix}

The Golden Rule of the Gradient Vector

In high-dimensional space, the Gradient Vector ∇L\nabla L ALWAYS points in the direction of STEEPEST ASCENT (maximum error increase). Therefore, −∇L-\nabla L ALWAYS points in the direction of STEEPEST DESCENT (fastest error reduction).

Worked Example: Deriving Linear Regression Gradients From Scratch

Consider a single training sample (x,y)(x, y) with prediction y^=wx+b\hat{y} = wx + b and Squared Error Loss L(w,b)=(wx+b−y)2L(w, b) = (wx + b - y)^2. Let e=(wx+b−y)e = (wx + b - y) be our prediction error:
Partial Derivative w.r.t. Weight ww
∂L∂w=2(wx+b−y)⋅x=2e⋅x\dfrac{\partial L}{\partial w} = 2(wx + b - y) \cdot x = 2e \cdot x
Partial Derivative w.r.t. Bias bb
∂L∂b=2(wx+b−y)⋅1=2e\dfrac{\partial L}{\partial b} = 2(wx + b - y) \cdot 1 = 2e

Look at how intuitive that math is: the weight gradient ∂L∂w=2ex\dfrac{\partial L}{\partial w} = 2ex is simply Error ×\times Input! If input x=0x = 0, that weight did not contribute to the mistake, so its gradient is 00.

⚡ Knowledge Check

Given the two-variable function L(w1,w2)=3w12+4w1w2+5w22L(w_1, w_2) = 3w_1^2 + 4w_1 w_2 + 5w_2^2, what is the partial derivative ∂L∂w1\dfrac{\partial L}{\partial w_1} evaluated at the point (w1=1,w2=2)(w_1 = 1, w_2 = 2)?

A) 14▼
✓ Correct!Treating w2w_2 as a constant, ∂L∂w1=6w1+4w2+0\dfrac{\partial L}{\partial w_1} = 6w_1 + 4w_2 + 0. Plugging in w1=1w_1 = 1 and w2=2w_2 = 2 gives 6(1)+4(2)=6+8=146(1) + 4(2) = 6 + 8 = 14.
B) 34▼
✕ Incorrect.When differentiating with respect to w1w_1, the term 5w225w_2^2 has no w1w_1 in it, so it acts like a constant and its derivative is 00 (not 10w210w_2).

  1. The Chain Rule: The Engine Behind Backpropagation

A neural network is a nested chain of functions: the weights ww compute a linear sum z=wx+bz = wx + b, which feeds into an activation a=σ(z)a = \sigma(z), which feeds into the Loss L(a)L(a). How do we find ∂L∂w\dfrac{\partial L}{\partial w} when ww is buried three layers deep?

We use the Chain Rule of Calculus, which states that the derivative of a composite chain of functions is simply the product of the local derivatives along the path:

∂L∂w=∂L∂a⋅∂a∂z⋅∂z∂w\dfrac{\partial L}{\partial w} = \dfrac{\partial L}{\partial a} \cdot \dfrac{\partial a}{\partial z} \cdot \dfrac{\partial z}{\partial w}
∂L∂a\dfrac{\partial L}{\partial a}
How much the Loss LL changes when the Activation aa changes.
∂a∂z\dfrac{\partial a}{\partial z}
How much the Activation aa changes when the Logit zz changes (σ′(z)\sigma^\prime(z)).
∂z∂w\dfrac{\partial z}{\partial w}
How much the Logit z=wx+bz = wx+b changes when Weight ww changes (=x= x).

What is Backpropagation? Backpropagation is literally just the Chain Rule applied from right to left (from the output Loss back to the early layers), caching intermediate derivatives so the GPU never recalculates the same derivative twice!

  1. Higher-Order Derivatives: Jacobians and Hessians

As you read advanced AI papers, you will encounter two matrix generalizations of partial derivatives:

1. The Jacobian Matrix (J\mathbf{J}) — 1st Order

When a function takes a vector x∈Rn\mathbf{x} \in \mathbb{R}^n and outputs a vector y∈Rm\mathbf{y} \in \mathbb{R}^m (like a Softmax layer), the Jacobian is the (m×n)(m \times n) grid of all first-order partial derivatives:

Ji,j=∂yi∂xjJ_{i,j} = \dfrac{\partial y_i}{\partial x_j}
2. The Hessian Matrix (H\mathbf{H}) — 2nd Order

The square (n×n)(n \times n) matrix of second-order partial derivatives of the Loss. It measures the curvature (bowl vs. saddle shape) of the loss landscape:

Hi,j=∂2L∂wi∂wjH_{i,j} = \dfrac{\partial^2 L}{\partial w_i \partial w_j}
⚡ Knowledge Check

In a neural network computational graph w→z→a→Lw \to z \to a \to L, suppose the local derivatives are ∂L∂a=4\dfrac{\partial L}{\partial a} = 4, ∂a∂z=0.5\dfrac{\partial a}{\partial z} = 0.5, and ∂z∂w=−3\dfrac{\partial z}{\partial w} = -3. What is the overall gradient ∂L∂w\dfrac{\partial L}{\partial w}?

A) -6▼
✓ Correct!By the Chain Rule, we multiply the local derivatives along the path: ∂L∂w=4×0.5×(−3)=−6\dfrac{\partial L}{\partial w} = 4 \times 0.5 \times (-3) = -6.
B) 1.5▼
✕ Incorrect.The Chain Rule multiplies derivatives along a sequential path (4×0.5×−3=−64 \times 0.5 \times -3 = -6); it only adds them when two separate branches merge together.

  1. Visual Explanation: Forward Pass vs. Backward Gradient Flow

Every training step in Deep Learning consists of two directional passes across the exact same computational graph: a Forward Pass to compute the prediction and Loss, followed by a Backward Pass using partial derivatives to send blame back to each parameter:

  1. Python Implementation: Numerical vs. Analytical & Autograd

In Python, we can verify our calculus formulas by comparing Analytical Partial Derivatives (exact formulas) against Finite Difference Numerical Approximation:

∂f∂w≈f(w+h)−f(w−h)2h\dfrac{\partial f}{\partial w} \approx \dfrac{f(w + h) - f(w - h)}{2h}
partial_derivatives_ai.pyPython 3 · NumPy
import numpy as np # 1. Define a Loss Function L(w, b) = (w * x + b - y) ** 2x, y = 2.0, 10.0                            # Sample input and true targetw, b = 3.0, 1.0                             # Current weights (yHat = 7.0) def computeLoss(wVal, bVal):    return (wVal * x + bVal - y) ** 2 # 2. Exact Analytical Partial Derivatives (Chain Rule)error = (w * x + b) - y                     # 7.0 - 10.0 = -3.0gradWExact = 2 * error * x                  # 2 * (-3.0) * 2.0 = -12.0gradBExact = 2 * error * 1.0                # 2 * (-3.0) * 1.0 = -6.0 # 3. Numerical Gradient Check (Centered Finite Difference)eps = 1e-5gradWNum = (computeLoss(w + eps, b) - computeLoss(w - eps, b)) / (2 * eps)gradBNum = (computeLoss(w, b + eps) - computeLoss(w, b - eps)) / (2 * eps) print("Analytical Gradient [dL/dw, dL/db]:", [gradWExact, gradBExact])print("Numerical Check:", [round(gradWNum, 4), round(gradBNum, 4)]) # 4. Take One Gradient Descent Step to Reduce Losslr = 0.05wNew = w - lr * gradWExact                  # 3.0 - 0.05 * (-12) = 3.6bNew = b - lr * gradBExact                  # 1.0 - 0.05 * (-6)  = 1.3print("Old Loss:", computeLoss(w, b), "-> New Loss:", round(computeLoss(wNew, bNew), 4))

Pro Tip (PyTorch Autograd): In PyTorch, you never have to write derivative formulas by hand. Setting requires_grad=True on your weight tensors and calling loss.backward() automatically traverses the computational graph and populates w.grad with the exact analytical partial derivatives!

Key Points

✓A derivative dfdx\dfrac{df}{dx} measures the instantaneous rate of change (slope) of an output with respect to a small nudge in its input.
✓A partial derivative ∂L∂wi\dfrac{\partial L}{\partial w_i} isolates the sensitivity of one specific weight wiw_i by treating all other parameters in the model as frozen constants.
✓The Gradient Vector ∇L\nabla L collects all partial derivatives into a single vector that points in the direction of steepest error increase; stepping in −∇L-\nabla L minimizes error fastest.
✓The Chain Rule (∂L∂w=∂L∂a∂a∂z∂z∂w\dfrac{\partial L}{\partial w} = \dfrac{\partial L}{\partial a} \dfrac{\partial a}{\partial z} \dfrac{\partial z}{\partial w}) multiplies local derivatives across nested layers and is the exact mathematical engine of Backpropagation.
✓Activation functions with tiny derivatives (like Sigmoid, where σ′(z)≤0.25\sigma^\prime(z) \le 0.25) cause vanishing gradients when multiplied across many layers, which is why modern networks prefer ReLU or GELU.
✓The Jacobian matrix stores first-order partial derivatives of vector-to-vector functions, while the Hessian matrix stores second-order derivatives that describe loss surface curvature.

Common Mistakes

✕ Differentiating with respect to the input data x instead of the parameters w.

During model training, your dataset features x\mathbf{x} and labels yy are fixed constants. Always take partial derivatives with respect to the learnable weights w\mathbf{w} and biases b\mathbf{b}.

✕ Adding the gradient instead of subtracting it during weight updates.

Because the gradient ∇L\nabla L points uphill toward maximum error, updating weights via w+α∇L\mathbf{w} + \alpha \nabla L performs Gradient Ascent and causes the model's loss to explode. Always subtract: w−α∇L\mathbf{w} - \alpha \nabla L.

✕ Confusing a zero gradient (dL/dw = 0) with a guaranteed global minimum.

A slope of zero simply means the surface is locally flat. In high-dimensional neural networks, ∇L=0\nabla L = \mathbf{0} can be a local minimum, a maximum peak, or most commonly a saddle point.

✕ Using finite differences (numerical derivatives) to train neural networks.

Computing L(w+ϵ)−L(w−ϵ)2ϵ\dfrac{L(w+\epsilon) - L(w-\epsilon)}{2\epsilon} requires two full forward passes per parameter—which would take billions of forward passes for one step in an LLM. Only use numerical gradients for unit-testing, and use analytical Backpropagation for training.

The Big Picture

Blind Trial-and-Error (No Calculus)

Guess Random Weights → Measure Error → Guess Again Blindly → Never Converges

Gradient-Guided Learning (Partial Derivatives + Chain Rule)

Forward Prediction → Compute Loss → Partial Derivatives via Chain Rule → Step Downhill (-∇L)

The important conceptual shift is that calculus turns training from a blind guessing game into a guided descent. Even when a model has 7070 billion weights, partial derivatives and the Chain Rule tell us the exact uphill or downhill slope for all 7070 billion knobs in a single backward pass.

Remember: Linear Algebra (vectors and matrices) powers the Forward Pass to make predictions, while Calculus (partial derivatives and the Chain Rule) powers the Backward Pass so the model can learn from its mistakes.