Activation Functions

Why neural networks need non-linear switches—from Sigmoid and Tanh to ReLU, Leaky ReLU, Softmax, and modern LLM activations like GELU and SwiGLU.

26 minBeginnerCode Examples

The Core Thesis: Without activation functions, even a 100-layer deep neural network is mathematically identical to a single-layer linear regression model. Activation functions are the non-linear switches placed after every matrix multiplication that bend, fold, and gate signals—allowing neural networks to learn complex curves, images, and human language.

  1. Beginner Foundations: What is an Activation Function and Why Do We Need It?

Inside every artificial neuron, the first step is a linear weighted sum: z=w⋅x+bz = \mathbf{w} \cdot \mathbf{x} + b. That raw number zz (called the pre-activation or logit) can be anything from −∞-\infty to +∞+\infty.

Before passing that signal to the next layer, the neuron feeds zz through an Activation Function σ(z)\sigma(z) to produce the final neuron output aa:

z=w⋅x+b⟹a=σ(z)z = \mathbf{w} \cdot \mathbf{x} + b \quad \Longrightarrow \quad a = \sigma(z)
Without Activation Functions (Linear Only)

Multiplying two linear matrices together just produces another linear matrix: W2(W1x)=Wcombinedx\mathbf{W}_2(\mathbf{W}_1\mathbf{x}) = \mathbf{W}_{\text{combined}}\mathbf{x}. No matter how many layers you stack, the network can only draw flat straight lines!

100 Linear Layers = 1 Single Linear Layer

With Non-Linear Activation Functions

Inserting a non-linear function σ(⋅)\sigma(\cdot) between layers breaks linearity: W2 σ(W1x)\mathbf{W}_2 \, \sigma(\mathbf{W}_1\mathbf{x}). Each layer can now bend and fold the coordinate grid to fit any curve or boundary!

Universal Function Approximator ✓

Why Not Use a Simple Step Function (00 or 11)?

Early 1950s Perceptrons used a hard step switch (11 if z≥0z \ge 0, else 00). Because a flat step has a derivative (slope) of 00 everywhere, Gradient Descent has no idea which way is downhill! Modern networks require smooth, differentiable activation functions so backpropagation can compute gradients.

  1. Beginner: Classical S-Curve Activations (Sigmoid & Tanh)

The first generation of deep neural networks used smooth, S-shaped squashing functions that compress any input z∈(−∞,+∞)z \in (-\infty, +\infty) into a fixed bounded range:

1. Sigmoid (Logistic Function)Range: (0, 1)

Squashes large negative numbers to 00, zero to 0.50.5, and large positive numbers to 11. Ideal for turning a final output logit into a binary probability!

σ(z)=11+e−z\sigma(z) = \dfrac{1}{1 + e^{-z}}
Derivative: σ′(z)=σ(z)(1−σ(z))≤0.25\sigma^\prime(z) = \sigma(z)\big(1 - \sigma(z)\big) \le 0.25
2. Tanh (Hyperbolic Tangent)Range: (-1, +1)

A rescaled Sigmoid centered at 00 (tanh⁡(0)=0\tanh(0) = 0). Because its outputs average around zero (zero-centered), it converges faster than Sigmoid in recurrent layers and LSTMs.

tanh⁡(z)=ez−e−zez+e−z\tanh(z) = \dfrac{e^z - e^{-z}}{e^z + e^{-z}}
Derivative: tanh⁡′(z)=1−tanh⁡2(z)≤1.0\tanh^\prime(z) = 1 - \tanh^2(z) \le 1.0

Why Sigmoid and Tanh Fail in Deep Hidden Layers: Saturation!

When ∣z∣|z| is large (for example, z=10z = 10 or z=−10z = -10), both Sigmoid and Tanh flatten out completely (saturate), making their derivative σ′(z)≈0\sigma^\prime(z) \approx 0. Even at its steepest center (z=0z = 0), Sigmoid's maximum slope is only 0.250.25. Multiplying 0.25×0.25×…0.25 \times 0.25 \times \dots across 1010 layers via the Chain Rule crushes the gradient to zero (Vanishing Gradients), freezing early layers!

⚡ Knowledge Check

Why is Sigmoid still widely used on the final output neuron of a Binary Classification model (like Spam vs. Not Spam), even though we avoid using it inside deep hidden layers?

A) Its output range (0, 1) directly represents a valid probability P(y = 1 | x)▼
✓ Correct!On the final output layer, Sigmoid maps any real logit zz cleanly into a probability between 00 and 11, and when paired with Binary Cross-Entropy loss, the log cancels the exponential so the output gradient y^−y\hat{y} - y never vanishes!
B) Because Sigmoid can output negative numbers▼
✕ Incorrect.Sigmoid outputs strictly positive numbers in (0,1)(0, 1); it is Tanh that outputs numbers in (−1,+1)(-1, +1).

  1. Intermediate: The ReLU Revolution & Fixing "Dying ReLUs"

In 2010–2012 (AlexNet), researchers solved the vanishing gradient problem and unlocked modern Deep Learning with an astonishingly simple function called ReLU (Rectified Linear Unit): if the input is positive, pass it through unchanged; if it is negative, output 00:

ReLU(z)=max⁡(0,z)⟹ReLU′(z)=1 (for z>0)or0 (for z<0)\text{ReLU}(z) = \max(0, z) \quad \Longrightarrow \quad \text{ReLU}^\prime(z) = 1 \text{ (for } z \gt 0\text{)} \quad \text{or} \quad 0 \text{ (for } z \lt 0\text{)}
1. No Positive Saturation

For any z>0z \gt 0, the derivative is a constant 1.01.0. Gradients flow backward through 100+100+ layers without shrinking (1×1×1=11 \times 1 \times 1 = 1)!

2. Blazing Fast on GPUs

Sigmoid and Tanh require expensive eze^z exponential hardware instructions. ReLU is just a simple if z > 0 comparison!

3. True Sparsity

Negative neurons output an exact 0.00.0, allowing only a sparse, specialized subset of neurons to activate for any given input.

The "Dying ReLU" Problem and Leaky ReLU / ELU

ReLU has one catch: for all negative inputs (z<0z \lt 0), its output is 00 and its slope is exactly 00. If a large gradient nudge pushes a neuron's bias strongly negative so that z<0z \lt 0 for every training sample, that neuron outputs 00 forever, receives 00 gradient forever, and permanently dies!

Leaky ReLU & PReLU (Non-Zero Negative Slope)

Instead of flat zero for z<0z \lt 0, Leaky ReLU gives negative inputs a tiny trickle slope α\alpha (typically α=0.01\alpha = 0.01) so dead neurons can recover! In PReLU, α\alpha is a learnable parameter.

LeakyReLU(z)=max⁡(αz,z)\text{LeakyReLU}(z) = \max(\alpha z, z)
ELU (Exponential Linear Unit)

Uses a smooth exponential curve α(ez−1)\alpha(e^z - 1) for negative values, pushing average activations closer to zero while avoiding a sharp corner at z=0z=0.

ELU(z)=z (for z>0)orα(ez−1) (for z≤0)\text{ELU}(z) = z \text{ (for } z \gt 0\text{)} \quad \text{or} \quad \alpha(e^z - 1) \text{ (for } z \le 0\text{)}
⚡ Knowledge Check

You set your learning rate too high (lr=1.0\text{lr} = 1.0) while training a deep ReLU network. After a few steps, 60%60\% of your hidden neurons output 0.00.0 for every single input in the dataset and never update again. What happened, and what is the easiest activation fix?

A) Dying ReLU problem; lower the learning rate or switch to Leaky ReLU / GELU▼
✓ Correct!A huge gradient step pushed the weights into the z<0z \lt 0 region where ReLU′(z)=0\text{ReLU}^\prime(z) = 0. Switching to Leaky ReLU (α=0.01\alpha = 0.01) or GELU ensures a non-zero gradient so neurons can climb back out.
B) Exploding gradients; switch all hidden layers to Sigmoid▼
✕ Incorrect.Constant 0.00.0 outputs with zero updates is the signature of dead ReLU neurons, and replacing ReLU with Sigmoid in deep networks reintroduces severe vanishing gradients.

  1. Advanced: Modern Transformer & LLM Activations (GELU, SiLU/Swish, and SwiGLU)

If you inspect the source code of BERT, GPT-4, Claude, Gemini, or Llama-3, you will notice they do not use standard ReLU! Why? Because ReLU has a sharp, non-differentiable corner at z=0z = 0 and completely zeroes out slightly negative values. Modern Large Language Models use smooth, probabilistic gated activations:

1. SiLU / Swish (Vision & Diffusion)

Discovered by Google Brain: multiplies zz by its own Sigmoid gate! Smooth everywhere and allows small negative values (around z≈−1.2z \approx -1.2) to pass through.

SiLU(z)=z⋅σ(z)=z1+e−z\text{SiLU}(z) = z \cdot \sigma(z) = \dfrac{z}{1 + e^{-z}}
2. GELU (GPT & BERT Standard)

Gaussian Error Linear Unit: weights input zz by the Gaussian Cumulative Distribution Φ(z)\Phi(z). Smoothly curves through zero with no sharp corner!

GELU(z)=z⋅Φ(z)≈z⋅σ(1.702z)\text{GELU}(z) = z \cdot \Phi(z) \approx z \cdot \sigma(1.702z)
3. SwiGLU (Llama-3, PaLM & DeepSeek)

Swish-Gated Linear Unit: splits the layer into two parallel projections (xWgate\mathbf{x}\mathbf{W}_{\text{gate}} and xWup\mathbf{x}\mathbf{W}_{\text{up}}) and multiplies them element-wise!

SwiGLU(x)=SiLU(xWg)⊙(xWu)\text{SwiGLU}(\mathbf{x}) = \text{SiLU}(\mathbf{x}\mathbf{W}_g) \odot (\mathbf{x}\mathbf{W}_u)

Why SwiGLU Dominates Modern LLMs: In standard ReLU or GELU, a neuron only gates based on its own single linear sum. In SwiGLU, the model learns a dedicated Gate Matrix Wg\mathbf{W}_g that decides how much of the Content Matrix Wu\mathbf{W}_u to let through via element-wise multiplication (⊙\odot), consistently achieving lower perplexity per FLOP in Transformer pretraining!

  1. Master Comparison: Choosing Hidden vs. Output Activations

In architecture design, you always make two separate choices: which activation to put inside Hidden Layers (to extract features) versus which activation to put on the Final Output Layer (to match your target format):

ActivationOutput RangeVanishing Gradient?Where to Use It in Practice
ReLU[0,+∞)[0, +\infty)No (for z>0z \gt 0)Default Hidden Layer for MLPs and CNNs.
GELU / SwiGLU(−0.17,+∞)(-0.17, +\infty)No (Smooth Gate)Default Hidden Layer for Transformers, ViTs, and LLMs.
Leaky ReLU(−∞,+∞)(-\infty, +\infty)No (Never Dies)Hidden Layers in GAN discriminators and deep nets prone to dead neurons.
Sigmoid(0,1)(0, 1)Yes (Severe)Output Layer ONLY for Binary Classification & Multi-Label Classification, or LSTM gates.
Softmax(0,1)(0, 1), ∑pk=1\sum p_k = 1No (with Cross-Entropy)Output Layer ONLY for Multi-Class Classification & LLM next-token probabilities, plus Attention weights.
Linear (Identity)(−∞,+∞)(-\infty, +\infty)N/AOutput Layer ONLY for Continuous Regression (predicting price, temperature).
⚡ Knowledge Check

You are building an image tagger where a single photo can contain BOTH a "Dog" AND a "Car" at the same time (Multi-Label Classification, where labels are not mutually exclusive). Should your final output layer use Softmax or Sigmoid?

A) Independent Sigmoid on each output neuron (so each tag gets its own 0-to-1 probability)▼
✓ Correct!Because Softmax forces all class probabilities to sum to 1.01.0, a high score for "Dog" (0.900.90) would force "Car" down (≤0.10\le 0.10). Applying Sigmoid independently to each output allows both "Dog" (0.950.95) and "Car" (0.920.92) to be true simultaneously!
B) Softmax across all output neurons▼
✕ Incorrect.Softmax is strictly for mutually exclusive classes (where only 1 class can be true at a time, like picking the next token in an LLM).

  1. Visual Explanation: Decision Tree for Picking the Right Activation

Follow this architectural flowchart whenever you design a new neural network layer:

  1. Python Implementation: Comparing Activations & Building SwiGLU in PyTorch

Here is a complete, runnable PyTorch implementation that compares gradient magnitudes across Sigmoid, ReLU, and GELU, and implements a production SwiGLU Feed-Forward Layer just like the one used inside Llama-3:

activation_functions_ai.pyPython 3.11+ · PyTorch 2.x
import torchimport torch.nn as nnimport torch.nn.functional as F # 1. Compare Forward Outputs & Backward Gradients at z = [-4.0, 0.0, 4.0]for name, actFn in [("Sigmoid", torch.sigmoid), ("ReLU", F.relu), ("GELU", F.gelu)]:    z = torch.tensor([-4.0, 0.0, 4.0], requires_grad=True)    out = actFn(z)    out.sum().backward()    print(name, "| Output:", out.detach().numpy().round(3), "| Grad:", z.grad.numpy().round(3)) # 2. Production Llama-3 Style SwiGLU MLP Blockclass SwiGLUMLP(nn.Module):    def init(self, dModel: int, dHidden: int):        super().init()        self.wGate = nn.Linear(dModel, dHidden, bias=False)        self.wUp   = nn.Linear(dModel, dHidden, bias=False)        self.wDown = nn.Linear(dHidden, dModel, bias=False)     def forward(self, x: torch.Tensor) -> torch.Tensor:        # SwiGLU(x) = (SiLU(x @ W_gate) * (x @ W_up)) @ W_down        gated = F.silu(self.wGate(x)) * self.wUp(x)        return self.wDown(gated) mlp = SwiGLUMLP(dModel=64, dHidden=172)sampleTokens = torch.randn(2, 10, 64)       # (Batch=2, SeqLen=10, Dim=64)outputTokens = mlp(sampleTokens)print("SwiGLU Output Shape:", outputTokens.shape)  # torch.Size([2, 10, 64])

Pro Tip (Look at the Gradient Output Above!): At z=4.0z = 4.0, Sigmoid's gradient drops to just 0.018 (nearly flat/saturated!), whereas ReLU's gradient is 1.000 and GELU's gradient is 1.000. At z=−0.5z = -0.5, ReLU's gradient is strictly 0.0, whereas GELU provides a smooth negative gradient so neurons never get permanently stuck!

Key Points

✓Activation functions introduce essential non-linearity between matrix multiplications; without them, any deep neural network collapses into a single linear matrix.
✓Sigmoid squashes values into (0,1)(0, 1) and Tanh squashes values into (−1,+1)(-1, +1), but both saturate at large ∣z∣|z| values, causing vanishing gradients in deep hidden layers.
✓ReLU (max⁡(0,z)\max(0, z)) solved vanishing gradients for positive inputs by maintaining a constant slope of 1.01.0, becoming the default choice for MLPs and CNNs.
✓Leaky ReLU (max⁡(αz,z)\max(\alpha z, z)) fixes the "Dying ReLU" problem by providing a small non-zero slope (e.g., α=0.01\alpha = 0.01) for negative inputs.
✓Modern Transformers and LLMs use smooth gated activations—specifically GELU (zΦ(z)z\Phi(z)), SiLU/Swish (zσ(z)z\sigma(z)), and SwiGLU—for smoother optimization landscapes and stronger representation capacity.
✓Always match your final output layer activation to your loss function: Linear for Regression, Sigmoid for Binary/Multi-Label Classification, and Softmax for mutually exclusive Multi-Class Classification.

Common Mistakes

✕ Using Sigmoid in the hidden layers of a deep neural network.

Because Sigmoid outputs are not zero-centered and its maximum derivative is 0.250.25, stacking Sigmoid hidden layers causes gradients to vanish rapidly. Use ReLU, Leaky ReLU, or GELU in hidden layers instead.

✕ Putting ReLU on the final output layer when predicting unrestricted continuous numbers.

ReLU clips all negative outputs to 0.00.0. If your target (such as temperature changes, financial returns, or normalized ZZ-scores) can be negative, a final ReLU makes negative predictions mathematically impossible!

✕ Using Softmax for Multi-Label Classification where multiple classes can be true at once.

Softmax forces outputs to compete so their sum equals 1.01.0. When multiple labels can co-occur in the same sample, use independent Sigmoid activations with nn.BCEWithLogitsLoss().

✕ Using an overly high learning rate with standard ReLU without gradient clipping.

A single huge gradient update can slam weights into the negative region for all training data, permanently killing a large fraction of your network's ReLU neurons.

The Big Picture

Matrix Multiplication Alone (Flat Geometry)

Input Space → Rotate & Scale Only → Layers Collapse into 1 Matrix → Only Straight Lines

Matrix Multiplication + Non-Linear Activations (Deep Geometry)

Linear Projection → Non-Linear Fold (ReLU / GELU / SwiGLU) → Repeat Across Depth → Models Any Function

The important conceptual shift is seeing how Linear Algebra and Activation Functions partner together inside every neural layer: the weight matrix W\mathbf{W} rotates and stretches the coordinate space, and the activation function σ(⋅)\sigma(\cdot) folds that space along a crease so the next layer can separate patterns that were previously tangled together.

Remember: Use ReLU / Leaky ReLU for fast hidden layers in MLPs and CNNs, GELU / SwiGLU for Transformers and LLMs, and save Sigmoid and Softmax for converting final output logits into probabilities.