Probability, Distributions, and Bayes' Rule

How AI quantifies uncertainty, models data with probability distributions, updates beliefs via Bayes' Theorem, and trains neural networks using Maximum Likelihood Estimation.

26 minBeginnerCode Examples

The Core Thesis: Real-world data is noisy, incomplete, and ambiguous—so an AI model never outputs absolute certainties. Instead, every modern classifier and Large Language Model outputs a Probability Distribution over possible answers, updates its beliefs using Bayes' Rule, and learns parameters by maximizing the likelihood of the training data.

  1. Why AI Speaks in Probabilities (Random Variables)

When ChatGPT generates the next word or a medical AI scans an X-ray, it does not simply output a hard-coded label. Instead, it models a Random Variable XX and assigns a probability score P(X=x)P(X = x) between 00 and 11 to every possible outcome:

0≤P(X=x)≤1and∑xP(X=x)=10 \le P(X = x) \le 1 \quad \text{and} \quad \sum_{x} P(X = x) = 1

To read any machine learning paper or loss function, you must distinguish between the three fundamental types of probability:

1. Marginal Probability
P(A)P(A)

The unconditional probability of a single event occurring on its own (e.g., the baseline percentage of all emails that are spam, P(Spam)=0.20P(\text{Spam}) = 0.20).

2. Joint Probability
P(A,B)P(A, B) or P(A∩B)P(A \cap B)

The probability that both event AA and event BB happen together (e.g., an email is spam AND contains the word "Lottery").

3. Conditional Probability
P(A∣B)P(A \mid B)

The probability of AA happening given that we already know BB occurred:

P(A∣B)=P(A,B)P(B)P(A \mid B) = \dfrac{P(A, B)}{P(B)}

AI Connection: Every supervised classification model and LLM is literally a Conditional Probability Engine! Given input features x\mathbf{x}, the model computes P(y∣x)P(y \mid \mathbf{x})—the conditional probability of label yy given input x\mathbf{x}.

  1. Expectation and Variance (Summarizing Distributions)

Before looking at specific probability distributions, we need two numbers that summarize any distribution's center and spread:

1. Expected Value / Mean (μ=E[X]\mu = \mathbb{E}[X])

The probability-weighted average of all possible values. It represents the center of gravity of the distribution:

E[X]=∑xx⋅P(X=x)\mathbb{E}[X] = \sum_{x} x \cdot P(X = x)
2. Variance (σ2=Var(X)\sigma^2 = \text{Var}(X))

Measures how far values spread out from the mean μ\mu. Standard deviation σ=Var(X)\sigma = \sqrt{\text{Var}(X)} puts this spread back in the original units:

Var(X)=E[(X−μ)2]=E[X2]−(E[X])2\text{Var}(X) = \mathbb{E}\big[(X - \mu)^2\big] = \mathbb{E}[X^2] - (\mathbb{E}[X])^2
⚡ Knowledge Check

When an autoregressive Large Language Model (like GPT-4) is given the prompt "The cat sat on the" and predicts the next token "mat", what mathematical expression is its output Softmax layer computing?

A) Conditional Probability: P(next_token | previous_tokens)▼
✓ Correct!An LLM computes the conditional probability distribution P(wt∣w1,w2,…,wt−1)P(w_t \mid w_1, w_2, \dots, w_{t-1}) over its entire vocabulary, conditioned on the preceding tokens in the context window.
B) Marginal Probability: P(next_token)▼
✕ Incorrect.Marginal probability P(wt)P(w_t) ignores the prompt completely and would just output the most common word in the English language ("the") every single time.

  1. The Four Essential Probability Distributions of AI

A Probability Distribution is a mathematical law that describes how total probability (1.01.0) is distributed across all possible outcomes. Discrete variables use a Probability Mass Function (PMF), while continuous variables use a Probability Density Function (PDF). Four distributions power almost all of Machine Learning:

1. Bernoulli DistributionDiscrete (Binary)

Models a single binary coin flip y∈{0,1}y \in \{0, 1\} with success probability pp. Powers Logistic Regression and Sigmoid Binary Classification!

P(y∣p)=py(1−p)1−yP(y \mid p) = p^y (1 - p)^{1 - y}
2. Categorical (Multinoulli)Discrete (KK Classes)

Models rolling a KK-sided die with class probabilities p=[p1,…,pK]\mathbf{p} = [p_1, \dots, p_K] where ∑pk=1\sum p_k = 1. Powers Softmax Multi-Class Classification and LLM Token Sampling!

P(y=k)=Softmax(z)k=ezk∑j=1KezjP(y = k) = \text{Softmax}(\mathbf{z})_k = \dfrac{e^{z_k}}{\sum_{j=1}^{K} e^{z_j}}
3. Gaussian (Normal) DistributionContinuous Bell Curve

Defined entirely by its mean μ\mu and variance σ2\sigma^2. Powers Linear Regression noise, Weight Initialization, LayerNorm, and Diffusion Models!

N(x∣μ,σ2)=12πσ2exp⁡(−(x−μ)22σ2)\mathcal{N}(x \mid \mu, \sigma^2) = \dfrac{1}{\sqrt{2\pi\sigma^2}} \exp\left(-\dfrac{(x - \mu)^2}{2\sigma^2}\right)
4. Uniform DistributionContinuous Flat

Every value in the interval [a,b][a, b] is equally likely. Used for Dropout masks, Random Hyperparameter Search, and Exploration in RL.

U(x∣a,b)=1b−afor x∈[a,b]\mathcal{U}(x \mid a, b) = \dfrac{1}{b - a} \quad \text{for } x \in [a, b]

  1. Bayes' Rule: How AI Updates Beliefs From Evidence

Often we know how likely a symptom or word is given a class—P(Evidence∣Hypothesis)P(\text{Evidence} \mid \text{Hypothesis})—but what we actually want to predict is the reverse: given the observed evidence, how likely is the hypothesis? Bayes' Theorem flips conditional probabilities around:

P(H∣E)=P(E∣H)⋅P(H)P(E)P(H \mid E) = \dfrac{P(E \mid H) \cdot P(H)}{P(E)}
Posterior: P(H∣E)P(H \mid E)
Updated belief in hypothesis HH after seeing evidence EE.
Likelihood: P(E∣H)P(E \mid H)
Probability of observing evidence EE if HH were true.
Prior: P(H)P(H)
Initial baseline belief in HH before seeing any evidence.
Evidence: P(E)P(E)
Total probability of EE across all possible hypotheses.

Step-by-Step Worked Example (The Base Rate Medical / Spam Trap):

Suppose 20%20\% of all emails are Spam (P(Spam)=0.20P(\text{Spam}) = 0.20) and 80%80\% are Normal (P(Ham)=0.80P(\text{Ham}) = 0.80). The word "Free" appears in 50%50\% of Spam emails (P(Free∣Spam)=0.50P(\text{Free} \mid \text{Spam}) = 0.50), and in 5%5\% of Normal emails (P(Free∣Ham)=0.05P(\text{Free} \mid \text{Ham}) = 0.05). If a new email contains the word "Free", what is the probability it is Spam?

P(Free)=P(Free∣Spam)P(Spam)+P(Free∣Ham)P(Ham)=(0.50)(0.20)+(0.05)(0.80)=0.10+0.04=0.14P(\text{Free}) = P(\text{Free} \mid \text{Spam})P(\text{Spam}) + P(\text{Free} \mid \text{Ham})P(\text{Ham}) = (0.50)(0.20) + (0.05)(0.80) = 0.10 + 0.04 = 0.14
P(Spam∣Free)=0.100.14≈0.714 (or 71.4%)P(\text{Spam} \mid \text{Free}) = \dfrac{0.10}{0.14} \approx 0.714 \text{ (or } 71.4\%\text{)}

Seeing the word "Free" jumped our belief that the email is Spam from a 20%20\% Prior up to a 71.4%71.4\% Posterior!

⚡ Knowledge Check

A rare condition affects 1%1\% of people (P(D)=0.01P(D) = 0.01). An AI screening test has a 90%90\% true positive rate (P(+∣D)=0.90P(+ \mid D) = 0.90) and a 9%9\% false positive rate on healthy people (P(+∣Healthy)=0.09P(+ \mid \text{Healthy}) = 0.09). If a random person tests positive, what is the actual Posterior probability P(D∣+)P(D \mid +)?

A) Only ~9.2% (because the Prior is so rare!)▼
✓ Correct!By Bayes' Rule: P(D∣+)=(0.90)(0.01)(0.90)(0.01)+(0.09)(0.99)=0.0090.009+0.0891≈9.17%P(D \mid +) = \dfrac{(0.90)(0.01)}{(0.90)(0.01) + (0.09)(0.99)} = \dfrac{0.009}{0.009 + 0.0891} \approx 9.17\%. Ignoring the 1%1\% Prior is the famous Base Rate Fallacy.
B) 90%▼
✕ Incorrect.90%90\% is the Likelihood P(+∣D)P(+ \mid D), not the Posterior P(D∣+)P(D \mid +). Confusing P(E∣H)P(E \mid H) with P(H∣E)P(H \mid E) ignores the fact that 99%99\% of the population is healthy.

  1. Maximum Likelihood Estimation (MLE): Why We Use MSE & Cross-Entropy

Have you ever wondered why Linear Regression uses Mean Squared Error (MSE) while Classification and LLMs use Cross-Entropy Loss? Neither formula was invented by accident—both come directly from Maximum Likelihood Estimation (MLE)!

Given a dataset of NN independent training examples, MLE chooses the model weights θ\boldsymbol{\theta} that maximize the joint probability of observing the true training labels:

θMLE=arg⁡max⁡θ∏i=1NP(yi∣xi;θ)\boldsymbol{\theta}_{\text{MLE}} = \arg\max_{\boldsymbol{\theta}} \prod_{i=1}^{N} P(y_i \mid \mathbf{x}_i; \boldsymbol{\theta})

Because multiplying thousands of tiny probabilities (like 0.01×0.02×…0.01 \times 0.02 \times \dots) causes numerical underflow to 0.00.0 in floating-point hardware, we take the Natural Logarithm (ln⁡\ln)—which turns products into sums—and negate it so we can minimize Negative Log-Likelihood (NLL):

LNLL(θ)=−∑i=1Nln⁡P(yi∣xi;θ)\mathcal{L}_{\text{NLL}}(\boldsymbol{\theta}) = -\sum_{i=1}^{N} \ln P(y_i \mid \mathbf{x}_i; \boldsymbol{\theta})
1. Gaussian Noise   ⟹  \implies MSE Loss

If you assume continuous targets have Gaussian noise y∼N(y^,σ2)y \sim \mathcal{N}(\hat{y}, \sigma^2), taking −ln⁡P(y∣y^)-\ln P(y \mid \hat{y}) cancels the exponential e−(y−y^)2e^{-(y-\hat{y})^2} and leaves Mean Squared Error (y−y^)2(y - \hat{y})^2!

2. Bernoulli/Categorical   ⟹  \implies Cross-Entropy

If you assume discrete labels follow a Bernoulli or Categorical distribution, taking −ln⁡P(y∣y^)-\ln P(y \mid \hat{y}) yields Cross-Entropy Loss −∑ykln⁡(y^k)-\sum y_k \ln(\hat{y}_k)!

⚡ Knowledge Check

Why do we minimize Negative Log-Likelihood −∑ln⁡P(yi∣xi)-\sum \ln P(y_i \mid \mathbf{x}_i) during neural network training instead of directly maximizing the product of probabilities ∏P(yi∣xi)\prod P(y_i \mid \mathbf{x}_i)?

A) Logs turn products into sums, preventing float underflow and simplifying derivatives▼
✓ Correct!Multiplying thousands of numbers between 00 and 11 underflows 32-bit floats to zero. Because ln⁡(a⋅b)=ln⁡(a)+ln⁡(b)\ln(a \cdot b) = \ln(a) + \ln(b) is strictly monotonic, minimizing −∑ln⁡P-\sum \ln P finds the exact same optimal weights safely.
B) Because probabilities can be negative▼
✕ Incorrect.Probabilities are strictly bounded between 00 and 11. Notice however that for P∈(0,1)P \in (0, 1), ln⁡(P)\ln(P) is negative, which is why we negate it (−ln⁡P-\ln P) to get a positive Loss value.

  1. Visual Explanation: From Logits to Probability Distributions

Look at how a neural network converts raw, unnormalized real numbers (Logits) into a valid Categorical Probability Distribution via Softmax, and then evaluates its prediction using Negative Log-Likelihood (Cross-Entropy):

  1. Python Implementation: Softmax Distributions, Bayes' Rule & NLL

Here is a complete NumPy implementation showing how to turn raw logits into a temperature-scaled probability distribution, compute Cross-Entropy (NLL) loss, and run a Bayesian belief update:

probability_and_bayes.pyPython 3 · NumPy
import numpy as np # 1. Numerically Stable Softmax (Logits -> Categorical Distribution)def softmax(logits, temperature=1.0):    scaled = logits / temperature    shifted = scaled - np.max(scaled)       # Subtract max for numerical stability    expScores = np.exp(shifted)    return expScores / np.sum(expScores) logits = np.array([2.0, 1.0, 0.1])probs = softmax(logits, temperature=1.0)print("Categorical Probabilities:", np.round(probs, 4))  # [0.659, 0.2424, 0.0986]print("Sum of Probabilities:", round(float(np.sum(probs)), 2))  # 1.0 # 2. Maximum Likelihood / Negative Log-Likelihood (Cross-Entropy Loss)trueClassIndex = 0                          # Ground-truth token is index 0nllLoss = -np.log(probs[trueClassIndex])print("Negative Log-Likelihood Loss:", round(float(nllLoss), 4))  # 0.417 # 3. Bayes Rule Posterior Update (Spam Filter Example)priorSpam = 0.20                            # P(Spam)priorHam = 0.80                             # P(Ham)likelihoodFreeGivenSpam = 0.50              # P("Free" | Spam)likelihoodFreeGivenHam = 0.05               # P("Free" | Ham) evidenceFree = (likelihoodFreeGivenSpam * priorSpam) + (likelihoodFreeGivenHam * priorHam)posteriorSpam = (likelihoodFreeGivenSpam * priorSpam) / evidenceFreeprint("Posterior P(Spam | Free):", round(posteriorSpam, 4))     # 0.7143

Pro Tip (LLM Temperature Scaling): Notice the temperature parameter in softmax(logits / temperature)! Lowering temperature (T→0.1T \to 0.1) sharpens the probability distribution so the top token gets ≈99%\approx 99\% probability (deterministic/greedy), while raising temperature (T→1.5T \to 1.5) flattens the distribution toward a Uniform distribution (more creative/random sampling).

Key Points

✓Supervised classifiers and Large Language Models are conditional probability engines that compute P(y∣x)P(y \mid \mathbf{x})—a probability distribution over outputs given input x\mathbf{x}.
✓Expected Value E[X]\mathbb{E}[X] measures the probability-weighted center of a distribution, while Variance Var(X)=E[(X−μ)2]\text{Var}(X) = \mathbb{E}[(X-\mu)^2] measures its spread.
✓Bernoulli distributions model binary classification (Sigmoid), Categorical distributions model multi-class classification and LLM vocabulary tokens (Softmax), and Gaussian distributions model continuous noise and weight initializations.
✓Bayes' Rule (P(H∣E)=P(E∣H)P(H)P(E)P(H \mid E) = \dfrac{P(E \mid H)P(H)}{P(E)}) updates a Prior belief P(H)P(H) with observed Likelihood P(E∣H)P(E \mid H) to produce an updated Posterior belief P(H∣E)P(H \mid E).
✓Maximum Likelihood Estimation (MLE) trains models by minimizing Negative Log-Likelihood (−∑ln⁡P(yi∣xi;θ)-\sum \ln P(y_i \mid \mathbf{x}_i; \boldsymbol{\theta})), which directly yields Mean Squared Error for Gaussian targets and Cross-Entropy Loss for categorical targets.

Common Mistakes

✕ Confusing P(A | B) with P(B | A) (The Prosecutor's / Base Rate Fallacy).

The probability of a positive test given a rare disease P(+∣D)P(+ \mid D) is completely different from the probability of having the disease given a positive test P(D∣+)P(D \mid +). Always weight by the Prior P(D)P(D) using Bayes' Rule.

✕ Assuming a continuous Gaussian PDF value f(x) is a probability bounded by 1.

For continuous random variables, f(x)f(x) is a probability density (height of the curve), which can exceed 1.01.0 when σ\sigma is very small. Only the area under the curve over an interval ∫abf(x)dx\int_a^b f(x)dx is a probability.

✕ Assuming features are conditionally independent when they are heavily correlated.

The Naive Bayes classifier assumes P(x1,x2∣y)=P(x1∣y)P(x2∣y)P(x_1, x_2 \mid y) = P(x_1 \mid y)P(x_2 \mid y). When input features are duplicate or strongly correlated, Naive Bayes double-counts evidence and produces overconfident probabilities near 00 or 11.

✕ Computing Softmax without subtracting the maximum logit first.

In Python, np.exp(1000) overflows to inf and produces NaN probabilities. Always shift logits by subtracting np.max(logits) before exponentiating.

The Big Picture

Deterministic Rule Systems (Brittle Logic)

Noisy Input → Hard Yes/No Thresholds → No Confidence Score → Fails on Ambiguity

Probabilistic AI (Distributions + Maximum Likelihood)

Input x → Model Outputs Distribution P(y | x) → Minimize Negative Log-Likelihood → Calibrated Predictions

The important conceptual shift is realizing that every loss function you use in Machine Learning—from Mean Squared Error to Cross-Entropy—is really just Maximum Likelihood Estimation fitting a probability distribution to your data.

Remember: Linear Algebra gives AI its compute engine, Calculus gives AI its learning compass, and Probability gives AI its language for uncertainty and loss functions.