PIXELBANKv9.1.0
Menu
Back to Foundations

Chapter 4: Probability & Statistics

Deep dive into probability theory, statistical distributions, estimation, hypothesis testing, Bayesian inference, and sampling methods for machine learning.

Chapter Overview

Probability and statistics are the mathematical language of uncertainty — and uncertainty is everywhere in machine learning. Every prediction carries a confidence level, every dataset is a sample from an unknown distribution, and every training procedure is an optimization over a probabilistic objective.

The Mathematical Foundations chapter introduced probability basics: distributions, Bayes' theorem, and expectation. This chapter takes you much deeper. You will build rigorous foundations in probability theory, master the full family of distributions that appear throughout ML, and learn the statistical tools that let you draw reliable conclusions from data.

We start with the axioms that make probability logically consistent, then systematically build up to the tools that practicing ML engineers use daily:

  • Probability Fundamentals: The axiomatic foundation — sample spaces, conditional probability, independence, and the law of total probability
  • Discrete Distributions: Bernoulli, Binomial, Poisson, and Geometric — the building blocks of classification, counting, and sequential models
  • Continuous Distributions: Gaussian, Exponential, Beta, Gamma, and the Central Limit Theorem — why the bell curve appears everywhere
  • Joint & Marginal Distributions: How multiple random variables interact — the theory behind multivariate models
  • Statistical Estimation: Point estimators, confidence intervals, and bootstrap — quantifying uncertainty in your estimates
  • Hypothesis Testing: P-values, significance tests, and A/B testing — making data-driven decisions with statistical rigor
  • Bayesian Inference: The full Bayesian framework — priors, posteriors, conjugate families, and the MAP-regularization connection
  • Sampling Methods: Monte Carlo, importance sampling, and MCMC — computational tools for intractable distributions

Chapter Roadmap

Click any topic to jump in

Explainer video
1
Fundamentals

Sample spaces, axioms, conditional probability, independence — the rules everything builds on.

Sample Spaces & EventsAxioms of ProbabilityConditional ProbabilityIndependence & Conditional IndependenceLaw of Total Probability
Two types of random variables

Discrete counts and continuous measurements

2
Discrete Distributions

Bernoulli, binomial, Poisson, geometric — modeling countable outcomes.

Bernoulli DistributionBinomial DistributionPoisson DistributionGeometric DistributionPMFs and CDFs
3
Continuous Distributions

PDFs, uniform, exponential, beta, gamma, CLT — modeling continuous variables.

Probability Density FunctionsUniform DistributionExponential DistributionBeta DistributionGamma DistributionCentral Limit Theorem
Combining multiple variables
4
Joint & Marginal

Joint distributions, marginalization, conditioning, multivariate Gaussian — reasoning about multiple variables.

Joint Probability DistributionsMarginal DistributionsConditional DistributionsChain Rule of ProbabilityMultivariate Gaussian
Learning from data

Estimation and testing

5
Estimation

Point estimators, bias-variance, confidence intervals, MLE vs MAP — learning from data.

Point EstimatorsBias-Variance of EstimatorsConfidence IntervalsBootstrap MethodMLE vs MAP Estimation
6
Hypothesis Testing

P-values, type I/II errors, t-tests, A/B testing — rigorous decisions from experiments.

Null & Alternative HypothesesP-ValuesType I & Type II ErrorsT-TestChi-Squared TestA/B Testing in ML
Advanced inference

Bayesian thinking and computational methods

7
Bayesian Inference

Prior, likelihood, posterior, conjugate priors, MAP — updating beliefs with evidence.

Prior, Likelihood & PosteriorConjugate PriorsMaximum A Posteriori (MAP) EstimationPosterior Predictive DistributionBayesian vs Frequentist Inference
8
Sampling Methods

Monte Carlo, rejection/importance sampling, MCMC — approximating intractable distributions.

Monte Carlo EstimationRejection SamplingImportance SamplingMarkov Chain Monte Carlo (MCMC)MCMC Diagnostics & Practical Considerations

A spam filter sees the word "winner" in an email. How worried should it be? The useful number is not the fraction of spam emails that contain "winner". It is the fraction of emails containing "winner" that turn out to be spam, and those two numbers can differ by a factor of ten. Confusing them is one of the most common reasoning errors in medicine, law, and machine learning. Avoiding it takes a precise language for uncertainty: which outcomes are possible, how probability is spread over them, and how that spread changes when new information arrives.

The previous chapter, Mathematical Foundations, used probabilities informally for expectation, variance, and entropy. This chapter rebuilds that material from the ground up. A classifier's softmax, a language model's next-token distribution, and a Bayesian posterior all obey the same three axioms.

We begin with sample spaces and events, state Kolmogorov's axioms, define conditional probability, separate independence from conditional independence, and finish with the law of total probability, which splits a hard probability into easy cases.

Definition

A probability space has three parts: a sample space Ω\Omega of outcomes, a collection of events that are subsets of Ω\Omega, and a probability measure PP. The measure assigns each event a number between 0 and 1, gives Ω\Omega probability 1, and adds over disjoint events. The conditional probability of AA given BB is P(A∩B)/P(B)P(A \cap B)/P(B) whenever P(B)>0P(B) > 0.

In this topic

1Sample Spaces & Events
2Axioms of Probability
3Conditional Probability
4Independence & Conditional Independence
5Law of Total Probability
1 of 5
Sample Spaces & Events

Ω={all possible outcomes},P(Ω)=1\Omega = \{\text{all possible outcomes}\}, \quad P(\Omega) = 1

Before you can assign a probability, you must say what could happen. The sample space Ω\Omega in the formula is the set of all outcomes of an experiment, and an event is any subset of it, such as "the die shows an even number", which is {2,4,6}\{2, 4, 6\}. Events combine with set operations. The complement AcA^c is everything outside AA, the union A∪BA \cup B means A or B, and the intersection A∩BA \cap B means both. P(Ω)=1P(\Omega) = 1 says some outcome always happens. The usual failure is a badly chosen Ω\Omega, such as counting "two heads, one of each, two tails" as three equally likely cases.

Mathematical Intuition

The sample space Ω\Omega is the universal set in measure theory — every probability statement is a measure on subsets of Ω\Omega. In ML, feature spaces are high-dimensional sample spaces where each point is one possible input configuration.

Example:

A card is drawn from a standard 52-card deck. Define the sample space and find P(drawing a face card).

2 of 5
Axioms of Probability

1.  P(A)≥02.  P(Ω)=13.  P(A∪B)=P(A)+P(B) if A∩B=∅\begin{aligned} 1.&\; P(A) \geq 0 \\ 2.&\; P(\Omega) = 1 \\ 3.&\; P(A \cup B) = P(A) + P(B) \text{ if } A \cap B = \emptyset \end{aligned}

Why only three rules? Kolmogorov showed in 1933 that non-negativity, normalization, and additivity over disjoint events are enough to derive everything else. In the formula, AA and BB are events, Ω\Omega is the sample space, and ∅\emptyset is the empty event. Additivity gives the complement rule P(Ac)=1−P(A)P(A^c) = 1 - P(A) and P(∅)=0P(\emptyset) = 0. Splitting A∪BA \cup B into disjoint pieces gives inclusion-exclusion, P(A∪B)=P(A)+P(B)−P(A∩B)P(A \cup B) = P(A) + P(B) - P(A \cap B). A softmax layer satisfies the axioms by construction, because exponentials are positive and the outputs are normalized. Independent sigmoid scores do not, since their class scores need not sum to 1.

Mathematical Intuition

Kolmogorov's axioms define probability as a measure with total mass 1. The additivity axiom generalizes to countable unions (σ\sigma-additivity), which is why we need σ\sigma-algebras. Softmax enforces axioms 1-2 by construction: softmax(z)i=ezi/∑jezj≥0\text{softmax}(z)_i = e^{z_i}/\sum_j e^{z_j} \geq 0 and sums to 1.

Example:

Verify the axioms for a fair six-sided die. Then use inclusion-exclusion: P(even OR greater than 4)?

3 of 5
Conditional Probability

P(A∣B)=P(A∩B)P(B),P(B)>0P(A|B) = \frac{P(A \cap B)}{P(B)}, \quad P(B) > 0

Learning that BB happened shrinks the world to the outcomes inside BB. The formula renormalizes: P(A∣B)P(A \mid B) is the share of BB's probability that also lies in AA, and it is defined only when P(B)>0P(B) > 0. Rearranging gives the product rule P(A∩B)=P(A∣B)P(B)P(A \cap B) = P(A \mid B) P(B), which the rest of the chapter uses constantly. Every classifier estimates a conditional probability such as P(spam∣words)P(\text{spam} \mid \text{words}), and a language model estimates P(next token∣context)P(\text{next token} \mid \text{context}). The classic failure is swapping P(A∣B)P(A \mid B) for P(B∣A)P(B \mid A), the prosecutor's fallacy from the spam example in the introduction.

Mathematical Intuition

Conditioning P(A∣B)=P(A∩B)/P(B)P(A|B) = P(A \cap B)/P(B) is geometrically a projection: you collapse the probability mass onto the subspace where BB holds, then renormalize. In neural nets, attention scores are conditional probabilities — each token attends to others proportionally to softmax(QKT/d)\text{softmax}(QK^T/\sqrt{d}).

Example:

In a class of 100 students: 40 study math, 30 study CS, 10 study both. If a student studies math, what is P(they also study CS)?

4 of 5
Independence & Conditional Independence

P(A∩B)=P(A)⋅P(B)⇔A⊥BP(A \cap B) = P(A) \cdot P(B) \quad \Leftrightarrow \quad A \perp B

Independence means one event carries no information about another. In the formula, A⊥BA \perp B holds exactly when the joint probability factors as P(A)P(B)P(A) P(B), which is equivalent to P(A∣B)=P(A)P(A \mid B) = P(A). Conditional independence applies the same test inside a context CC: P(A∩B∣C)=P(A∣C)P(B∣C)P(A \cap B \mid C) = P(A \mid C) P(B \mid C). Neither kind implies the other. Two symptoms can be dependent overall yet independent given the disease that causes both. Naive Bayes assumes words are conditionally independent given the class, which cuts its parameter count from exponential to linear in the vocabulary. A frequent error is confusing independent with disjoint: disjoint events with positive probability are always dependent.

Mathematical Intuition

Independence means the joint distribution factors: p(x,y)=p(x)p(y)p(x,y) = p(x)p(y). This is a massive dimensionality reduction — from O(n2)O(n^2) parameters to O(n)O(n). Naive Bayes exploits this: p(x1,…,xd∣y)=∏ip(xi∣y)p(x_1,\ldots,x_d|y) = \prod_i p(x_i|y) reduces exponential parameter space to linear.

Example:

A fair coin is flipped twice. Are the two flips independent? Verify using the definition.

5 of 5
Law of Total Probability

P(B)=∑i=1nP(B∣Ai)⋅P(Ai)P(B) = \sum_{i=1}^{n} P(B|A_i) \cdot P(A_i)

Some probabilities are hard to compute directly but easy once you know which case you are in. If A1,…,AnA_1, \ldots, A_n partition the sample space, meaning they are disjoint and together cover it, the formula writes P(B)P(B) as a weighted average of the case probabilities P(B∣Ai)P(B \mid A_i), with weights P(Ai)P(A_i). It follows from the product rule of conditional probability plus additivity. The same move marginalizes a latent variable, P(x)=∑zP(x∣z)P(z)P(x) = \sum_z P(x \mid z) P(z), which is how mixture models score data, and it produces the denominator of Bayes' theorem. The failure mode is a set of cases that overlap or leave gaps, so the weights no longer sum to 1.

Mathematical Intuition

Total probability is marginalization in disguise: P(B)=∑iP(B∣Ai)P(Ai)P(B) = \sum_i P(B|A_i)P(A_i) is just ∫p(b∣a)p(a)da\int p(b|a)p(a)da in continuous form. In mixture models like GMMs, this is exactly how you compute the data likelihood — sum over all cluster assignments weighted by mixing coefficients.

Example:

Factory has 3 machines: M1 (50% of output, 2% defect), M2 (30%, 3% defect), M3 (20%, 5% defect). Find P(defective item).

Theory Exercise

Problem:

A medical test has 95% sensitivity (P(+|disease) = 0.95) and 90% specificity (P(-|healthy) = 0.90). The disease prevalence is 1%. Using the law of total probability, find P(+). Then use Bayes' theorem to find P(disease|+).

Hints:
  • P(+) = P(+|disease)P(disease) + P(+|healthy)P(healthy)
  • P(+|healthy) = 1 - specificity = 0.10
  • P(disease|+) = P(+|disease)P(disease) / P(+)

Coding Exercise

Problem:

Verify Bayes' theorem and the law of total probability by Monte Carlo simulation. A disease has 1% prevalence. A test has 95% sensitivity and 90% specificity. Simulate 1,000,000 patients, then empirically estimate P(disease | positive test) and compare it to the analytic Bayes answer.

Hints:
  • Draw the disease status with np.random.rand(N) < 0.01.
  • Generate test results conditionally: sick patients test positive with prob 0.95, healthy with prob 0.10 (1 - specificity).
  • P(disease | +) is simply the fraction of positive testers who are actually sick.