PIXELBANKv9.1.0
Menu
Back to ML Study Plan
Week 3-4

Chapter 2: Linear Regression

Master the foundational algorithm of machine learning. Learn to predict continuous values with linear models, optimize using gradient descent, handle multiple features with matrix operations, and prevent overfitting with L1/L2 regularization.

Chapter Overview

Linear regression is arguably the most important algorithm to understand deeply. Not because it's the most powerful—it's not—but because it introduces nearly every concept you'll need for more complex models: loss functions, optimization, overfitting, regularization, and the bias-variance tradeoff.

The core idea is beautifully simple: model the relationship between inputs and outputs as a weighted sum plus a bias term. The weights tell us how much each feature contributes to the prediction. The challenge is finding the optimal weights—those that minimize the prediction errors.

For simple problems, we can solve for optimal weights analytically using calculus. But this approach doesn't scale. Gradient descent provides an iterative alternative: start with random weights, compute how the loss changes with small weight adjustments (the gradient), then nudge weights in the direction that decreases loss. Repeat until convergence.

Real problems have many features interacting in complex ways. Multiple regression extends the model to handle this, fitting a hyperplane in high-dimensional space. The normal equation gives a closed-form solution, but gradient descent is more practical for large datasets.

Regularization addresses a critical problem: models can fit training data too well, capturing noise rather than signal. By penalizing large weights, L2 (Ridge) regularization shrinks coefficients toward zero, while L1 (Lasso) can eliminate features entirely. This tradeoff between fitting the data and keeping the model simple is fundamental to all of machine learning.

This chapter covers:

  • Simple Linear Regression: Fitting a line with MSE loss and the closed-form solution
  • Gradient Descent: Iterative optimization, variants (batch, stochastic, mini-batch), and learning rate schedules
  • Multiple Regression: Matrix formulation, feature scaling, and handling many features
  • Polynomial Features: Capturing nonlinear relationships while staying in the linear regression framework
  • Regularization: L1 (Lasso), L2 (Ridge), and Elastic Net for controlling complexity
  • Model Assumptions & Diagnostics: When linear regression is appropriate and how to verify

Chapter Roadmap

Click any topic to jump in

Explainer video
1
Simple Linear Regression

The starting point — fit a line through data by minimizing squared errors with a closed-form solution.

Linear ModelResidualsMean Squared Error (MSE)Why MSE?Closed-Form SolutionR² (Coefficient of Determination)
Scaling the approach

Optimization for large data and extending to multiple features

2
Gradient Descent

Iterative optimization when closed-form solutions don't scale — follow the negative gradient to minimize loss.

GradientUpdate RuleMSE Gradient for Linear RegressionLearning Rate (α)Batch Gradient DescentStochastic Gradient Descent (SGD)Mini-Batch Gradient DescentConvergence
3
Multiple Regression

Extend to many features with matrix notation, the normal equation, and feature scaling for stable training.

Multiple FeaturesMatrix FormNormal EquationFeature Scaling: StandardizationFeature Scaling: NormalizationFeature ImportanceMulticollinearity
Handling complexity

Nonlinear patterns and controlling model complexity

4
Polynomial Features

Capture nonlinear patterns while staying inside the linear regression framework via feature transformations.

Polynomial RegressionFeature TransformationInteraction TermsDegree SelectionCurse of DimensionalitySklearn PolynomialFeatures
5
Regularization

Prevent overfitting by penalizing large weights — L1 for sparsity, L2 for shrinkage, Elastic Net for both.

Ridge Regression (L2)Lasso Regression (L1)Why L1 Produces SparsityElastic NetRegularization Strength (λ)Standardize Before Regularization
Validating your model
6
Assumptions & Diagnostics

When linear regression is appropriate and how to verify — residual analysis, influential points, and model checks.

Linearity AssumptionIndependenceHomoscedasticityNormality of ResidualsNo MulticollinearityResidual AnalysisInfluential Points

You have two columns of numbers — house size and sale price, study hours and exam score — and you are asked what the second column would be for a value of the first that nobody recorded. Looking it up is impossible; you have to invent a rule that turns any input into a plausible output, and fit that rule to the data you do have.

That rule is a model: a formula containing a few unknown numbers called parameters. Choosing them from data is fitting, and "best fit" only means something once you define a loss — one number saying how wrong the model currently is. Chapter 1's Introduction to Machine Learning named these words; this is the first topic where you see all of them made concrete on a model you can solve by hand.

The arc follows the order you would actually derive it. We write down the straight-line model, define the residual (one point's error), average the squared residuals into MSE, defend the squaring, solve for the parameters exactly with calculus, and close with R2R^2 — the score that says whether fitting the line was worth it at all.

Definition

Simple linear regression models a numeric target yy as an affine function of one numeric input xx, y^=w0+w1x\hat{y} = w_0 + w_1 x, choosing the parameters w0w_0 and w1w_1 to minimise the mean squared error between predictions and observations. Because that objective is a convex quadratic, the minimiser is unique and available in closed form — no iterative search needed.

In this topic

1Linear Model
2Residuals
3Mean Squared Error (MSE)
4Why MSE?
5Closed-Form Solution
6R² (Coefficient of Determination)
1 of 6
Linear Model

y^=w0+w1x\hat{y} = w_0 + w_1 x

Without a model you can only repeat values you have already observed. A model is a formula with unknown numbers. Here xx is the single input feature and y^\hat{y} — "y-hat" — is the predicted output, hatted to separate it from an observed yiy_i. The parameters w0w_0 and w1w_1 are what we choose: the intercept w0w_0 is the prediction at x=0x = 0, the slope w1w_1 is how far y^\hat{y} moves per one-unit increase in xx. This is chapter 0's y=mx+by = mx + b in ML notation. Its assumption is strong: if the truth curves, no (w0,w1)(w_0, w_1) repairs it and the leftover errors are systematic, not random.

Mathematical Intuition

The model y^=w0+w1x\hat{y} = w_0 + w_1 x defines an affine mapping from input space R\mathbb{R} to output space R\mathbb{R}. The parameter w1w_1 is the rate of change ∂y^∂x\frac{\partial \hat{y}}{\partial x}, meaning each unit increase in xx shifts the prediction by exactly w1w_1. The intercept w0w_0 anchors the line at x=0x = 0. With nn data points, we have 2 free parameters and nn constraints — the system is overdetermined when n>2n > 2, which is why we need a loss function to define "best fit" rather than exact interpolation.

Example:

Model: ŷ = 50 + 10x. What's the prediction when x=3?

2 of 6
Residuals

ei=yi−y^ie_i = y_i - \hat{y}_i

A model can only be fitted if you can measure how wrong it is at each point; the residual is that per-point measurement. For training example ii, yiy_i is the observed target, y^i=w0+w1xi\hat{y}_i = w_0 + w_1 x_i is what the line predicts at that example's input, and eie_i is the signed vertical gap: positive means the model underpredicted, negative means it overpredicted. Residuals are not properties of the data: move the line and every eie_i changes, which is what fitting exploits. They measure vertical error only, so xx is assumed noise-free; residuals that curve or fan out warn that the linear model above is the wrong shape.

Mathematical Intuition

The residual ei=yi−y^ie_i = y_i - \hat{y}_i measures the signed vertical distance from observation ii to the fitted line. For the OLS estimator, ∑i=1nei=0\sum_{i=1}^{n} e_i = 0 and ∑i=1nxiei=0\sum_{i=1}^{n} x_i e_i = 0 — these are the normal equations. The first says residuals balance out on average; the second says residuals are uncorrelated with the predictor. Together, these two conditions uniquely determine w0w_0 and w1w_1. Violations of these orthogonality conditions indicate the model has not converged or was computed incorrectly.

Example:

Actual y=100, predicted ŷ=85. What's the residual? Interpretation?

3 of 6
Mean Squared Error (MSE)

MSE=1n∑i=1n(yi−y^i)2MSE = \frac{1}{n} \sum_{i=1}^{n} (y_i - \hat{y}_i)^2

Residuals give nn numbers, but choosing parameters needs one number to minimise — and simply adding residuals fails, because positive and negative misses cancel and a terrible line can total zero. Squaring destroys the signs, so the mean of the squared residuals over the nn training examples is zero only for a perfect fit. That number is the loss. Dividing by nn makes it per-example, so datasets of different sizes stay comparable. Its units are the square of yy's units, which is why people report MSE\sqrt{\text{MSE}} (RMSE) instead. The price of squaring is outlier sensitivity: one point ten units off costs as much as a hundred points one unit off.

Mathematical Intuition

MSE =1n∑(yi−y^i)2= \frac{1}{n}\sum(y_i - \hat{y}_i)^2 is a convex, differentiable function of the weights. Squaring amplifies large errors: an error of 4 contributes 16 to the sum, while four errors of 1 contribute only 4. This makes MSE sensitive to outliers — a single point far from the line can dominate the entire loss. Statistically, minimizing MSE is equivalent to maximum likelihood estimation when the noise ϵi∼N(0,σ2)\epsilon_i \sim N(0, \sigma^2), connecting the geometric idea of "closest line" to the probabilistic idea of "most likely parameters."

Example:

Residuals: [2, -3, 1, -2]. Calculate MSE.

4 of 6
Why MSE?

Squaring residuals is a choice, not a law, with three defences. Probabilistic: assume every observation is the line plus independent noise ϵi∼N(0,σ2)\epsilon_i \sim \mathcal{N}(0, \sigma^2), where σ2\sigma^2 is the noise variance. Each point's likelihood is proportional to e−ei2/2σ2e^{-e_i^2 / 2\sigma^2}, the dataset's likelihood is their product, and taking the logarithm turns that product into −12σ2∑ei2-\frac{1}{2\sigma^2}\sum e_i^2 plus constants — so maximising likelihood is literally minimising MSE. Computational: e2e^2 is differentiable everywhere, while ∣e∣|e| has a corner at zero. Behavioural: doubling an error quadruples its cost. That last is also the flaw — MSE estimates the conditional mean, so one outlier drags the line, whereas MAE estimates the median and shrugs.

Mathematical Intuition

The MSE loss surface J(w0,w1)=1n∑(yi−w0−w1xi)2J(w_0, w_1) = \frac{1}{n}\sum(y_i - w_0 - w_1 x_i)^2 is a paraboloid in parameter space — it has exactly one global minimum and no local minima. The Hessian matrix H=2nXTXH = \frac{2}{n} X^T X is positive semi-definite, guaranteeing convexity. The absolute error ∣ei∣|e_i| has a corner at zero where the derivative is undefined, making gradient-based optimization problematic. MSE's smoothness everywhere means the gradient ∇J\nabla J exists at every point, providing a well-defined direction toward the optimum.

Example:

Why not use |error| (MAE) instead of error²?

5 of 6
Closed-Form Solution

w1=∑(xi−xˉ)(yi−yˉ)∑(xi−xˉ)2=Cov(x,y)Var(x)w_1 = \frac{\sum(x_i - \bar{x})(y_i - \bar{y})}{\sum(x_i - \bar{x})^2} = \frac{Cov(x,y)}{Var(x)}

Because MSE is a convex bowl in (w0,w1)(w_0, w_1), its one flat point is the minimum: set both partial derivatives to zero and solve. Differentiating J=1n∑(yi−w0−w1xi)2J = \frac{1}{n}\sum(y_i - w_0 - w_1 x_i)^2 with respect to w0w_0 gives −2n∑ei=0-\frac{2}{n}\sum e_i = 0 — residuals must sum to zero — which rearranges to w0=yˉ−w1xˉw_0 = \bar{y} - w_1 \bar{x}, where xˉ\bar{x} and yˉ\bar{y} are sample means. Substituting back and differentiating with respect to w1w_1 produces the formula shown: the covariance over the variance of xx. It fails only when ∑(xi−xˉ)2=0\sum(x_i - \bar{x})^2 = 0 — every xix_i identical, so there is no slope to estimate; near-constant xx makes it numerically fragile.

Mathematical Intuition

Setting ∂J∂w1=0\frac{\partial J}{\partial w_1} = 0 yields w1=Cov(X,Y)Var(X)w_1 = \frac{\text{Cov}(X, Y)}{\text{Var}(X)}. This has a clean geometric interpretation: the slope equals the correlation coefficient times the ratio of standard deviations, w1=rxy⋅sysxw_1 = r_{xy} \cdot \frac{s_y}{s_x}. When features are perfectly correlated (r=±1r = \pm 1), the line passes through every point. When r=0r = 0, the slope is zero and the best prediction is just yˉ\bar{y}. The closed-form solution requires O(n)O(n) time for simple regression, making it efficient for small feature counts.

Example:

Cov(x,y) = 15, Var(x) = 5, x̄ = 10, ȳ = 25. Find the line.

6 of 6
R² (Coefficient of Determination)

R2=1−∑(yi−y^i)2∑(yi−yˉ)2R^2 = 1 - \frac{\sum(y_i - \hat{y}_i)^2}{\sum(y_i - \bar{y})^2}

MSE answers "how wrong?" in squared units of yy, so its value alone cannot say whether a fit is good. R2R^2 fixes that by comparing against the laziest model: ignore xx and always predict yˉ\bar{y}. The denominator ∑(yi−yˉ)2\sum(y_i - \bar{y})^2 is that baseline's squared error — the total variation in yy; the numerator ∑(yi−y^i)2\sum(y_i - \hat{y}_i)^2 is what your line leaves. Their ratio is the fraction of baseline error still unexplained, so R2R^2 is the fraction removed: 0.8 means four fifths of the variation in yy is explained by xx. It goes negative when a model beats nothing, and never falls when you add a feature, even a useless one.

Mathematical Intuition

R2=1−SSresSStotR^2 = 1 - \frac{SS_{res}}{SS_{tot}} decomposes total variance into explained and unexplained parts: SStot=SSreg+SSresSS_{tot} = SS_{reg} + SS_{res}. For simple regression, R2=rxy2R^2 = r_{xy}^2 — the square of the Pearson correlation. Adding features can only increase R2R^2 (or keep it the same), which is why adjusted R2=1−(1−R2)(n−1)n−p−1R^2 = 1 - \frac{(1-R^2)(n-1)}{n-p-1} penalizes model complexity by accounting for the number of parameters pp. A negative R2R^2 means the model fits worse than a horizontal line at yˉ\bar{y}.

Example:

SSres = 200, SStot = 1000. Calculate and interpret R².

Code Examples

Fitting by hand, and checking the two conditions the optimum must satisfy

import numpy as np

x = np.array([1.0, 2.0, 3.0, 4.0, 5.0])
y = np.array([2.0, 4.0, 5.0, 4.0, 5.0])

xb, yb = x.mean(), y.mean()
w1 = np.sum((x - xb) * (y - yb)) / np.sum((x - xb) ** 2)
w0 = yb - w1 * xb
e = y - (w0 + w1 * x)

print(f"w0 = {w0:.4f}   w1 = {w1:.4f}")
print(f"sum of residuals   = {e.sum():.2e}")
print(f"sum of x_i * e_i   = {np.dot(x, e):.2e}")

mse = lambda a, b: np.mean((y - (a + b * x)) ** 2)
print(f"MSE at optimum     = {mse(w0, w1):.6f}")
for d in (-0.2, 0.2):
    print(f"MSE at w1 {d:+.1f}     = {mse(w0, w1 + d):.6f}")

The same five points as the theory exercise, so you can check the hand arithmetic: it prints w0 = 2.2000 and w1 = 0.6000. The two sums come out at the 1e-16 level — not merely small. They are the stationarity conditions from the derivation, exact up to floating point. MSE at the optimum is 0.480000, and nudging the slope by either -0.2 or +0.2 raises it to 0.920000: moving away from the minimum costs you in both directions and by the same amount, which is what a symmetric convex bowl looks like numerically.

Outliers, negative R², and why R² never goes down

import numpy as np
from sklearn.linear_model import LinearRegression
from sklearn.metrics import r2_score

rng = np.random.default_rng(1)
x = rng.uniform(0, 10, 60)
y = 2.0 * x + 1.0 + rng.normal(0, 1.0, 60)
X1 = x.reshape(-1, 1)

base = LinearRegression().fit(X1, y)
print(f"clean slope        = {base.coef_[0]:.3f}")
y_out = y.copy(); y_out[x.argmax()] += 60.0   # corrupt the highest-x point
print(f"slope w/ 1 outlier = {LinearRegression().fit(X1, y_out).coef_[0]:.3f}")

X2 = np.column_stack([x, rng.normal(0, 1, 60)])   # add a pure-noise feature
fit2 = LinearRegression().fit(X2, y)
print(f"R2, 1 feature      = {r2_score(y, base.predict(X1)):.6f}")
print(f"R2, + noise column = {r2_score(y, fit2.predict(X2)):.6f}")
print(f"R2 of always-zero  = {r2_score(y, np.zeros_like(y)):.3f}")

Three of the topic's claims, verified. The clean slope is 2.031, close to the true 2.0. Adding 60 to the single highest-x target drags it to 2.597 — one corrupted point in sixty, and the fit chases it. That is MSE's outlier sensitivity, and the point is corrupted at the edge of the x-range because leverage is highest there. Appending a column of pure random noise moves R² from 0.980120 to 0.980122: a rise in the sixth decimal, and never a fall, because the extra parameter can only reduce training error. That is exactly why later topics need adjusted R² and regularisation. The always-zero predictor scores -3.961, far below the mean baseline of 0.

Theory Exercise

Problem:

Given data points (1,2), (2,4), (3,5), (4,4), (5,5), calculate the best-fit line using the closed-form solution.

Hints:
  • First calculate means: x̄ and ȳ
  • Use the formula for w₁ involving covariance and variance
  • Then calculate w₀ = ȳ - w₁x̄

Coding Exercise

Problem:

Generate noisy data from y = 2.5x + 1 with a fixed seed, then implement simple linear regression from scratch using the closed-form (least-squares) formulas for slope and intercept. Verify your weights against sklearn's LinearRegression and report the R² of your fit.

Hints:
  • Slope = sum((x - x̄)(y - ȳ)) / sum((x - x̄)²); intercept = ȳ - slope·x̄.
  • Compute predictions ŷ = slope·x + intercept, then use sklearn.metrics.r2_score(y, ŷ).
  • Fit LinearRegression on x.reshape(-1, 1) and compare your slope to sk.coef_[0] and intercept to sk.intercept_.