PIXELBANKv9.1.0
Menu
Back to ML Study Plan
Week 5-6

Chapter 3: Classification

Learn to predict discrete categories with machine learning. Master logistic regression for binary classification, understand the sigmoid function and cross-entropy loss, extend to multiclass problems with softmax, and use feature engineering to handle non-linear patterns.

Chapter Overview

Classification is the task of predicting which category an input belongs to. Unlike regression which outputs continuous values, classification outputs discrete class labels: spam or not spam, cat or dog or bird, malignant or benign.

The foundational algorithm is logistic regression, which despite its name is a classification method. It works by computing a linear combination of features, then passing the result through the sigmoid function to produce a probability between 0 and 1. If this probability exceeds a threshold (typically 0.5), we predict the positive class.

The key to understanding logistic regression is the loss function. We can't use mean squared error because sigmoid creates a non-convex loss surface. Instead, we use cross-entropy (log loss), which heavily penalizes confident wrong predictions and creates a smooth, convex optimization landscape.

The decision boundary in logistic regression is a hyperplane—the set of points where the predicted probability equals 0.5. Points on one side are classified as positive, points on the other as negative. This boundary is linear in the original feature space, but we can create non-linear boundaries through feature engineering.

For problems with more than two classes, we extend to multiclass classification using techniques like one-vs-rest (train K binary classifiers) or softmax regression (directly model probabilities for all K classes).

This chapter covers:

  • Binary Classification: The sigmoid function, decision thresholds, and linear decision boundaries
  • Logistic Regression: Maximum likelihood, cross-entropy loss, and why it works
  • Multiclass: One-vs-rest, one-vs-one, and softmax approaches
  • Naive Bayes: A probabilistic classifier based on Bayes' theorem
  • Class Imbalance: Handling datasets where classes have very different frequencies
  • Feature Engineering: Polynomial features and transformations to handle non-linear patterns

Chapter Roadmap

Click any topic to jump in

1
Binary Classification

The core task — separate inputs into two classes using sigmoid probabilities and decision thresholds.

Classification vs RegressionSigmoid FunctionSigmoid PropertiesDecision ThresholdLinear Decision BoundaryThreshold Tuning
Two algorithmic approaches

Discriminative vs. generative classifiers for binary problems

2
Logistic Regression

The workhorse classifier — maximum likelihood estimation with cross-entropy loss and linear decision boundaries.

Logistic ModelCross-Entropy LossMLE ConnectionGradient for Logistic RegressionOdds and Log-OddsRegularized Logistic RegressionWeight Interpretation
3
Naive Bayes

A probabilistic approach using Bayes' theorem with conditional independence — fast, interpretable, strong baseline.

Bayes' TheoremNaive Independence AssumptionClassification RuleGaussian Naive BayesMultinomial Naive BayesLaplace Smoothing
Beyond two classes
4
Multiclass Classification

Extend binary methods to K classes with one-vs-rest, one-vs-one, or softmax regression.

One-vs-Rest (OvR)One-vs-One (OvO)Softmax (Multinomial)Softmax PropertiesCross-Entropy (Multiclass)Multiclass vs Multilabel
Real-world challenges

Handling imbalanced data and engineering better features

5
Class Imbalance

Handle skewed class distributions with resampling, class weights, threshold tuning, and proper metrics.

The ProblemResampling: OversamplingResampling: UndersamplingClass WeightsThreshold AdjustmentMetrics for Imbalanced Data
6
Feature Engineering

Transform raw inputs into informative features — encoding, scaling, selection, and dimensionality reduction.

Polynomial FeaturesFeature ScalingOne-Hot EncodingLabel EncodingTarget EncodingFeature Selection: Filter MethodsFeature Selection: Wrapper MethodsDimensionality Reduction

A bank must decide whether to approve a loan, a hospital whether a scan shows a tumor, an email service whether a message is spam. Each question has exactly two answers, and getting it wrong in one direction usually costs far more than the other. The previous chapter, Linear Regression, predicted continuous numbers such as prices. Its straight-line output is unbounded, so it cannot directly say how likely a yes is, and a number like 1.7 or minus 0.4 is not a probability.

This chapter, Classification, adapts the linear model to discrete outcomes, and this first topic sets up the two-class case. We begin by separating classification from regression. We then introduce the sigmoid function, which squashes any score into the range 0 to 1, and study its symmetry and saturation. Next we turn probabilities into decisions with a threshold, see that a linear score defines a flat decision boundary, and finish by choosing the threshold from the real costs of each kind of mistake rather than defaulting to 0.5.

Definition

Binary classification is the task of learning a function that maps an input feature vector to one of two classes, usually labeled 0 and 1. A probabilistic binary classifier outputs an estimate of the probability that the label is 1, and a decision threshold converts that probability into a hard prediction.

In this topic

1Classification vs Regression
2Sigmoid Function
3Sigmoid Properties
4Decision Threshold
5Linear Decision Boundary
6Threshold Tuning
1 of 6
Classification vs Regression

The first modeling decision is what kind of output the target is, because it fixes the loss, the metrics, and the final layer of the model. Regression predicts a continuous quantity, such as a price or a temperature, and errors are measured as distances. Classification predicts a discrete category, such as yes or no, or cat, dog, or bird, and errors are counted as right or wrong. Most classifiers first produce a probability for each class and then pick one. A common trap is treating ordered categories, such as ratings from 1 to 5, as regression or as nominal classes without thinking about which errors matter.

Mathematical Intuition

Regression models E[y∣x]∈RE[y|\mathbf{x}] \in \mathbb{R} — a continuous conditional expectation. Classification models P(y=k∣x)P(y = k | \mathbf{x}) for discrete k∈{0,1,…,K−1}k \in \{0, 1, \ldots, K-1\} — a probability distribution over categories. Using MSE for classification creates a loss surface with flat regions (gradient ≈0\approx 0) where the sigmoid saturates, stalling gradient descent. Cross-entropy fixes this because −log⁡(σ(z))-\log(\sigma(z)) has gradient σ(z)−1\sigma(z) - 1, which remains large even when σ(z)\sigma(z) is near 0 or 1.

Example:

Predict house price vs predict if house sells within 30 days. Which is regression, which is classification?

2 of 6
Sigmoid Function

σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}}

Linear regression gives a score zz anywhere on the real line, but a probability must lie between 0 and 1. The sigmoid fixes this by mapping zz to 1/(1+e−z)1/(1 + e^{-z}). Large positive scores give outputs near 1, large negative scores give outputs near 0, and z=0z = 0 gives exactly 0.5. The output is read as the probability that the label is 1. Its derivative has the convenient form σ(z)(1−σ(z))\sigma(z)(1 - \sigma(z)), which peaks at 0.25 when z=0z = 0. That form makes gradients cheap to compute, and it is why the sigmoid pairs so neatly with cross-entropy loss later in this chapter.

Mathematical Intuition

The sigmoid σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}} maps R→(0,1)\mathbb{R} \to (0, 1). Its derivative is σ′(z)=σ(z)(1−σ(z))\sigma'(z) = \sigma(z)(1 - \sigma(z)), which peaks at z=0z = 0 (value 0.250.25) and vanishes exponentially as ∣z∣→∞|z| \to \infty. The sigmoid is the canonical link function for Bernoulli-distributed responses in generalized linear models. It arises naturally from the log-odds: if log⁡P(y=1)P(y=0)=z\log \frac{P(y=1)}{P(y=0)} = z (a linear function of features), then solving for P(y=1)P(y=1) gives exactly the sigmoid. The inverse is the logit function: σ−1(p)=log⁡(p/(1−p))\sigma^{-1}(p) = \log(p/(1-p)).

Example:

Calculate σ(0) and σ(2).

3 of 6
Sigmoid Properties

Three properties of the sigmoid shape how classifiers behave. First, its range: outputs approach 0 and 1 but never reach them, so a logistic model never claims total certainty. Second, symmetry: σ(−z)=1−σ(z)\sigma(-z) = 1 - \sigma(z), so flipping the sign of the score swaps the two class probabilities, and the two classes are treated alike. Third, saturation: for scores beyond about plus or minus 5, the curve is almost flat and its gradient is near zero. Saturation is harmless with cross-entropy loss, which cancels it, but with squared-error loss it stalls learning on badly wrong examples, which is one reason logistic regression does not use MSE.

Mathematical Intuition

The symmetry σ(−z)=1−σ(z)\sigma(-z) = 1 - \sigma(z) means the sigmoid is symmetric about the point (0,0.5)(0, 0.5). This implies that if the linear score zz flips sign, the predicted probability for class 1 and class 0 swap. Saturation at the extremes (σ(z)≈0\sigma(z) \approx 0 for z≪0z \ll 0 and σ(z)≈1\sigma(z) \approx 1 for z≫0z \gg 0) means the function is approximately linear only in a narrow band around z=0z = 0 where σ(z)≈0.5+0.25z\sigma(z) \approx 0.5 + 0.25z. Outside this band, changes in zz barely affect the output.

Example:

If σ(z) = 0.73, what is σ(-z)?

4 of 6
Decision Threshold

y^=1 if σ(z)≥t, else 0\hat{y} = 1 \text{ if } \sigma(z) \geq t, \text{ else } 0

A classifier's probability is not yet a decision. The threshold tt converts it: predict class 1 when the probability is at least tt, otherwise class 0. Because the sigmoid is monotonic, this is the same as comparing the raw score zz with the logit of tt, which is log⁡(t/(1−t))\log(t/(1-t)). The default t=0.5t = 0.5 corresponds to z=0z = 0 and minimizes error count only when both mistakes cost the same and probabilities are calibrated. Lowering tt labels more examples positive, raising recall and lowering precision, and raising it does the reverse. The threshold is chosen after training, on validation data, without retraining the model.

Mathematical Intuition

The decision rule y^=1\hat{y} = 1 if σ(z)≥t\sigma(z) \geq t is equivalent to y^=1\hat{y} = 1 if z≥log⁡t1−tz \geq \log\frac{t}{1-t} (the logit of tt). At t=0.5t = 0.5, the boundary is z=0z = 0. Lowering tt to 0.1 shifts the boundary to z=log⁡(1/9)≈−2.2z = \log(1/9) \approx -2.2, dramatically expanding the positive prediction region. The optimal threshold minimizes expected cost: t∗=CFPCFP+CFNt^* = \frac{C_{FP}}{C_{FP} + C_{FN}}, where CFPC_{FP} and CFNC_{FN} are the costs of false positives and false negatives respectively.

Example:

Probabilities: [0.3, 0.6, 0.45, 0.8]. Predictions at t=0.5? At t=0.4?

5 of 6
Linear Decision Boundary

z=wTx+b=0z = \mathbf{w}^T \mathbf{x} + b = 0

The set of points where the classifier is undecided is its decision boundary. For a linear score z=wTx+bz = \mathbf{w}^T \mathbf{x} + b, the boundary is where z=0z = 0, which is a line in two dimensions, a plane in three, and a hyperplane in general. The weight vector w\mathbf{w} is perpendicular to it and points toward class 1. The value of zz divided by the length of w\mathbf{w} is the signed distance to the boundary, so points far from it get confident probabilities. The limitation is that a single hyperplane cannot separate classes arranged in rings or in XOR patterns, which motivates the feature engineering at the end of this chapter.

Mathematical Intuition

The decision boundary {x:wTx+b=0}\{\mathbf{x} : \mathbf{w}^T\mathbf{x} + b = 0\} is a hyperplane in Rp\mathbb{R}^p with normal vector w\mathbf{w}. The signed distance from any point x0\mathbf{x}_0 to this hyperplane is wTx0+b∥w∥\frac{\mathbf{w}^T \mathbf{x}_0 + b}{\|\mathbf{w}\|}, which is proportional to the log-odds. Points farther from the boundary have higher confidence. The margin (distance between the nearest points of each class and the boundary) determines how robust the classifier is to small perturbations in the input.

Example:

Boundary: 2x₁ + 3x₂ - 6 = 0. Is point (1, 2) class 0 or 1?

6 of 6
Threshold Tuning

The right threshold comes from the cost of each mistake, not from habit. If a false positive costs CFPC_{FP} and a false negative costs CFNC_{FN}, and the probabilities are calibrated, predicting positive is cheaper on average whenever the probability exceeds CFP/(CFP+CFN)C_{FP}/(C_{FP} + C_{FN}). Equal costs give 0.5. In cancer screening a miss is far worse than a false alarm, so the threshold drops sharply. For a spam filter, losing a real email is worse, so it rises. When costs are hard to price, teams pick the threshold that meets a recall or precision target on validation data. This only works if the model's probabilities are calibrated.

Mathematical Intuition

The ROC curve plots true positive rate TPR=TP/(TP+FN)TPR = TP/(TP+FN) against false positive rate FPR=FP/(FP+TN)FPR = FP/(FP+TN) as the threshold varies from 0 to 1. The area under this curve (AUC) measures discrimination ability independent of threshold. The precision-recall curve is more informative for imbalanced data: precision =TP/(TP+FP)= TP/(TP+FP) and recall =TP/(TP+FN)= TP/(TP+FN) directly reflect performance on the minority class. The F1F_1 score =2⋅precision⋅recallprecision+recall= 2 \cdot \frac{\text{precision} \cdot \text{recall}}{\text{precision} + \text{recall}} summarizes the tradeoff as a single number.

Example:

Cancer detection: missing cancer costs 1,000,000 dollars, a false alarm costs 1,000 dollars. How to set threshold?

Theory Exercise

Problem:

A medical test has P(positive|disease) = 0.95 and P(positive|healthy) = 0.05. If 1% of people have the disease, what is P(disease|positive)?

Hints:
  • This is Bayes' theorem
  • P(disease) = 0.01, P(healthy) = 0.99
  • Calculate P(positive) using total probability

Coding Exercise

Problem:

Implement the sigmoid function in numpy, fit a LogisticRegression on the breast cancer dataset, then sweep the decision threshold from 0.5 down to 0.3 and report how precision and recall trade off.

Hints:
  • sigmoid(z) = 1 / (1 + np.exp(-z)); verify sigmoid(0)==0.5.
  • Use clf.predict_proba(X_test)[:, 1] to get positive-class probabilities, then apply (probs >= t).astype(int) for each threshold t.
  • Use precision_score and recall_score from sklearn.metrics and observe that lowering the threshold raises recall but lowers precision.