Chapter 1: Introduction to ML
Discover what makes machine learning transformative: how computers learn from data instead of explicit programming. Master the three learning paradigms (supervised, unsupervised, reinforcement), understand the complete ML pipeline from data to deployment, and grasp the fundamental bias-variance tradeoff that governs all model selection.
Chapter Overview
Machine Learning (ML) is a subset of artificial intelligence that enables computers to learn from data without being explicitly programmed. Instead of writing rules like "if email contains 'free money', mark as spam," we provide examples of spam and non-spam emails, and let algorithms discover the patterns that distinguish them.
The power of ML lies in its ability to handle complexity that would be impossible to encode manually. Consider recognizing faces in photos—there's no set of rules you could write to capture all the variations in lighting, angle, expression, and age. But given thousands of examples, ML algorithms can learn to perform this task with superhuman accuracy.
ML represents a fundamental shift in how we build intelligent systems. Traditional programming requires developers to anticipate every scenario. ML systems, by contrast, improve automatically as they see more data, making them ideal for domains where patterns are complex, evolving, or difficult to articulate.
The field has evolved rapidly—from perceptrons in the 1950s to deep learning systems that now power self-driving cars, language translation, medical diagnosis, and scientific discovery. Understanding ML fundamentals positions you to work with these transformative technologies.
This chapter introduces the foundational concepts you'll use throughout your ML journey:
- What is ML? Understanding the paradigm shift from rule-based to data-driven systems
- Types of Learning: Supervised, unsupervised, semi-supervised, self-supervised, and reinforcement learning paradigms
- Data in ML: Understanding data types, representations, and the importance of data quality
- The ML Pipeline: The end-to-end workflow from raw data collection through model deployment and monitoring
- Bias-Variance Tradeoff: The fundamental tension between model simplicity and complexity that determines generalization
- No Free Lunch Theorem: Why no single algorithm works best for all problems
Chapter Roadmap
Click any topic to jump in
What Is ML?
The paradigm shift from rule-based to data-driven systems — when and why ML outperforms manual programming.
How machines learn and what they learn from
Types of Learning
Supervised, unsupervised, semi-supervised, self-supervised, and reinforcement learning paradigms.
Data in ML
Data types, quality issues, feature engineering, and the train/test split — where 80% of ML work happens.
ML Pipeline
The end-to-end workflow: data collection, preprocessing, training, evaluation, deployment, and monitoring.
The tradeoffs that govern model selection
Bias-Variance Tradeoff
The fundamental tension between underfitting and overfitting that governs all model selection decisions.
No Free Lunch
No single algorithm dominates all problems — match algorithm assumptions to your specific data structure.
Try writing a program that recognizes a cat in a photo. You might start with rules for pointed ears and whiskers, then discover cats photographed from behind, cats in the dark, and cartoon cats, and the rule list never ends. The same wall appears for spam filtering, speech recognition, and fraud detection: the pattern is real, but nobody can write it down as explicit instructions.
The previous chapter, Mathematical Foundations, built the vectors, probability, and optimization tools this plan relies on. This chapter, Introduction to ML, explains what those tools are for, and this first topic defines machine learning itself.
We start by contrasting traditional programming, where people write the rules, with machine learning, where an algorithm infers the rules from examples. We then cover when ML is the right tool and, just as important, when it is not. Next comes the core vocabulary of features, labels, models, training, and inference. We then look at samples drawn from a population distribution, the reason models can fail after deployment, and finish with the split between parametric and non-parametric models.
Definition
Machine learning is the study of algorithms that improve at a task through experience: given example data, they fit the parameters of a model so that it maps inputs to useful outputs, and the fitted model then generalizes to new inputs drawn from the same distribution, without anyone writing the decision rules by hand.
In this topic
Traditional Programming vs ML
Some tasks resist explicit rules, and machine learning changes who writes them. In traditional programming, a person supplies rules and data, and the program produces outputs. In machine learning, we supply data together with the desired outputs, and a training algorithm produces the rules, stored as a model's learned parameters. The model then runs like ordinary code on new inputs. This shifts effort from writing logic to collecting good examples and choosing a model family. The tradeoff is that learned rules are statistical: they are usually right rather than always right, and they are often harder to inspect than hand-written code.
Traditional programming maps , while ML inverts this to learn . The learned function minimizes empirical risk over the training set. The key insight is that is parameterized (e.g., by weights ), and learning means adjusting to minimize this loss — converting the "find the rules" problem into a continuous optimization problem that calculus can solve.
How would you detect cats in photos with traditional programming vs ML?
When to Use ML
Machine learning costs data collection, training, and monitoring, so it should earn its place. It pays off in four situations. First, the rules are too complex to write, as in vision and language. Second, the patterns change over time, such as shopping preferences, so hand-written rules go stale. Third, there is plenty of data but no clear theory linking inputs to outputs. Fourth, the system must personalize for millions of users, where per-user rules are impossible. Each situation asks for patterns that are learned rather than specified. Without enough representative data, though, none of these justifies ML, because the model can only learn what the examples show.
ML is appropriate when the mapping exists but is too complex to specify manually. The formal criterion: the problem has a pattern ( exists), it cannot be pinned down mathematically (no closed-form), and data is available (). The learning bound shows that the gap between true and empirical risk shrinks as data grows — more data makes ML more reliable, while rule-based systems do not improve with more data.
Netflix recommendation system: rules or ML?
When NOT to Use ML
Reaching for ML by default adds cost and risk where a simpler tool works better. Skip it when explicit rules already solve the problem exactly, such as tax brackets or unit conversions. Skip it when you cannot collect enough representative, correctly labeled data. Be cautious when every decision must be explained and reproduced exactly, for example under regulation, because learned models are statistical and harder to audit. Where errors are very costly, such as medical dosing, use ML only with human review or hard safety rules around it. A useful test: if you can write a short, correct specification of the answer, write code, not a model.
ML adds unnecessary complexity when the problem has a known exact solution. For deterministic rules like tax calculation, can be coded directly with zero error. ML would introduce approximation error by construction, since it learns from finite samples. The bias-variance decomposition shows that ML always has non-zero error — only justified when the alternative (manual rules) has even higher error.
Should a tax calculator use ML?
ML Terminology
A shared vocabulary makes every later topic easier to follow. Features, written , are the input variables the model sees; each row of is one example. The label, written , is the value we want to predict. A model is a function with parameters, chosen so that approximates . Training is the process of fitting those parameters to labeled examples, usually by minimizing a loss. Inference is applying the trained model to new examples whose labels are unknown. A common mistake is to treat training performance as the goal, when the real goal is accuracy at inference time on data the model has never seen.
The ML setup: given dataset with features and labels , learn model parameterized by . Training minimizes . Inference computes for unseen data. The number of parameters relative to training samples controls model capacity — when , the model can memorize rather than generalize.
House price prediction: identify features, labels, and what inference means
Sample, Population, Distribution
A model never sees the whole world, only a sample of it, and that gap explains many real failures. The population is every example the model might ever face, generated by an unknown probability distribution. The training set is a finite sample from it. Learning works only if the sample represents that distribution, so the patterns found in training also hold at deployment. When the sample is skewed, such as one region, one season, or one camera, the model can score well in testing and still fail in production. This is called distribution shift, and checking that training data matches deployment data is a basic first step in any project.
We observe a sample drawn i.i.d. from an unknown population distribution . ML approximates the true conditional using only the sample. The generalization error depends on , not just . Distribution shift occurs when the training distribution differs from deployment , violating the core assumption. Domain adaptation techniques attempt to learn representations that are invariant to the shift.
Model trained on US customers, deployed globally. What could go wrong?
Parametric vs Non-Parametric
Model families differ in what they keep after training, and that drives memory, speed, and flexibility. A parametric model summarizes the data in a fixed number of parameters chosen in advance; linear regression with features keeps numbers no matter how many examples it saw. It is fast and compact but limited by its assumed shape. A non-parametric model lets its complexity grow with the data: k-nearest neighbors stores every training example, and decision trees add splits as data grows. These adapt to almost any pattern, but they cost more memory and prediction time and need more data to avoid overfitting.
Parametric models have a fixed number of parameters regardless of data size: linear regression has parameters whether trained on 100 or 1M samples. Non-parametric models grow with data: k-NN stores all training points, costing per prediction. The bias-variance tradeoff manifests differently: parametric models may underfit if the fixed form is wrong (high bias), while non-parametric models may overfit with small (high variance). Kernel methods and Gaussian processes are non-parametric models with elegant theoretical guarantees.
Linear regression with 10 features: how many parameters? k-NN with 1M samples?
Theory Exercise
Problem:
A bank wants to detect fraudulent transactions. Why is ML more suitable than a rule-based system? What challenges might you face?
Hints:
- Think about how fraud patterns evolve
- Consider the class imbalance problem
- Think about false positives vs false negatives
Coding Exercise
Problem:
On the breast cancer dataset, compare a hand-written threshold rule (predict 'malignant' when 'worst radius' exceeds its training median) against a fitted LogisticRegression. Report both test accuracies to show the learned model generalizes better than the manual rule.
Hints:
- Load the data with load_breast_cancer() and split into train/test with a fixed random_state.
- For the rule, find the index of the 'worst radius' feature, compute its median on the TRAIN set only, and threshold the test column.
- Fit LogisticRegression(max_iter=5000, random_state=42) on the training data and compare accuracy_score on the test set for both approaches.
Related Problems on PixelBank
Suppose you have a million X-ray images. If a radiologist has labeled every one, you can train a model to copy those labels. If only a hundred are labeled, or none at all, or the task is to make decisions whose outcome only shows up later, the same images call for a completely different training setup. The kind of feedback available decides how a model can learn.
The previous topic, What is ML?, defined machine learning as learning rules from data. This topic sorts the main paradigms by the feedback they use, which is usually the first decision in designing any ML system.
We start with supervised learning, where every example comes with a label. Unsupervised learning finds structure such as clusters with no labels at all. Semi-supervised learning combines a few labels with many unlabeled examples. Self-supervised learning creates its own labels from raw data, the recipe behind BERT and modern foundation models. Reinforcement learning learns from rewards earned by acting in an environment. We finish with transfer learning, which reuses knowledge from one task to learn another from far less data.
Definition
Learning paradigms are classified by the feedback a model receives: supervised learning uses an input-output label for every example, unsupervised learning uses inputs only, semi-supervised learning mixes a few labels with many unlabeled inputs, self-supervised learning derives labels from the data itself, and reinforcement learning uses reward signals from interaction.
In this topic
Supervised Learning
Learn from labeled examples
When we know the right answer for many past examples, the most direct approach is to learn to reproduce it. Supervised learning fits a function from inputs to outputs using labeled pairs, by minimizing a loss that measures disagreement between predictions and labels. If is a category, such as spam or not spam, the task is classification; if is a number, such as a price, it is regression. Supervised learning is the most reliable paradigm when labels are good, but labels are often expensive, slow, or require experts, and noisy or biased labels are learned just as faithfully as correct ones.
Supervised learning minimizes over labeled pairs . For classification with classes, the softmax output converts logits to probabilities, and cross-entropy loss measures prediction quality. The sample complexity for learning with error and confidence scales as — more features require proportionally more labeled data.
Email spam detection: what are X, y? Classification or regression?
Unsupervised Learning
Find structure in without labels
Most data has no labels, yet it still has structure worth finding. Unsupervised learning works with inputs alone and looks for that structure. Clustering, such as k-means or DBSCAN, groups similar examples. Dimensionality reduction, such as PCA from the previous chapter or t-SNE, compresses many features into a few. Anomaly detection flags examples unlike the rest, and density estimation models the data distribution itself. Because there is no correct answer to compare with, results are harder to evaluate: clusters can reflect irrelevant factors such as feature scale, so they need checking by people who know the domain and by tests on downstream tasks.
Unsupervised learning discovers structure in without labels by optimizing objectives like clustering ( for k-means) or dimensionality reduction ( s.t. for PCA). The fundamental challenge is that "structure" is ill-defined without labels — different objectives reveal different patterns. Information-theoretic approaches maximize mutual information between data and learned representation, while generative approaches model the data distribution directly.
Given 100K customer purchase histories with no labels, what can unsupervised learning do?
Semi-Supervised Learning
Learn from few labels + many unlabeled
Labels are often the bottleneck: unlabeled X-rays are plentiful, but each label needs a radiologist. Semi-supervised learning uses a small labeled set together with a large unlabeled set. Pseudo-labeling trains on the labeled examples, predicts the unlabeled ones, and adds high-confidence predictions as new training labels. Consistency regularization requires the model to give the same prediction for different augmented versions of one unlabeled input, which smooths its decision boundary. Both rely on the assumption that nearby or similar inputs share a label. When the model's early mistakes are confident, pseudo-labeling reinforces them, so confidence thresholds and a clean validation set are essential.
Semi-supervised learning leverages labeled and unlabeled examples. The smoothness assumption states that if and are close in input space, their labels should also be close: . Consistency regularization enforces — predictions should be stable under data augmentation. Pseudo-labeling assigns to unlabeled points when the model is confident (), then trains on these as if they were true labels.
100 labeled X-rays (500 dollars each to label), 10,000 unlabeled. How to use all data?
Self-Supervised Learning
Create labels from the data itself
Human labels cannot scale to billions of examples, but raw data can supervise itself. Self-supervised learning creates a prediction task from the structure of unlabeled data: hide part of the input and predict it from the rest. BERT masks words and predicts them, GPT predicts the next token, and contrastive methods such as SimCLR learn that two augmented views of one image belong together. The learned representations are then fine-tuned for downstream tasks with few labels. This is how modern foundation models are pretrained. The pretext task must require real understanding; one solvable by a shortcut, such as a compression artifact, teaches nothing useful.
Self-supervised learning constructs surrogate labels from the data itself, avoiding the cost of human annotation. The pretext task defines a label function applied to the input: masked language modeling sets = masked tokens, contrastive learning sets = whether two augmented views came from the same image. The key insight is that solving the pretext task forces the model to learn useful representations. Formally, if the pretext task requires capturing mutual information between views, the learned representation must encode the shared semantic content.
How does BERT create labels from raw text?
Reinforcement Learning
Learn policy to maximize reward
Some tasks have no labeled correct action, only outcomes that arrive later, such as winning a game. Reinforcement learning handles this. An agent observes a state , picks an action using its policy , receives a reward, and moves to a new state. Its goal is to maximize total reward over time, so it must link delayed rewards back to earlier actions and balance exploring new actions against exploiting known good ones. A value function estimates how much future reward a state promises. RL powers game agents and robotics, but it needs many interactions and is sensitive to reward design: a badly specified reward gets exploited.
RL maximizes the expected cumulative reward where is the discount factor. The Bellman equation recursively defines the value of taking action in state . The exploration-exploitation tradeoff is fundamental: the agent must try new actions (explore) to discover better strategies while also using known good actions (exploit). -greedy chooses randomly with probability and greedily otherwise.
Chess AI: define state, action, reward
Transfer Learning
Adapt knowledge from source to target domain
Training a large model from scratch needs data and compute that most teams lack, but much of what a model learns is reusable. Transfer learning starts from a model pretrained on a large source task and adapts it to a smaller target task. Early layers of an image network learn edges and textures that are useful almost everywhere, while later layers become task-specific. You can train only a new output head on frozen layers, or fine-tune everything with a small learning rate. It works best when source and target are related; transferring from everyday photos to very different data, such as spectrograms, helps less and sometimes not at all.
Transfer learning leverages a model pre-trained on source task to improve performance on target task . The hypothesis is that low-level features (edges, textures in vision; syntax, word associations in NLP) are universal across tasks. Formally, if the feature extractor learned on produces representations useful for , then fine-tuning only the classifier head on data achieves . The amount of fine-tuning data needed drops dramatically — from millions to hundreds of samples.
You have 500 medical images. Train from scratch or transfer?
Theory Exercise
Problem:
You have 100 labeled images and 10,000 unlabeled images for medical diagnosis. What learning approaches could you use?
Hints:
- Consider combining labeled and unlabeled data
- Think about pre-trained models
- Self-supervised pretraining might help
Coding Exercise
Problem:
On the Iris dataset, train a supervised classifier (LogisticRegression) using the labels and separately run unsupervised KMeans (k=3) WITHOUT labels. Show that supervised learning recovers the classes far more accurately, while clustering finds structure but doesn't know the true class identities.
Hints:
- Use load_iris(); keep X and y. Split X/y for the supervised model with a fixed random_state.
- Supervised: fit LogisticRegression(max_iter=1000) and score test accuracy with accuracy_score.
- Unsupervised: fit KMeans(n_clusters=3, n_init=10, random_state=0) on X (ignoring y) and measure how well clusters match true labels with adjusted_rand_score.
Related Problems on PixelBank
Two teams train the same algorithm on the same task. One reaches 95 percent accuracy and the other 70 percent, and the difference is entirely in the data: one has clean, representative, correctly typed examples, while the other has duplicates, wrong labels, and test rows that leaked into training. In practice the data usually decides results more than the choice of algorithm does.
The previous topic, Types of Learning, sorted paradigms by the kind of feedback available. Whatever the paradigm, everything a model knows comes from its data, so this topic looks closely at what data is and how it goes wrong.
We start with data types, numerical and categorical, ordinal and nominal, which decide how each column must be encoded. We then contrast structured tables with unstructured images and text. Next we separate features from labels and see why the choice of features matters so much. We catalog common data quality issues, from missing values to label noise. We then split data into training, validation, and test sets, and finish with the IID assumption that makes those splits trustworthy.
Definition
In machine learning, data is a collection of examples, each described by features of specific types, numerical, categorical, text, image, or sequence, and in supervised settings paired with a label. Its quality, quantity, representativeness, and the independence of its train, validation, and test splits bound how well any model trained on it can generalize.
In this topic
Data Types
Models only understand numbers, and how a column becomes numbers depends on its type. Numerical data is either continuous, like a temperature of 23.5, or discrete counts, like 3 bedrooms. Categorical data is either nominal, with no order, like colors, or ordinal, with a natural order, like small, medium, and large. Text, images, and time series are sequences or grids with their own structure. The type decides the encoding: nominal categories need one-hot encoding, ordinal ones can map to ordered integers, and text needs tokenization. Treating a nominal category as a number, such as red as 1 and blue as 2, invents an order that does not exist.
Different data types require different mathematical representations. Numerical features can be directly used in computations. Categorical features with categories map to one-hot vectors , expanding the feature space by dimensions. Ordinal features preserve order: . The choice of encoding affects model behavior — one-hot encoding allows each category to have independent effects, while ordinal encoding assumes evenly spaced categories (which may not hold).
Classify: temperature, shirt size (S/M/L), color (red/blue), customer review text
Structured vs Unstructured
Different kinds of data reward different model families, so knowing which kind you have narrows the choices early. Structured data lives in tables with defined columns, each holding one typed attribute, as in SQL databases and CSV files. Unstructured data, such as images, audio, and free text, has no fixed columns; meaning lives in patterns across many raw values, such as pixels or words. Gradient-boosted trees and linear models do very well on structured data, while deep learning dominates unstructured data because it learns its own features. Semi-structured formats such as JSON logs sit in between and usually need parsing into columns first.
Structured data fits in tables with fixed schemas — each row has the same features in the same order. Unstructured data (images, text, audio) has variable length and spatial/sequential structure. The representation gap means different architectures excel: gradient boosting (XGBoost, LightGBM) dominates on tabular data because tree splits align naturally with feature boundaries, while neural networks dominate on unstructured data because convolutions (images) and attention (text) exploit spatial and sequential patterns.
Which is structured: customer database, Instagram photos, financial transactions?
Feature vs Label
(features) → (label/target)
Every supervised problem starts by deciding what the model sees and what it must predict. Features are the inputs available at prediction time; the label is the target. Choosing features well usually matters more than choosing an algorithm: good features make the pattern easy to learn, and missing features make it impossible. Feature engineering builds new features from raw ones, such as ratios or date parts, adding domain knowledge. The most damaging mistake is target leakage, a feature that is only known after the label is decided. It inflates offline accuracy and then fails in production, where that information is not yet available.
The feature matrix has samples and features. The label vector (regression) or (classification) contains what we predict. Feature engineering creates new informative features from raw ones: interaction terms , polynomial features , or domain-specific transformations. The information bottleneck principle suggests that good features should maximize (predictive power) while minimizing (compression).
Predict loan default. What features? What label?
Data Quality Issues
Real datasets are messy, and a model trained on them learns the mess along with the signal. Common problems include missing values, impossible or extreme outliers, duplicate rows, inconsistent formats such as mixed units or spellings, label noise where the target itself is wrong, and class imbalance where one outcome is rare. Each fails differently: duplicates that cross the train-test split inflate scores, placeholder values such as 999 distort averages, and label noise caps achievable accuracy. Cleaning often takes most of a project's time. Fix the source where possible, and document every cleaning rule so the same steps run identically at inference.
Missing data reduces the effective sample size and can introduce bias if the missingness pattern is not random. Mean imputation underestimates variance. MICE (Multiple Imputation by Chained Equations) models each feature's conditional distribution and samples from it, preserving the uncertainty. Outliers with leverage (where is the diagonal of the hat matrix ) can disproportionately influence the model — robust methods like Huber loss limit their impact.
Age column has: [25, 30, -5, 999, NaN, 'thirty']. What issues?
Train/Validation/Test Split
Typically 60/20/20 or 70/15/15
If we judge a model on the data it trained on, we measure memory, not generalization, so data is split three ways. The training set fits the parameters. The validation set compares models and tunes hyperparameters such as learning rate or tree depth. The test set is held back until the end and used once to estimate real-world performance. Every decision made after looking at test results quietly fits the test set, making the final number optimistic. Leakage between splits, such as duplicates or the same patient in two sets, has the same effect, so split by group when examples are related.
The three-way split serves distinct purposes: training estimates by minimizing , validation selects hyperparameters by minimizing , and test estimates true risk exactly once. The validation set prevents "meta-overfitting" to hyperparameters — if we tuned on the test set, our test error would be optimistically biased. With total samples, the bias-variance tradeoff of the split itself matters: too much test data wastes training signal, too little gives noisy estimates.
10,000 samples. How many in each set for 70/15/15 split?
IID Assumption
Independent and Identically Distributed
Random splits and test scores are trustworthy only under an assumption that is easy to forget. IID means independent and identically distributed: each example is drawn independently of the others, and all come from the same distribution. Under IID, a random test set behaves like future data, so test accuracy predicts deployed accuracy. Time series break independence, because today's value depends on yesterday's. Distribution shift breaks the identical part, as when training data comes from one hospital and deployment from another. When IID fails, a random split leaks future information into training and overstates performance, so split by time or by group.
The i.i.d. assumption independently means each sample provides fresh information: the effective sample size equals . Violation by temporal correlation means consecutive samples carry redundant information — the effective sample size is where is the autocorrelation. For time series, this means a random train/test split leaks future information. The correct approach is temporal splitting: train on and test on , preserving the causal structure.
Stock prices: is tomorrow's price independent of today's?
Theory Exercise
Problem:
Your dataset has 1000 samples: 950 negative and 50 positive examples. What problems might arise and how would you address them?
Hints:
- This is a class imbalance problem
- Think about what accuracy would mean here
- Consider sampling strategies
Coding Exercise
Problem:
Demonstrate data leakage. With 100 samples, 2000 random features, and RANDOM binary labels (true accuracy should be ~0.5), select the top 20 features using ALL the data before cross-validation, then compare to selecting features INSIDE a Pipeline (per fold). Show the leaky version reports an impossibly high score.
Hints:
- Generate X = rng.randn(100, 2000) and y = random 0/1 labels with a fixed RandomState — there is genuinely no signal, so honest accuracy must be ~0.5.
- Leaky way: run SelectKBest(f_classif, k=20).fit_transform(X, y) on the FULL data, then cross_val_score the classifier on the reduced data.
- Honest way: put SelectKBest and LogisticRegression in make_pipeline and pass the RAW X to cross_val_score so selection happens separately on each training fold.
Related Problems on PixelBank
A model that scores well in a notebook is not yet a product. Between raw data and a dependable prediction service sit many steps, and most real failures happen in the unglamorous ones: a preprocessing step fit on the full dataset, a comparison run on different splits, a model that quietly decays months after launch.
The previous topic, Data in ML, covered what data looks like and how it goes wrong. This topic arranges the full workflow around that data into a repeatable pipeline, the sequence a practitioner follows on every project.
We start with data collection and exploratory analysis, getting to know the data before modeling. Preprocessing handles missing values, encodings, and scaling. Feature engineering adds domain knowledge as new inputs. The train-test split, done before fitting anything, protects the evaluation, and cross-validation makes that evaluation more reliable. Model selection and training then compare candidates fairly. Evaluation and error analysis find where the model fails and why. We finish with deployment and monitoring, because data drifts and a model that is never watched slowly gets worse.
Definition
An ML pipeline is the end-to-end, repeatable sequence of steps that turns raw data into a deployed model: collection and exploration, preprocessing, feature engineering, splitting, training with validation, evaluation, deployment, and monitoring. Every transformation is fit on training data only and applied identically at inference time.
In this topic
Data Collection & Exploration
Modeling data you have not looked at is guesswork, so the pipeline starts by collecting and exploring. Collection means gathering examples that represent the deployment population, with labels you trust. Exploratory data analysis, or EDA, then checks size, column types, missing values, summary statistics, distributions, and correlations. In pandas this is shape, info, describe, and corr, plus histograms and scatter plots. EDA catches problems cheaply: a column that is 90 percent missing, a target that is almost all one class, impossible values, or a feature suspiciously correlated with the label, which often signals leakage. Do the exploration on training data so the test set stays untouched.
Exploratory Data Analysis (EDA) characterizes the dataset through summary statistics (, , quantiles), distribution shapes (skewness, kurtosis), and pairwise relationships (correlation matrix ). The correlation matrix reveals multicollinearity — when for , features are redundant and the Gram matrix becomes ill-conditioned, inflating variance in linear models. Detection of these patterns in EDA guides all downstream decisions.
First look at a dataset. What EDA would you do?
Preprocessing
Raw columns rarely suit a model directly, so preprocessing converts them into clean numbers. Missing values are imputed, for example with the median, or rows are dropped when few are affected. Categorical columns are encoded, one-hot for nominal and integers for ordinal. Numeric features are scaled, by standardization to mean 0 and standard deviation 1 or by min-max scaling, so gradient-based and distance-based models treat them evenly. Every preprocessing step that learns statistics, such as a median or a mean, must be fit on the training split only and then applied to validation and test. A scikit-learn Pipeline makes this the default.
Standardization ensures all features have zero mean and unit variance. This is critical for gradient-based methods because the loss landscape curvature becomes more uniform — the condition number of the Hessian decreases, enabling a larger learning rate. For one-hot encoding of a categorical feature with categories, we use indicator variables (full rank encoding) to avoid the multicollinearity trap where the indicators sum to 1.
Features: age (10% missing), city (categorical), income (range 20K-500K). Preprocess?
Feature Engineering
Models learn more easily when the inputs already express the relevant pattern, and feature engineering builds those inputs from raw data using domain knowledge. A timestamp can become hour of day, day of week, and a holiday flag. Two columns can combine into a ratio, such as debt to income. Text can become TF-IDF vectors, and numeric features can be crossed into interaction or polynomial terms. Good features let simple models compete with complex ones and make results easier to explain. Every engineered feature must be computable at prediction time from information available then, or it leaks future data.
Feature engineering constructs from raw features to make patterns explicit. Polynomial features let linear models fit curves. Interaction features capture synergistic effects that individual features miss. The kernel trick computes dot products in an implicit high-dimensional feature space without explicitly constructing — RBF kernels implicitly use infinite-dimensional feature spaces.
Predicting flight delays. Raw feature: departure_time='2024-01-15 14:30'. Create useful features.
Train-Test Split
Never evaluate on training data!
The test set exists to estimate performance on unseen data, and that estimate is ruined if any information from it reaches training. So the split happens first, before imputation, scaling, feature selection, or EDA-driven decisions that use statistics of the data. For classification, a stratified split keeps class proportions the same in every split, which matters for rare classes. For time-ordered data, split by time so the model never trains on the future. When examples are grouped, such as several images from one patient, split by group so near-duplicates cannot sit on both sides.
Splitting before preprocessing prevents data leakage: if we compute and on all data including test, the standardized test features carry information from the training set. Formally, the test error estimate should be an unbiased estimator of . Data leakage introduces optimistic bias: , making the model appear better than it actually is. Stratified splitting ensures class proportions match: for 5% positive class, both train and test get approximately 5%.
Dataset: 90% class A, 10% class B. Why use stratified split?
Cross-Validation
K-Fold: Train on K-1 folds, validate on 1
A single validation split gives one noisy number, which can mislead when data is limited. K-fold cross-validation divides the training data into equal folds, trains times, each time holding out a different fold for validation and training on the other , and averages the scores. Every example is validated exactly once, and the spread across folds shows how stable the estimate is. Use stratified folds for classification, time-series splits for temporal data, and group folds when examples are related. Cross-validation costs training runs, and the final test set must still stay outside it.
K-fold CV partitions data into equal folds, trains on and validates on the remaining fold times. The CV estimate has lower variance than a single split because each sample appears in exactly one validation fold. Leave-one-out CV () is nearly unbiased but has high variance because the training sets overlap almost completely. The standard choice or balances bias (each fold trains on 80-90% of data) and variance (5 or 10 validation scores to average, though they are correlated because training sets overlap).
5-fold CV with 1000 samples. How many times is each sample used for validation?
Model Selection & Training
Choosing among algorithms is only meaningful if the comparison is fair. Start with a simple baseline, such as predicting the majority class or a logistic regression, so every later model has a number to beat. Then try a few diverse candidates and tune each one's hyperparameters, using grid search, random search, or Bayesian optimization. Use identical cross-validation splits for all candidates and compare mean and spread, not one lucky run. Give each model a similar tuning budget, since an untuned model against a tuned one compares effort, not algorithms. Once a winner is chosen, retrain it on all training data before the final test.
Model selection compares candidates by their validation performance. Grid search evaluates all hyperparameter combinations — exponential in the number of hyperparameters. Random search samples combinations and empirically finds near-optimal settings in fewer evaluations because most hyperparameters have low effective dimensionality (only 1-2 matter). Bayesian optimization models as a Gaussian process and acquires points that maximize expected improvement, efficiently exploring the hyperparameter space.
Comparing RandomForest vs XGBoost. How to ensure fair comparison?
Evaluation & Error Analysis
A single accuracy number hides where a model fails, and those failures point to what to fix next. Evaluation first picks metrics that match the real cost of mistakes: recall when missing a positive is costly, precision when false alarms are, and calibrated probabilities when downstream decisions use confidence. Error analysis then looks directly at the misclassified examples and groups them by attributes such as lighting, device, or customer segment. Errors concentrated in one slice are systematic and usually come from missing data, a missing feature, or label noise in that slice. Random errors spread evenly suggest the model has reached the data's limits.
For classification with classes, the confusion matrix counts instances of true class predicted as class . Precision measures false positive rate, recall measures false negative rate, and is their harmonic mean. The ROC curve plots vs across classification thresholds, and AUC measures the probability that a random positive is scored higher than a random negative.
Model fails on images with poor lighting. What does this suggest?
Deployment & Monitoring
Deployment is not the end of the pipeline, because the world keeps changing after training. The model is served behind an API or run in batch jobs, and the same preprocessing code must run exactly as it did in training. Monitoring then watches for data drift, where the input distribution shifts, such as a new customer region, and for concept drift, where the relationship between inputs and the label changes, such as fraudsters adopting new tactics. Track input statistics, prediction rates, and real accuracy once labels arrive. When metrics degrade, retrain on recent data, on a schedule or triggered by alerts.
In production, the model serves predictions via API with latency constraints. Data drift occurs when — the input distribution shifts. Concept drift occurs when itself changes — the relationship between features and labels evolves. Monitoring tracks KL divergence on feature distributions and performance metrics on labeled samples. Retraining triggers when drift exceeds a threshold or performance degrades below acceptable levels.
Fraud model deployed in 2020. By 2023, fraud rate increases. What happened?
Theory Exercise
Problem:
You fit a StandardScaler on your entire dataset, then split into train/test. Why is this a problem?
Hints:
- What statistics does StandardScaler compute?
- What if test data was from a different time period?
- This is called 'data leakage'
Coding Exercise
Problem:
Build a proper scikit-learn Pipeline that chains StandardScaler with LogisticRegression on the breast cancer dataset, then evaluate it with 5-fold cross_val_score so scaling is re-fit on each training fold. Report the mean and standard deviation of the fold accuracies.
Hints:
- Load X, y with load_breast_cancer(return_X_y=True).
- Chain the steps with make_pipeline(StandardScaler(), LogisticRegression(max_iter=5000, random_state=1)) so the whole pipeline is treated as one estimator.
- Pass the pipeline and raw X, y to cross_val_score(..., cv=5) and report scores.mean() and scores.std().
Related Problems on PixelBank
A model can fail in two opposite ways. It can be too rigid to capture the real pattern, scoring poorly everywhere, or so flexible that it memorizes the noise in its training data, scoring brilliantly in training and poorly on anything new. Diagnosing which failure you have decides the fix, and the two fixes point in opposite directions.
The previous topic, ML Pipeline, built the workflow from data to deployment, including the validation and cross-validation that measure generalization. This topic explains what those measurements reveal about a model's errors.
We start with bias, the error from assumptions too simple for the data, which shows up as underfitting. Variance is the error from sensitivity to the particular training sample, which shows up as overfitting. The total error decomposition shows that expected test error splits into squared bias, variance, and irreducible noise. Model complexity is the dial that trades one against the other. Regularization offers a finer-grained way to turn that dial, and learning curves show which error dominates and whether collecting more data would help.
Definition
The bias-variance tradeoff describes how a model's expected squared prediction error on new data splits into bias squared, error from overly simple assumptions; variance, error from sensitivity to the particular training sample; and irreducible noise. Increasing model complexity usually lowers bias and raises variance, so good generalization requires balancing the two.
In this topic
Bias (Underfitting)
Error from wrong assumptions
Some models fail because their assumptions cannot express the true pattern, however much data they get. That systematic error is bias. Formally, bias is the gap between the model's average prediction, over many possible training sets, and the true value. A straight line fit to a curved relationship misses in the same way every time. The symptom is high error on both training and validation data, with the two close together; more data does not help, because the model already uses all it can represent. Fixes add capacity: more features, nonlinear models, or less regularization.
Bias measures the systematic error from wrong model assumptions: , where the expectation is over different training sets. A linear model fitting a quadratic function has high bias because no matter how much data it sees, it can never capture the curvature. Bias is a property of the model class, not the specific training run. Reducing bias requires increasing model capacity — adding polynomial features, deeper networks, or switching to a more flexible model family.
Linear model on y=x². Train error: 30%, Test error: 32%. Diagnosis?
Variance (Overfitting)
Error from sensitivity to training data
Other models fail because they fit their particular training sample too closely, noise included. That sensitivity is variance: how much the model's predictions change when it is trained on a different sample from the same distribution. A very deep decision tree can carve out a region for every training point, reaching zero training error, but a fresh sample produces a very different tree. The symptom is a large gap between low training error and much higher validation error. Fixes reduce effective flexibility or add information: limit depth or prune, regularize, use ensembles that average many models, or gather more training data.
Variance measures sensitivity to the training data: . High variance means different training sets produce wildly different models. A degree-20 polynomial fit to 15 points will oscillate dramatically between data points, and a slightly different sample gives a completely different curve. Variance increases with model complexity and decreases with training set size. The relationship is approximately .
Decision tree (depth=50). Train error: 0%, Test error: 40%. Diagnosis?
Total Error Decomposition
To fix error, we need to know where it comes from, and for squared error the expected test error splits exactly into three parts. Bias squared is the systematic error of the model's average prediction. Variance is how much predictions scatter across different training sets. Irreducible noise is the randomness in the labels themselves, which no model can predict. The decomposition shows why fixes have side effects: lowering bias by adding complexity usually raises variance, and the reverse. It also sets a floor: once bias and variance are small, error cannot go below the noise, so further model work stops paying.
The expected prediction error decomposes as . The noise term is irreducible — it reflects inherent randomness in the data that no model can eliminate. The optimal model minimizes , which involves a tradeoff: reducing one typically increases the other. This decomposition explains why a model can be "too good" on training data (zero bias, high variance) and why simpler models sometimes generalize better.
If Bias²=4, Variance=5, Noise=1. What is total error? What to reduce?
Model Complexity
Complexity is the main dial between bias and variance. Simple models, such as linear models or shallow trees, make strong assumptions: high bias, low variance. Complex models, such as high-degree polynomials, deep trees, and large networks, can fit almost anything: low bias, high variance. As complexity grows, training error keeps falling, while validation error usually falls and then rises, a U-shape whose bottom is the sweet spot. Where that bottom sits depends on how much data you have and how complex the true pattern is. Very large networks can show a second descent past the point of fitting the training data exactly.
Model complexity can be measured by the number of effective parameters, VC dimension, or Rademacher complexity. For polynomial regression of degree with samples, the effective complexity is roughly . When , the model is constrained (high bias, low variance). When , the model can interpolate the training data exactly (zero training error, high variance). The "double descent" phenomenon shows that (overparameterized regime) can surprisingly reduce test error again, which modern deep learning exploits.
100 samples: polynomial degree 1, 5, and 20. Which is best?
Regularization
Instead of switching model families, we can adjust complexity smoothly by penalizing it. Regularization adds a penalty to the training loss . With the L1 norm of the weights, as in Lasso, many weights become exactly zero, a form of feature selection. With the squared L2 norm, as in Ridge or weight decay, all weights shrink toward zero. The strength sets the tradeoff: at zero the model fits freely, with low bias and high variance; as grows, variance falls and bias rises. Because penalties act on weight sizes, features must be standardized first, and is chosen by cross-validation.
Regularization adds a complexity penalty: . L2 regularization adds to each eigenvalue of the Hessian, improving conditioning and reducing the effective degrees of freedom from to approximately where are Hessian eigenvalues. The regularization path shows how the optimal weights change as varies: at we get MLE, and as , all weights shrink to zero. Cross-validation over finds the optimal bias-variance tradeoff.
Model overfitting. λ=0.001 gives test error 25%. Try λ=0.1 and λ=10?
Learning Curves
Before spending money on more data or a bigger model, we want to know which will help, and learning curves answer that. Plot training and validation error against training set size. With high bias, both curves flatten at a high error close together, so more data will not help and more capacity will. With high variance, training error stays low while validation error sits well above it, and the gap narrows as data grows, so more data or regularization will help. Learning curves cost several training runs at different sizes, and their noise at small sizes makes cross-validated points more reliable.
Learning curves plot error versus training set size . For high bias: both training and validation errors converge to a high value as increases — more data does not help because the model is fundamentally wrong. For high variance: training error is low and validation error is high, but the gap closes with more data because the model can better distinguish signal from noise. The gap between curves at any approximates the variance term, and the level where both curves converge approximates .
Train error: 5% at 1K samples, 5% at 10K samples. Val error: 30% at 1K, 15% at 10K. More data help?
Theory Exercise
Problem:
A model achieves 99% accuracy on training data but only 60% on test data. What is the likely problem and how would you fix it?
Hints:
- Compare train vs test performance
- Think about model complexity
- Consider regularization or more data
Coding Exercise
Problem:
On noisy data generated from a sine curve, fit polynomial regressions of degree 1, 4, and 15 and report training vs test mean squared error for each. Show that low degree underfits (high bias), high degree overfits (high variance), and a middle degree generalizes best.
Hints:
- Make X uniform in [-3, 3] and y = sin(X) + small Gaussian noise using a fixed RandomState; split into train/test.
- For each degree, build make_pipeline(PolynomialFeatures(degree), LinearRegression()) and fit on the training set.
- Compute mean_squared_error on both train and test sets and compare how the gap grows as degree increases.
Related Problems on PixelBank
Every few months a new algorithm is announced as the best: a boosted tree that tops a leaderboard, a neural network that sets a record. It is tempting to adopt the current winner everywhere, but practitioners keep seeing the champion of one benchmark lose to a plain logistic regression on their own data. Is there a best learning algorithm at all?
The previous topic, Bias-Variance Tradeoff, showed that the right model complexity depends on the data. This topic, which closes the Introduction to ML chapter, generalizes that lesson: which algorithm is right depends on the problem itself.
We start with the No Free Lunch theorem, which proves that averaged over all possible problems no algorithm beats any other. We then draw out its implication for practitioners: benchmark candidates on your own data. Occam's razor offers a tie-breaker that favors simpler models. Inductive bias names the assumptions that make each algorithm good on some problems and bad on others. We then lay out a practical model selection strategy, and finish by looking at when deep learning wins and when gradient-boosted trees still do better.
Definition
The No Free Lunch theorem states that, averaged uniformly over all possible target functions, every learning algorithm has the same expected performance on unseen data. An algorithm can only outperform others by making assumptions, its inductive bias, that match the structure of the problems it is actually applied to.
In this topic
No Free Lunch Theorem
It is natural to search for one universally best learning algorithm, and the No Free Lunch theorem shows none can exist. Wolpert proved that, averaged uniformly over every possible target function, all algorithms have the same expected accuracy on points they have not seen. The reason is that, without assumptions, the training data says nothing about the unseen points: for each labeling of them where one algorithm guesses right, another labeling exists where it guesses wrong. An algorithm wins only on problems that match its assumptions. Real-world problems are not uniformly random, which is why learning works at all, but the theorem means every success depends on assumptions.
The NFL theorem states that averaged over all possible target functions , every learning algorithm has the same expected error. Formally, for any algorithms and : . This means any superiority of algorithm on one class of problems is exactly compensated by inferiority on another class. The practical implication is that algorithm selection must be guided by problem structure — there is no universally best algorithm to default to.
Algorithm A beats B on 1000 problems. Does this mean A is always better?
Implication for Practitioners
If no algorithm is best everywhere, then 'which algorithm should I use?' has no universal answer, only an empirical one for your problem. The practical consequences follow from the theorem. Benchmark several algorithms with different assumptions on your own data, using the same validation splits. Treat published leaderboard results as evidence about problems like those benchmarks, not as guarantees for yours. Use domain knowledge to choose candidates whose assumptions fit, such as convolutions for images or monotone constraints for credit scores. The failure mode is cargo-culting: adopting whatever is fashionable without testing whether its assumptions match your data's structure.
NFL means benchmark results on standard datasets do not guarantee performance on your specific problem. The distribution of real-world problems is not uniform over all possible functions — real problems have structure (smoothness, sparsity, hierarchy) that certain algorithms exploit. The practical strategy is to match algorithm inductive biases to known problem properties: smooth functions favor kernel methods, sparse signals favor regularization, hierarchical patterns favor deep networks. Always validate on your own data.
Paper claims 'XGBoost achieves SOTA on all benchmarks'. Should you always use XGBoost?
Occam's Razor
Among models with similar performance, prefer the simpler one
When two models perform about equally, we need a principled tie-breaker, and Occam's razor supplies it: prefer the simpler one. Simpler models have fewer parameters, so they tend to have lower variance and generalize more reliably, as the bias-variance topic showed. They are also faster to train and serve, easier to explain to stakeholders and regulators, and easier to debug when something breaks. Simplicity is a preference among models of similar accuracy, not a rule to accept worse ones. Before calling a difference small, check whether it exceeds the noise in the evaluation, and weigh it against the real cost of the extra errors.
Occam's Razor prefers simpler models among those with similar performance. This has formal backing: MDL (Minimum Description Length) theory shows that the best model minimizes — the data fit plus the model description length. Simpler models have shorter descriptions , so they need stronger data evidence to be displaced by complex models. Bayesian model comparison automatically implements Occam's Razor through the marginal likelihood , which penalizes complex models for spreading probability mass too thinly.
Logistic regression: 89% accuracy. Deep neural net: 90% accuracy. Which to deploy?
Inductive Bias
The No Free Lunch theorem says every algorithm wins only by assuming something, and those assumptions are its inductive bias. Linear models assume the output changes linearly with each feature. Decision trees assume the data splits well along one feature at a time. k-nearest neighbors assumes nearby points share labels. Convolutional networks assume local patterns matter and can appear anywhere in an image, giving translation equivariance through weight sharing. Recurrent networks and transformers assume sequential or contextual structure. Choosing a model is choosing assumptions, so match them to what you know about the domain. A mismatched bias needs far more data to overcome, or never overcomes it.
Every learning algorithm embeds assumptions (inductive biases) about what functions are likely. Linear models assume — the decision boundary is a hyperplane. CNNs assume translation equivariance: a pattern in one image location should be detected everywhere. Transformers assume that relevant context can be selected by attention: . The right inductive bias reduces the hypothesis space from all possible functions to a manageable subset that includes the true function, dramatically reducing sample complexity.
Image classification: why do CNNs work better than linear models?
Model Selection Strategy
Since no algorithm is best everywhere, we need a disciplined process for finding the best one here. First, set a baseline, such as predicting the majority class or the mean, then a simple model such as logistic regression, to know what any real model must beat. Second, try diverse algorithms with different inductive biases, such as linear models, tree ensembles, and neural networks. Third, tune the strongest two or three with an equal budget on the same folds. Fourth, consider an ensemble of the best if it adds meaningful accuracy. Finally, choose by weighing performance against complexity, cost, and interpretability, and confirm once on the held-out test set.
Systematic model selection: (1) Establish a baseline — majority-class classifier gives accuracy = . (2) Try diverse model families to cover different inductive biases. (3) For each, tune hyperparameters via cross-validation. (4) Compare best variants using the same CV folds (paired comparison reduces variance). (5) Ensemble the top models: the ensemble error is bounded by when models are uncorrelated, providing a systematic way to reduce variance.
New classification problem. What's your experiment plan?
When Deep Learning Wins
Deep learning is the default for some problems and a poor choice for others, and the No Free Lunch view explains the difference. Deep networks win on unstructured data, such as images, audio, and text, where they learn features no one could engineer by hand. They also win when data is plentiful, often tens of thousands of examples or more, or when a pretrained model can be fine-tuned. On small to medium structured tables, gradient-boosted trees such as XGBoost and LightGBM usually match or beat them, with less tuning and compute. Grinsztajn and colleagues traced this to tree-friendly properties of tabular data, such as uninformative features and irregular target functions.
Deep learning excels when: (1) data is abundant (), so the high variance of complex models is controlled, (2) data is unstructured (images, text), where learned features outperform handcrafted ones, (3) transfer learning is available, providing strong initialization. For tabular data with samples, gradient boosting typically wins because tree ensembles have lower variance at small and naturally handle mixed feature types. The crossover point depends on feature quality — with strong domain features, simple models suffice even at large .
Two tasks: (A) 1M images, (B) 1K rows of tabular data. Best algorithm for each?
Theory Exercise
Problem:
You have a tabular dataset with 500 samples and 20 features. A colleague insists on using a deep neural network. What would you recommend?
Hints:
- Consider the data size vs model complexity
- Think about what algorithms excel on tabular data
- Remember the bias-variance tradeoff
Coding Exercise
Problem:
Show that no single model is best everywhere. On a linearly-separable dataset and on the make_moons dataset, compare LogisticRegression (linear) vs an RBF-kernel SVM. Demonstrate the linear model wins (or ties) on linear data while the RBF-SVM wins on the curved moons data.
Hints:
- Build two datasets: make_classification with class_sep large for linearly separable data, and make_moons(noise=0.2) for non-linear data — both with a fixed random_state.
- For each dataset, split train/test and fit both LogisticRegression(max_iter=1000) and SVC(kernel='rbf').
- Compare test accuracies per dataset and note which model wins where — the ranking flips between datasets.