PIXELBANKv9.1.0
Menu
Back to ML Study Plan
Week 1-2

Chapter 1: Introduction to ML

Discover what makes machine learning transformative: how computers learn from data instead of explicit programming. Master the three learning paradigms (supervised, unsupervised, reinforcement), understand the complete ML pipeline from data to deployment, and grasp the fundamental bias-variance tradeoff that governs all model selection.

Chapter Overview

Machine Learning (ML) is a subset of artificial intelligence that enables computers to learn from data without being explicitly programmed. Instead of writing rules like "if email contains 'free money', mark as spam," we provide examples of spam and non-spam emails, and let algorithms discover the patterns that distinguish them.

The power of ML lies in its ability to handle complexity that would be impossible to encode manually. Consider recognizing faces in photos—there's no set of rules you could write to capture all the variations in lighting, angle, expression, and age. But given thousands of examples, ML algorithms can learn to perform this task with superhuman accuracy.

ML represents a fundamental shift in how we build intelligent systems. Traditional programming requires developers to anticipate every scenario. ML systems, by contrast, improve automatically as they see more data, making them ideal for domains where patterns are complex, evolving, or difficult to articulate.

The field has evolved rapidly—from perceptrons in the 1950s to deep learning systems that now power self-driving cars, language translation, medical diagnosis, and scientific discovery. Understanding ML fundamentals positions you to work with these transformative technologies.

This chapter introduces the foundational concepts you'll use throughout your ML journey:

  • What is ML? Understanding the paradigm shift from rule-based to data-driven systems
  • Types of Learning: Supervised, unsupervised, semi-supervised, self-supervised, and reinforcement learning paradigms
  • Data in ML: Understanding data types, representations, and the importance of data quality
  • The ML Pipeline: The end-to-end workflow from raw data collection through model deployment and monitoring
  • Bias-Variance Tradeoff: The fundamental tension between model simplicity and complexity that determines generalization
  • No Free Lunch Theorem: Why no single algorithm works best for all problems

Chapter Roadmap

Click any topic to jump in

Explainer video
1
What Is ML?

The paradigm shift from rule-based to data-driven systems — when and why ML outperforms manual programming.

Traditional Programming vs MLWhen to Use MLWhen NOT to Use MLML TerminologySample, Population, DistributionParametric vs Non-Parametric
Learning paradigms and data foundations

How machines learn and what they learn from

2
Types of Learning

Supervised, unsupervised, semi-supervised, self-supervised, and reinforcement learning paradigms.

Supervised LearningUnsupervised LearningSemi-Supervised LearningSelf-Supervised LearningReinforcement LearningTransfer Learning
3
Data in ML

Data types, quality issues, feature engineering, and the train/test split — where 80% of ML work happens.

Data TypesStructured vs UnstructuredFeature vs LabelData Quality IssuesTrain/Validation/Test SplitIID Assumption
Putting it all together
4
ML Pipeline

The end-to-end workflow: data collection, preprocessing, training, evaluation, deployment, and monitoring.

Data Collection & ExplorationPreprocessingFeature EngineeringTrain-Test SplitCross-ValidationModel Selection & TrainingEvaluation & Error AnalysisDeployment & Monitoring
Fundamental principles

The tradeoffs that govern model selection

5
Bias-Variance Tradeoff

The fundamental tension between underfitting and overfitting that governs all model selection decisions.

Bias (Underfitting)Variance (Overfitting)Total Error DecompositionModel ComplexityRegularizationLearning Curves
6
No Free Lunch

No single algorithm dominates all problems — match algorithm assumptions to your specific data structure.

No Free Lunch TheoremImplication for PractitionersOccam's RazorInductive BiasModel Selection StrategyWhen Deep Learning Wins

Try writing a program that recognizes a cat in a photo. You might start with rules for pointed ears and whiskers, then discover cats photographed from behind, cats in the dark, and cartoon cats, and the rule list never ends. The same wall appears for spam filtering, speech recognition, and fraud detection: the pattern is real, but nobody can write it down as explicit instructions.

The previous chapter, Mathematical Foundations, built the vectors, probability, and optimization tools this plan relies on. This chapter, Introduction to ML, explains what those tools are for, and this first topic defines machine learning itself.

We start by contrasting traditional programming, where people write the rules, with machine learning, where an algorithm infers the rules from examples. We then cover when ML is the right tool and, just as important, when it is not. Next comes the core vocabulary of features, labels, models, training, and inference. We then look at samples drawn from a population distribution, the reason models can fail after deployment, and finish with the split between parametric and non-parametric models.

Definition

Machine learning is the study of algorithms that improve at a task through experience: given example data, they fit the parameters of a model so that it maps inputs to useful outputs, and the fitted model then generalizes to new inputs drawn from the same distribution, without anyone writing the decision rules by hand.

In this topic

1Traditional Programming vs ML
2When to Use ML
3When NOT to Use ML
4ML Terminology
5Sample, Population, Distribution
6Parametric vs Non-Parametric
1 of 6
Traditional Programming vs ML

Some tasks resist explicit rules, and machine learning changes who writes them. In traditional programming, a person supplies rules and data, and the program produces outputs. In machine learning, we supply data together with the desired outputs, and a training algorithm produces the rules, stored as a model's learned parameters. The model then runs like ordinary code on new inputs. This shifts effort from writing logic to collecting good examples and choosing a model family. The tradeoff is that learned rules are statistical: they are usually right rather than always right, and they are often harder to inspect than hand-written code.

Mathematical Intuition

Traditional programming maps f:Rules×Data→Outputf: \text{Rules} \times \text{Data} \to \text{Output}, while ML inverts this to learn f:Data×Output→Rulesf: \text{Data} \times \text{Output} \to \text{Rules}. The learned function f^\hat{f} minimizes empirical risk 1n∑i=1nL(f^(xi),yi)\frac{1}{n}\sum_{i=1}^n L(\hat{f}(x_i), y_i) over the training set. The key insight is that f^\hat{f} is parameterized (e.g., by weights θ\theta), and learning means adjusting θ\theta to minimize this loss — converting the "find the rules" problem into a continuous optimization problem that calculus can solve.

Example:

How would you detect cats in photos with traditional programming vs ML?

2 of 6
When to Use ML

Machine learning costs data collection, training, and monitoring, so it should earn its place. It pays off in four situations. First, the rules are too complex to write, as in vision and language. Second, the patterns change over time, such as shopping preferences, so hand-written rules go stale. Third, there is plenty of data but no clear theory linking inputs to outputs. Fourth, the system must personalize for millions of users, where per-user rules are impossible. Each situation asks for patterns that are learned rather than specified. Without enough representative data, though, none of these justifies ML, because the model can only learn what the examples show.

Mathematical Intuition

ML is appropriate when the mapping f:X→Yf: X \to Y exists but is too complex to specify manually. The formal criterion: the problem has a pattern (ff exists), it cannot be pinned down mathematically (no closed-form), and data is available ({(xi,yi)}\{(x_i, y_i)\}). The learning bound ∣R(f^)−R^(f^)∣≤O(complexity/n)|R(\hat{f}) - \hat{R}(\hat{f})| \leq O(\sqrt{\text{complexity}/n}) shows that the gap between true and empirical risk shrinks as data nn grows — more data makes ML more reliable, while rule-based systems do not improve with more data.

Example:

Netflix recommendation system: rules or ML?

3 of 6
When NOT to Use ML

Reaching for ML by default adds cost and risk where a simpler tool works better. Skip it when explicit rules already solve the problem exactly, such as tax brackets or unit conversions. Skip it when you cannot collect enough representative, correctly labeled data. Be cautious when every decision must be explained and reproduced exactly, for example under regulation, because learned models are statistical and harder to audit. Where errors are very costly, such as medical dosing, use ML only with human review or hard safety rules around it. A useful test: if you can write a short, correct specification of the answer, write code, not a model.

Mathematical Intuition

ML adds unnecessary complexity when the problem has a known exact solution. For deterministic rules like tax calculation, f(income)=taxf(\text{income}) = \text{tax} can be coded directly with zero error. ML would introduce approximation error ϵ>0\epsilon > 0 by construction, since it learns from finite samples. The bias-variance decomposition E[(f^(x)−f(x))2]=Bias2+Variance+NoiseE[(\hat{f}(x) - f(x))^2] = \text{Bias}^2 + \text{Variance} + \text{Noise} shows that ML always has non-zero error — only justified when the alternative (manual rules) has even higher error.

Example:

Should a tax calculator use ML?

4 of 6
ML Terminology

A shared vocabulary makes every later topic easier to follow. Features, written XX, are the input variables the model sees; each row of XX is one example. The label, written yy, is the value we want to predict. A model is a function ff with parameters, chosen so that f(X)f(X) approximates yy. Training is the process of fitting those parameters to labeled examples, usually by minimizing a loss. Inference is applying the trained model to new examples whose labels are unknown. A common mistake is to treat training performance as the goal, when the real goal is accuracy at inference time on data the model has never seen.

Mathematical Intuition

The ML setup: given dataset D={(xi,yi)}i=1n\mathcal{D} = \{(\mathbf{x}_i, y_i)\}_{i=1}^n with features xi∈Rd\mathbf{x}_i \in \mathbb{R}^d and labels yiy_i, learn model f^θ:Rd→Y\hat{f}_\theta: \mathbb{R}^d \to \mathcal{Y} parameterized by θ\theta. Training minimizes R^(θ)=1n∑L(f^θ(xi),yi)\hat{R}(\theta) = \frac{1}{n}\sum L(\hat{f}_\theta(x_i), y_i). Inference computes y^=f^θ(xnew)\hat{y} = \hat{f}_\theta(x_{\text{new}}) for unseen data. The number of parameters ∣θ∣|\theta| relative to training samples nn controls model capacity — when ∣θ∣≫n|\theta| \gg n, the model can memorize rather than generalize.

Example:

House price prediction: identify features, labels, and what inference means

5 of 6
Sample, Population, Distribution

A model never sees the whole world, only a sample of it, and that gap explains many real failures. The population is every example the model might ever face, generated by an unknown probability distribution. The training set is a finite sample from it. Learning works only if the sample represents that distribution, so the patterns found in training also hold at deployment. When the sample is skewed, such as one region, one season, or one camera, the model can score well in testing and still fail in production. This is called distribution shift, and checking that training data matches deployment data is a basic first step in any project.

Mathematical Intuition

We observe a sample D={(xi,yi)}\mathcal{D} = \{(x_i, y_i)\} drawn i.i.d. from an unknown population distribution P(X,Y)P(X, Y). ML approximates the true conditional P(Y∣X)P(Y|X) using only the sample. The generalization error R(f^)=E(x,y)∼P[L(f^(x),y)]R(\hat{f}) = E_{(x,y) \sim P}[L(\hat{f}(x), y)] depends on PP, not just D\mathcal{D}. Distribution shift occurs when the training distribution PtrainP_{\text{train}} differs from deployment PdeployP_{\text{deploy}}, violating the core assumption. Domain adaptation techniques attempt to learn representations that are invariant to the shift.

Example:

Model trained on US customers, deployed globally. What could go wrong?

6 of 6
Parametric vs Non-Parametric

Model families differ in what they keep after training, and that drives memory, speed, and flexibility. A parametric model summarizes the data in a fixed number of parameters chosen in advance; linear regression with dd features keeps d+1d + 1 numbers no matter how many examples it saw. It is fast and compact but limited by its assumed shape. A non-parametric model lets its complexity grow with the data: k-nearest neighbors stores every training example, and decision trees add splits as data grows. These adapt to almost any pattern, but they cost more memory and prediction time and need more data to avoid overfitting.

Mathematical Intuition

Parametric models have a fixed number of parameters ∣θ∣|\theta| regardless of data size: linear regression has d+1d+1 parameters whether trained on 100 or 1M samples. Non-parametric models grow with data: k-NN stores all nn training points, costing O(nd)O(nd) per prediction. The bias-variance tradeoff manifests differently: parametric models may underfit if the fixed form is wrong (high bias), while non-parametric models may overfit with small nn (high variance). Kernel methods and Gaussian processes are non-parametric models with elegant theoretical guarantees.

Example:

Linear regression with 10 features: how many parameters? k-NN with 1M samples?

Theory Exercise

Problem:

A bank wants to detect fraudulent transactions. Why is ML more suitable than a rule-based system? What challenges might you face?

Hints:
  • Think about how fraud patterns evolve
  • Consider the class imbalance problem
  • Think about false positives vs false negatives

Coding Exercise

Problem:

On the breast cancer dataset, compare a hand-written threshold rule (predict 'malignant' when 'worst radius' exceeds its training median) against a fitted LogisticRegression. Report both test accuracies to show the learned model generalizes better than the manual rule.

Hints:
  • Load the data with load_breast_cancer() and split into train/test with a fixed random_state.
  • For the rule, find the index of the 'worst radius' feature, compute its median on the TRAIN set only, and threshold the test column.
  • Fit LogisticRegression(max_iter=5000, random_state=42) on the training data and compare accuracy_score on the test set for both approaches.

Related Problems on PixelBank