PIXELBANKv9.1.0
Menu
Back to ML Study Plan
Week 25-26

Chapter 13: Generative & Production ML

Master generative models that can create new data—from autoencoders for compression to VAEs for sampling and GANs for adversarial generation. Complete your ML journey by learning MLOps: the practices and tools needed to deploy, monitor, and maintain ML systems in production.

Chapter Overview

Generative models represent one of the most exciting frontiers in machine learning: systems that can create new data indistinguishable from real examples. Unlike discriminative models that learn boundaries between classes, generative models learn the underlying distribution of the data itself.

The progression from autoencoders to VAEs to GANs represents increasingly sophisticated approaches to generation. Autoencoders learn compressed representations useful for reconstruction. VAEs add probabilistic structure that enables sampling. GANs use adversarial training to produce highly realistic outputs.

Equally important is understanding how to deploy ML models in production. MLOps (Machine Learning Operations) encompasses the practices needed to reliably deploy and maintain ML systems. This includes experiment tracking, model versioning, continuous training, serving infrastructure, and monitoring for data drift.

The gap between a working notebook and a production system is substantial. Models in production face real-world challenges: changing data distributions, latency requirements, scaling concerns, and the need for reproducibility. Understanding MLOps is essential for any practicing ML engineer.

This chapter covers:

  • Autoencoders: Neural networks that learn to compress and reconstruct data through a bottleneck
  • VAEs: Variational autoencoders that learn probabilistic latent spaces enabling generation of new samples
  • GANs: Generative Adversarial Networks where generator and discriminator compete in a minimax game
  • MLOps: The full lifecycle of production ML including experiment tracking, deployment, monitoring, and retraining

Chapter Roadmap

Click any topic to jump in

1
Autoencoders

Learning compressed representations through encoder-decoder bottlenecks — dimensionality reduction, denoising, and pre-training.

Encoder-DecoderBottleneckDenoising AutoencoderSparse Autoencoder
From compression to generation

Probabilistic and adversarial approaches

2
Variational Autoencoders

Probabilistic latent spaces with the reparameterization trick — sampling new data and smooth latent interpolation.

Probabilistic EncodingELBO LossReparameterization TrickGeneration
3
GANs

Adversarial training between generator and discriminator — minimax game, mode collapse, and Wasserstein distance.

Minimax GameTraining ProcessMode CollapsePopular Variants
Deploying generative models in production
4
MLOps & Production

The full ML lifecycle — experiment tracking, model registry, serving infrastructure, and drift monitoring.

Experiment TrackingModel RegistryModel ServingMonitoring & Drift

The previous chapter, Reinforcement Learning, trained agents from reward signals. Every model so far in this plan has needed some external target: a label, a next token, or a reward. Most of the world's data carries none of these. A factory logs millions of sensor readings and a hospital stores millions of scans, but nobody has annotated them. Can a network still learn useful features from raw inputs alone, with no target except the input itself?

This chapter, Generative and Production ML, closes the plan in two halves: models that learn the structure of data well enough to compress and generate it, and the engineering that keeps any model working after deployment. This first topic, Autoencoders, answers the question above with a simple trick: train a network to reproduce its input through a narrow middle layer. It starts with the encoder-decoder structure, then shows why the bottleneck is what forces learning. Two variants follow: denoising autoencoders, which reconstruct clean inputs from corrupted ones, and sparse autoencoders, which keep most hidden units silent.

Definition

An autoencoder is a neural network trained to reconstruct its own input. An encoder fθf_\theta maps an input xx to a latent code z=fθ(x)z = f_\theta(x), a decoder gϕg_\phi maps the code back to x^=gϕ(z)\hat{x} = g_\phi(z), and training minimizes a reconstruction loss between xx and x^\hat{x}. A constraint such as a narrow code, input noise, or sparsity prevents simple copying.

In this topic

1Encoder-Decoder
2Bottleneck
3Denoising Autoencoder
4Sparse Autoencoder
1 of 4
Encoder-Decoder

Labels are expensive, but an input can serve as its own target. The encoder fθf_\theta maps an input xx with DD features to a code zz with dd features; the decoder gϕg_\phi maps zz back to a reconstruction x^\hat{x} with DD features. Training minimizes a reconstruction loss such as mean squared error, or binary cross-entropy when pixels lie in [0, 1]. With linear layers and squared error, the best autoencoder spans the same subspace as PCA's top dd components, so nonlinear layers are what let it follow curved data manifolds. The code becomes a feature vector for later tasks such as clustering or classification.

Mathematical Intuition

An autoencoder minimizes reconstruction loss L=∥x−x^∥2L = \|\mathbf{x} - \hat{\mathbf{x}}\|^2 where x^=gϕ(fθ(x))\hat{\mathbf{x}} = g_\phi(f_\theta(\mathbf{x})), with encoder fθ:RD→Rdf_\theta: \mathbb{R}^D \to \mathbb{R}^d and decoder gϕ:Rd→RDg_\phi: \mathbb{R}^d \to \mathbb{R}^D (d≪Dd \ll D). For linear activations, the optimal encoder learns the PCA subspace — the columns of the encoder weight matrix span the same space as the top dd eigenvectors of the data covariance matrix. Non-linear activations enable capturing non-linear manifolds that PCA cannot represent.

Example:

MNIST image (784 pixels) → latent code (32 dims) → reconstruction. What loss function? What's the compression ratio?

2 of 4
Bottleneck

Reconstruction alone teaches nothing if copying is allowed. When the code dimension dd is at least the input dimension DD, the network can learn the identity map: perfect reconstruction with features that are no more useful than raw pixels. A bottleneck with dd much smaller than DD forces the encoder to keep only the information that best explains the data and to discard noise and redundancy. Choosing dd is a trade-off: too small loses real structure and blurs reconstructions; too large drifts back toward copying. Reconstruction error always falls as dd grows, so pick dd by downstream usefulness, not by reconstruction loss alone.

Mathematical Intuition

The bottleneck dimension dd controls the information bottleneck: the encoder must discard D−dD - d dimensions of information. By the rate-distortion theorem, there exists a minimum distortion achievable at any given rate (bottleneck size). Setting dd too small loses important structure (underfitting); setting dd too large allows the identity mapping (no compression). Cross-validation on reconstruction error helps select dd, but the optimal value depends on the intrinsic dimensionality of the data manifold.

Example:

Autoencoder with latent dim = input dim. What happens? Why is this bad?

3 of 4
Denoising Autoencoder

A denoising autoencoder blocks copying without a narrow code. Each training input xx is corrupted to x~\tilde{x}, for example by zeroing a random 30 percent of pixels or adding Gaussian noise, and the network must output the clean xx. Copying x~\tilde{x} now gives a poor loss, so the network must learn how pixels depend on one another in order to fill in the gaps. Vincent and colleagues showed that stacked denoising features pre-trained deep networks well. The noise level is a hyperparameter: too little allows near-copying, while too much destroys the information needed. Later, denoising score matching connected this idea to diffusion models.

Mathematical Intuition

A denoising autoencoder minimizes L=Ex~∼q(x~∣x)[∥x−g(f(x~))∥2]L = \mathbb{E}_{\tilde{\mathbf{x}} \sim q(\tilde{\mathbf{x}} \mid \mathbf{x})} [\Vert \mathbf{x} - g(f(\tilde{\mathbf{x}})) \Vert^2] where qq is a corruption process such as Gaussian noise, masking, or salt-and-pepper noise. Vincent (2011) showed that with small Gaussian noise of variance σ2\sigma^2, this objective is equivalent to denoising score matching: the optimal reconstruction satisfies r(x~)−x~≈σ2∇xlog⁡p(x~)r(\tilde{\mathbf{x}}) - \tilde{\mathbf{x}} \approx \sigma^2 \nabla_{\mathbf{x}} \log p(\tilde{\mathbf{x}}), so the network points toward regions of higher data density. That link underlies score-based diffusion models. Because the input is corrupted, copying is no longer optimal, so the network learns useful features even when d≥Dd \geq D.

Example:

Training: input MNIST with 30% pixels set to 0 (dropout noise). Target: original clean image. Why does this help?

4 of 4
Sparse Autoencoder

A sparse autoencoder can use a code wider than the input, called overcomplete, yet still avoid copying, because a penalty keeps most code units inactive for any given input. One form adds λ\lambda times the sum over units jj of KL(ρ ∥ ρ^j)\mathrm{KL}(\rho \,\Vert\, \hat{\rho}_j), where ρ^j\hat{\rho}_j is unit jj's average activation over the data and ρ\rho is a small target such as 0.05; another adds an L1 penalty on activations. Each unit then specializes in a pattern that appears in only a few inputs, which tends to make units more interpretable. The penalty weight λ\lambda trades reconstruction for sparsity; set too high, units go permanently dead.

Mathematical Intuition

Sparse autoencoders add a sparsity penalty Ω(h)=λ∑jKL(ρ∥ρ^j)\Omega(\mathbf{h}) = \lambda \sum_j \text{KL}(\rho \| \hat{\rho}_j) where ρ^j=1N∑ihj(xi)\hat{\rho}_j = \frac{1}{N}\sum_i h_j(\mathbf{x}_i) is the average activation of neuron jj and ρ≈0.05\rho \approx 0.05 is the target sparsity. The KL divergence penalty drives most neurons to be inactive for most inputs, encouraging each neuron to specialize in a specific feature. This produces overcomplete representations (d>Dd > D) that are still useful because only a sparse subset activates for any given input.

Example:

Latent dim=100, but add penalty: average activation per neuron should be ~0.05. What does the network learn?

Theory Exercise

Problem:

A linear autoencoder (no non-linear activations) is trained to minimize squared reconstruction error with a bottleneck of dimension d. What does its learned representation correspond to, and what does this imply about when a non-linear autoencoder is actually worth it?

Hints:
  • Think about what subspace minimizes reconstruction error for a linear projection
  • Recall the optimal rank-d approximation of a centered data matrix
  • Consider whether the data lies on a flat or curved manifold

Coding Exercise

Problem:

Build a small fully-connected autoencoder in PyTorch on synthetic data, train it to reconstruct the input through a tight bottleneck, and report how reconstruction error drops over training.

Hints:
  • Encoder shrinks input_dim -> hidden -> latent; decoder mirrors it back to input_dim
  • Use nn.MSELoss() as the reconstruction objective and Adam as the optimizer
  • Track the loss each epoch to confirm the bottleneck still learns a useful compression

Related Problems on PixelBank