PIXELBANKv9.1.0
Menu
Back to ML Study Plan
Week 25-26

Chapter 13: Generative & Production ML

Master generative models that can create new data—from autoencoders for compression to VAEs for sampling and GANs for adversarial generation. Complete your ML journey by learning MLOps: the practices and tools needed to deploy, monitor, and maintain ML systems in production.

Chapter Overview

Generative models represent one of the most exciting frontiers in machine learning: systems that can create new data indistinguishable from real examples. Unlike discriminative models that learn boundaries between classes, generative models learn the underlying distribution of the data itself.

The progression from autoencoders to VAEs to GANs represents increasingly sophisticated approaches to generation. Autoencoders learn compressed representations useful for reconstruction. VAEs add probabilistic structure that enables sampling. GANs use adversarial training to produce highly realistic outputs.

Equally important is understanding how to deploy ML models in production. MLOps (Machine Learning Operations) encompasses the practices needed to reliably deploy and maintain ML systems. This includes experiment tracking, model versioning, continuous training, serving infrastructure, and monitoring for data drift.

The gap between a working notebook and a production system is substantial. Models in production face real-world challenges: changing data distributions, latency requirements, scaling concerns, and the need for reproducibility. Understanding MLOps is essential for any practicing ML engineer.

This chapter covers:

  • Autoencoders: Neural networks that learn to compress and reconstruct data through a bottleneck
  • VAEs: Variational autoencoders that learn probabilistic latent spaces enabling generation of new samples
  • GANs: Generative Adversarial Networks where generator and discriminator compete in a minimax game
  • MLOps: The full lifecycle of production ML including experiment tracking, deployment, monitoring, and retraining

Chapter Roadmap

Click any topic to jump in

1
Autoencoders

Learning compressed representations through encoder-decoder bottlenecks — dimensionality reduction, denoising, and pre-training.

Encoder-DecoderBottleneckDenoising AutoencoderSparse Autoencoder
From compression to generation

Probabilistic and adversarial approaches

2
Variational Autoencoders

Probabilistic latent spaces with the reparameterization trick — sampling new data and smooth latent interpolation.

Probabilistic EncodingELBO LossReparameterization TrickGeneration
3
GANs

Adversarial training between generator and discriminator — minimax game, mode collapse, and Wasserstein distance.

Minimax GameTraining ProcessMode CollapsePopular Variants
Deploying generative models in production
4
MLOps & Production

The full ML lifecycle — experiment tracking, model registry, serving infrastructure, and drift monitoring.

Experiment TrackingModel RegistryModel ServingMonitoring & Drift

Autoencoders learn to compress data into a lower-dimensional representation and reconstruct it. The bottleneck forces the network to learn meaningful features.

Applications: Dimensionality reduction, denoising, anomaly detection, pre-training.

In this topic

1Encoder-Decoder
2Bottleneck
3Denoising Autoencoder
4Sparse Autoencoder
1 of 4
Encoder-Decoder

Encoder: input → latent code z. Decoder: z → reconstruction. Train to minimize reconstruction error.

Mathematical Intuition

An autoencoder minimizes reconstruction loss L=∥x−x^∥2L = \|\mathbf{x} - \hat{\mathbf{x}}\|^2 where x^=gϕ(fθ(x))\hat{\mathbf{x}} = g_\phi(f_\theta(\mathbf{x})), with encoder fθ:RD→Rdf_\theta: \mathbb{R}^D \to \mathbb{R}^d and decoder gϕ:Rd→RDg_\phi: \mathbb{R}^d \to \mathbb{R}^D (d≪Dd \ll D). For linear activations, the optimal encoder learns the PCA subspace — the columns of the encoder weight matrix span the same space as the top dd eigenvectors of the data covariance matrix. Non-linear activations enable capturing non-linear manifolds that PCA cannot represent.

Example:

MNIST image (784 pixels) → latent code (32 dims) → reconstruction. What loss function? What's the compression ratio?

2 of 4
Bottleneck

Latent dimension < input dimension forces compression. Network must learn efficient representation.

Mathematical Intuition

The bottleneck dimension dd controls the information bottleneck: the encoder must discard D−dD - d dimensions of information. By the rate-distortion theorem, there exists a minimum distortion achievable at any given rate (bottleneck size). Setting dd too small loses important structure (underfitting); setting dd too large allows the identity mapping (no compression). Cross-validation on reconstruction error helps select dd, but the optimal value depends on the intrinsic dimensionality of the data manifold.

Example:

Autoencoder with latent dim = input dim. What happens? Why is this bad?

3 of 4
Denoising Autoencoder

Add noise to input, train to reconstruct clean version. Learns robust features, good for pre-training.

Mathematical Intuition

A denoising autoencoder minimizes L=Ex~∼q(x~∣x)[∥x−g(f(x~))∥2]L = \mathbb{E}_{\tilde{\mathbf{x}} \sim q(\tilde{\mathbf{x}}|\mathbf{x})} [\|\mathbf{x} - g(f(\tilde{\mathbf{x}}))\|^2] where qq is a corruption process (Gaussian noise, masking, salt-and-pepper). Vincent et al. showed this is equivalent to learning the score function ∇xlog⁡p(x)\nabla_{\mathbf{x}} \log p(\mathbf{x}) — the gradient of the log data density — which connects denoising autoencoders to score-based diffusion models. The corruption prevents identity learning even when d≥Dd \geq D.

Example:

Training: input MNIST with 30% pixels set to 0 (dropout noise). Target: original clean image. Why does this help?

4 of 4
Sparse Autoencoder

Add sparsity penalty on latent activations. Forces network to learn disentangled features.

Mathematical Intuition

Sparse autoencoders add a sparsity penalty Ω(h)=λ∑jKL(ρ∥ρ^j)\Omega(\mathbf{h}) = \lambda \sum_j \text{KL}(\rho \| \hat{\rho}_j) where ρ^j=1N∑ihj(xi)\hat{\rho}_j = \frac{1}{N}\sum_i h_j(\mathbf{x}_i) is the average activation of neuron jj and ρ≈0.05\rho \approx 0.05 is the target sparsity. The KL divergence penalty drives most neurons to be inactive for most inputs, encouraging each neuron to specialize in a specific feature. This produces overcomplete representations (d>Dd > D) that are still useful because only a sparse subset activates for any given input.

Example:

Latent dim=100, but add penalty: average activation per neuron should be ~0.05. What does the network learn?

Theory Exercise

Problem:

A linear autoencoder (no non-linear activations) is trained to minimize squared reconstruction error with a bottleneck of dimension d. What does its learned representation correspond to, and what does this imply about when a non-linear autoencoder is actually worth it?

Hints:
  • Think about what subspace minimizes reconstruction error for a linear projection
  • Recall the optimal rank-d approximation of a centered data matrix
  • Consider whether the data lies on a flat or curved manifold

Coding Exercise

Problem:

Build a small fully-connected autoencoder in PyTorch on synthetic data, train it to reconstruct the input through a tight bottleneck, and report how reconstruction error drops over training.

Hints:
  • Encoder shrinks input_dim -> hidden -> latent; decoder mirrors it back to input_dim
  • Use nn.MSELoss() as the reconstruction objective and Adam as the optimizer
  • Track the loss each epoch to confirm the bottleneck still learns a useful compression

Related Problems on PixelBank