Chapter 13: Generative & Production ML
Master generative models that can create new data—from autoencoders for compression to VAEs for sampling and GANs for adversarial generation. Complete your ML journey by learning MLOps: the practices and tools needed to deploy, monitor, and maintain ML systems in production.
Chapter Overview
Generative models represent one of the most exciting frontiers in machine learning: systems that can create new data indistinguishable from real examples. Unlike discriminative models that learn boundaries between classes, generative models learn the underlying distribution of the data itself.
The progression from autoencoders to VAEs to GANs represents increasingly sophisticated approaches to generation. Autoencoders learn compressed representations useful for reconstruction. VAEs add probabilistic structure that enables sampling. GANs use adversarial training to produce highly realistic outputs.
Equally important is understanding how to deploy ML models in production. MLOps (Machine Learning Operations) encompasses the practices needed to reliably deploy and maintain ML systems. This includes experiment tracking, model versioning, continuous training, serving infrastructure, and monitoring for data drift.
The gap between a working notebook and a production system is substantial. Models in production face real-world challenges: changing data distributions, latency requirements, scaling concerns, and the need for reproducibility. Understanding MLOps is essential for any practicing ML engineer.
This chapter covers:
- Autoencoders: Neural networks that learn to compress and reconstruct data through a bottleneck
- VAEs: Variational autoencoders that learn probabilistic latent spaces enabling generation of new samples
- GANs: Generative Adversarial Networks where generator and discriminator compete in a minimax game
- MLOps: The full lifecycle of production ML including experiment tracking, deployment, monitoring, and retraining
Chapter Roadmap
Click any topic to jump in
Autoencoders
Learning compressed representations through encoder-decoder bottlenecks — dimensionality reduction, denoising, and pre-training.
Probabilistic and adversarial approaches
Variational Autoencoders
Probabilistic latent spaces with the reparameterization trick — sampling new data and smooth latent interpolation.
GANs
Adversarial training between generator and discriminator — minimax game, mode collapse, and Wasserstein distance.
MLOps & Production
The full ML lifecycle — experiment tracking, model registry, serving infrastructure, and drift monitoring.
Autoencoders learn to compress data into a lower-dimensional representation and reconstruct it. The bottleneck forces the network to learn meaningful features.
Applications: Dimensionality reduction, denoising, anomaly detection, pre-training.
In this topic
Encoder-Decoder
Encoder: input → latent code z. Decoder: z → reconstruction. Train to minimize reconstruction error.
An autoencoder minimizes reconstruction loss where , with encoder and decoder (). For linear activations, the optimal encoder learns the PCA subspace — the columns of the encoder weight matrix span the same space as the top eigenvectors of the data covariance matrix. Non-linear activations enable capturing non-linear manifolds that PCA cannot represent.
MNIST image (784 pixels) → latent code (32 dims) → reconstruction. What loss function? What's the compression ratio?
Bottleneck
Latent dimension < input dimension forces compression. Network must learn efficient representation.
The bottleneck dimension controls the information bottleneck: the encoder must discard dimensions of information. By the rate-distortion theorem, there exists a minimum distortion achievable at any given rate (bottleneck size). Setting too small loses important structure (underfitting); setting too large allows the identity mapping (no compression). Cross-validation on reconstruction error helps select , but the optimal value depends on the intrinsic dimensionality of the data manifold.
Autoencoder with latent dim = input dim. What happens? Why is this bad?
Denoising Autoencoder
Add noise to input, train to reconstruct clean version. Learns robust features, good for pre-training.
A denoising autoencoder minimizes where is a corruption process (Gaussian noise, masking, salt-and-pepper). Vincent et al. showed this is equivalent to learning the score function — the gradient of the log data density — which connects denoising autoencoders to score-based diffusion models. The corruption prevents identity learning even when .
Training: input MNIST with 30% pixels set to 0 (dropout noise). Target: original clean image. Why does this help?
Sparse Autoencoder
Add sparsity penalty on latent activations. Forces network to learn disentangled features.
Sparse autoencoders add a sparsity penalty where is the average activation of neuron and is the target sparsity. The KL divergence penalty drives most neurons to be inactive for most inputs, encouraging each neuron to specialize in a specific feature. This produces overcomplete representations () that are still useful because only a sparse subset activates for any given input.
Latent dim=100, but add penalty: average activation per neuron should be ~0.05. What does the network learn?
Theory Exercise
Problem:
A linear autoencoder (no non-linear activations) is trained to minimize squared reconstruction error with a bottleneck of dimension d. What does its learned representation correspond to, and what does this imply about when a non-linear autoencoder is actually worth it?
Hints:
- Think about what subspace minimizes reconstruction error for a linear projection
- Recall the optimal rank-d approximation of a centered data matrix
- Consider whether the data lies on a flat or curved manifold
Coding Exercise
Problem:
Build a small fully-connected autoencoder in PyTorch on synthetic data, train it to reconstruct the input through a tight bottleneck, and report how reconstruction error drops over training.
Hints:
- Encoder shrinks input_dim -> hidden -> latent; decoder mirrors it back to input_dim
- Use nn.MSELoss() as the reconstruction objective and Adam as the optimizer
- Track the loss each epoch to confirm the bottleneck still learns a useful compression
Related Problems on PixelBank
VAEs are generative models that learn a continuous latent space. Unlike regular autoencoders, VAEs can generate new samples by sampling from the latent distribution.
Key insight: Encode to distribution parameters (μ, σ), sample z, decode. Regularize latent space to be Gaussian.
In this topic
Probabilistic Encoding
Encoder outputs mean and variance. Sample z using reparameterization: z = μ + σ·ε, ε~N(0,1).
The VAE maximizes the evidence lower bound: . The first term is reconstruction quality; the KL term regularizes the approximate posterior to be close to the prior . The reparameterization trick with enables backpropagation through the sampling step since the randomness is externalized.
Regular autoencoder: image → z = [2.3, -1.5]. VAE: image → μ=[2.3, -1.5], σ=[0.1, 0.2]. What's different?
ELBO Loss
Reconstruction loss + KL divergence. KL term regularizes latent space toward prior N(0,I).
For Gaussian and prior , the KL divergence has a closed form: . This penalizes latent dimensions that deviate from the standard normal — pushing means toward 0 and variances toward 1. The 'posterior collapse' problem occurs when the KL dominates and the encoder ignores the input entirely, outputting for all inputs.
Encoder outputs μ=[3.0, 0], σ=[0.5, 0.5]. Prior is N(0,1). KL divergence high or low? What happens during training?
Reparameterization Trick
z = μ + σ·ε allows backprop through sampling. ε is random, μ and σ are learned.
The VAE latent space is continuous and smooth: nearby points in decode to similar outputs because the KL regularization prevents holes in the latent space. Linear interpolation between two encoded points produces a smooth transition in output space. The prior means we can generate new samples by simply drawing and decoding — the KL term ensures the decoder sees this distribution during training.
Without trick: z ~ N(μ,σ²). Can't backprop through sampling! With trick: z = μ + σ·ε. How does gradient flow?
Generation
Sample z ~ N(0,I), decode to generate new data. Interpolate in z-space for smooth transitions.
A conditional VAE (CVAE) conditions both encoder and decoder on auxiliary information : and . The ELBO becomes . This enables controlled generation: for image generation conditioned on class labels, the CVAE learns to separate style () from content (), allowing generation of new images for any specified class.
Face VAE: z₁ encodes 'smiling woman', z₂ encodes 'frowning man'. Generate face at z = 0.5·z₁ + 0.5·z₂. Result?
Theory Exercise
Problem:
In a VAE, why is the reparameterization trick z = mu + sigma * epsilon necessary? What goes wrong if you instead sample z directly from N(mu, sigma^2) inside the forward pass?
Hints:
- Think about what backpropagation needs to compute gradients of the loss with respect to mu and sigma
- A sampling operation is not a differentiable function of its parameters
- Where does the randomness live in the reparameterized form?
Coding Exercise
Problem:
Implement the closed-form Gaussian KL term of the VAE loss and the reparameterization trick in PyTorch, then confirm that the KL is zero when the encoder outputs exactly match the standard-normal prior.
Hints:
- The encoder produces mu and logvar (log of variance) for numerical stability
- KL = -0.5 * sum(1 + logvar - mu^2 - exp(logvar)) per the closed form for diagonal Gaussians
- Reparameterize with z = mu + exp(0.5 * logvar) * epsilon, epsilon ~ N(0, I)
Related Problems on PixelBank
GANs consist of two networks: a Generator that creates fake data and a Discriminator that tries to distinguish real from fake. They compete in a minimax game.
Key challenge: Training can be unstable. Many tricks needed for good results.
In this topic
Minimax Game
D maximizes: correctly classify real/fake. G minimizes: fool D.
The GAN objective has the global optimum . At equilibrium, and everywhere. The minimax value at optimality equals , and the divergence being minimized is the Jensen-Shannon divergence .
D outputs P(real). For real image, D(x)=0.9. For G's fake, D(G(z))=0.3. What's each term? Who's winning?
Training Process
Alternate: 1) Train D on real + fake. 2) Train G to fool D. Balance is critical.
GAN training alternates between steps of discriminator optimization and 1 step of generator optimization. The discriminator gradient is , maximizing classification accuracy. The generator gradient uses the modified objective instead of because the latter has vanishing gradients when is confident — the modified version provides stronger learning signal early in training.
D gets too strong (always outputs 0 for fakes). What happens to G's gradient?
Mode Collapse
G produces limited variety of outputs. Fixes: minibatch discrimination, feature matching, progressive training.
Mode collapse occurs when maps many different values to the same output region. Formally, the support of is much smaller than the support of . This is a Nash equilibrium failure: G finds a mode that maximally fools D, and D cannot distinguish it from real data in that mode. Minibatch discrimination adds a feature that measures diversity within a batch, penalizing generators that produce identical outputs.
MNIST GAN: G only outputs '1' digits (gets D(G(z))=0.5 for all). Why? How to detect?
Popular Variants
DCGAN: Conv layers. WGAN: Wasserstein distance. StyleGAN: State-of-the-art faces. BigGAN: Large scale.
WGAN replaces the JS divergence with the Wasserstein-1 distance , estimated via the Kantorovich-Rubinstein dual: . The Lipschitz constraint is enforced via weight clipping or gradient penalty . Wasserstein distance provides gradients everywhere, even when distributions have non-overlapping support — the key advantage over JS divergence.
WGAN uses Wasserstein distance instead of JS divergence. Why does this help training stability?
Theory Exercise
Problem:
The original GAN generator loss log(1 - D(G(z))) suffers from vanishing gradients early in training. Explain why, and describe the 'non-saturating' fix and why it provides a stronger learning signal.
Hints:
- Early on, the discriminator easily rejects fakes, so D(G(z)) is near 0
- Examine the gradient of log(1 - D(G(z))) as D(G(z)) -> 0
- Consider the alternative objective of maximizing log(D(G(z)))
Coding Exercise
Problem:
Demonstrate why the minimax GAN generator loss saturates by numerically comparing the gradient magnitude of log(1 - D(G(z))) versus the non-saturating -log(D(G(z))) as a function of the discriminator's output on fakes.
Hints:
- Treat D(G(z)) as a single scalar d in (0, 1) and use autograd to get gradients
- Compute both the saturating loss log(1 - d) and the non-saturating loss -log(d)
- Evaluate at small d (fakes easily detected) to see which loss still has a usable gradient
Related Problems on PixelBank
MLOps (Machine Learning Operations) applies DevOps principles to ML systems. It covers the entire lifecycle: data versioning, experiment tracking, model deployment, monitoring, and retraining.
Key challenge: ML systems have unique challenges - data drift, model decay, reproducibility.
In this topic
Experiment Tracking
Log parameters, metrics, artifacts for reproducibility. Tools: MLflow, Weights & Biases, Neptune.
Experiment tracking logs the mapping where is hyperparameters, is the data version hash, is metrics, and is artifacts (model weights, configs). The search space for hyperparameters grows as for hyperparameters with values each (grid search) or is explored in random trials to find an -optimal configuration (random search). Bayesian optimization uses a Gaussian process surrogate to select the next trial in for past observations, but amortizes this with better sample efficiency.
Run 50 experiments varying lr, batch_size, architecture. 3 months later, which config gave best F1? How do you reproduce it?
Model Registry
Version and stage models (staging, production). Track lineage: which data and code produced which model.
A model registry maintains a DAG (directed acyclic graph) of model lineage: . Versioning enables reproducibility: any model can be re-created from its lineage. The staging workflow (development -> staging -> production -> archived) ensures only validated models serve traffic. Canary deployments route of traffic to the new model and monitor metrics before full rollout.
Production model v2.3 has bug. Need to rollback. What metadata should registry have?
Model Serving
Deploy as REST API, batch prediction, or edge. Tools: TensorFlow Serving, TorchServe, Triton, BentoML.
Serving latency decomposes as . For a transformer with layers, dimensions, and input length : . Dynamic batching groups requests arriving within a time window to amortize GPU kernel launch overhead — optimal batch size balances throughput () against latency ( for the first request in the batch). Quantization (INT8) reduces model size by and inference latency by - with accuracy loss.
Model needs: <50ms latency, 1000 req/s, A/B testing between 2 models. What serving architecture?
Monitoring & Drift
Track prediction distributions, detect data drift and model degradation. Trigger retraining when needed.
Data drift is detected by comparing feature distributions: Population Stability Index measures divergence between training () and serving () distributions, with indicating significant drift. Concept drift occurs when changes even if is stable — detectable only with labeled data. Monitoring tracks prediction distribution entropy: a sudden increase suggests the model is uncertain about new data patterns. Automated retraining triggers when drift exceeds thresholds for consecutive measurement windows.
Credit scoring model trained on 2023 data. In 2024, fraud patterns changed (new attack vectors). How do you detect this?
Theory Exercise
Problem:
Your deployed model's accuracy dropped from 95% to 85% over 3 months. What could cause this, and how would you detect and fix it?
Hints:
- Think about what changes over time
- Consider data distribution shifts
- What monitoring would help?
Coding Exercise
Problem:
Implement a lightweight data-drift check using the Population Stability Index (PSI). Compare a 'training' feature distribution against a shifted 'serving' distribution and flag drift when PSI crosses the standard 0.2 threshold.
Hints:
- Bin both distributions using the SAME bin edges (derived from the reference/training data)
- PSI = sum over bins of (p_serve - p_train) * ln(p_serve / p_train); add a tiny epsilon to avoid log(0)
- PSI < 0.1 = stable, 0.1-0.2 = moderate shift, > 0.2 = significant drift