Chapter 13: Generative & Production ML
Master generative models that can create new data—from autoencoders for compression to VAEs for sampling and GANs for adversarial generation. Complete your ML journey by learning MLOps: the practices and tools needed to deploy, monitor, and maintain ML systems in production.
Chapter Overview
Generative models represent one of the most exciting frontiers in machine learning: systems that can create new data indistinguishable from real examples. Unlike discriminative models that learn boundaries between classes, generative models learn the underlying distribution of the data itself.
The progression from autoencoders to VAEs to GANs represents increasingly sophisticated approaches to generation. Autoencoders learn compressed representations useful for reconstruction. VAEs add probabilistic structure that enables sampling. GANs use adversarial training to produce highly realistic outputs.
Equally important is understanding how to deploy ML models in production. MLOps (Machine Learning Operations) encompasses the practices needed to reliably deploy and maintain ML systems. This includes experiment tracking, model versioning, continuous training, serving infrastructure, and monitoring for data drift.
The gap between a working notebook and a production system is substantial. Models in production face real-world challenges: changing data distributions, latency requirements, scaling concerns, and the need for reproducibility. Understanding MLOps is essential for any practicing ML engineer.
This chapter covers:
- Autoencoders: Neural networks that learn to compress and reconstruct data through a bottleneck
- VAEs: Variational autoencoders that learn probabilistic latent spaces enabling generation of new samples
- GANs: Generative Adversarial Networks where generator and discriminator compete in a minimax game
- MLOps: The full lifecycle of production ML including experiment tracking, deployment, monitoring, and retraining
Chapter Roadmap
Click any topic to jump in
Autoencoders
Learning compressed representations through encoder-decoder bottlenecks — dimensionality reduction, denoising, and pre-training.
Probabilistic and adversarial approaches
Variational Autoencoders
Probabilistic latent spaces with the reparameterization trick — sampling new data and smooth latent interpolation.
GANs
Adversarial training between generator and discriminator — minimax game, mode collapse, and Wasserstein distance.
MLOps & Production
The full ML lifecycle — experiment tracking, model registry, serving infrastructure, and drift monitoring.
The previous chapter, Reinforcement Learning, trained agents from reward signals. Every model so far in this plan has needed some external target: a label, a next token, or a reward. Most of the world's data carries none of these. A factory logs millions of sensor readings and a hospital stores millions of scans, but nobody has annotated them. Can a network still learn useful features from raw inputs alone, with no target except the input itself?
This chapter, Generative and Production ML, closes the plan in two halves: models that learn the structure of data well enough to compress and generate it, and the engineering that keeps any model working after deployment. This first topic, Autoencoders, answers the question above with a simple trick: train a network to reproduce its input through a narrow middle layer. It starts with the encoder-decoder structure, then shows why the bottleneck is what forces learning. Two variants follow: denoising autoencoders, which reconstruct clean inputs from corrupted ones, and sparse autoencoders, which keep most hidden units silent.
Definition
An autoencoder is a neural network trained to reconstruct its own input. An encoder maps an input to a latent code , a decoder maps the code back to , and training minimizes a reconstruction loss between and . A constraint such as a narrow code, input noise, or sparsity prevents simple copying.
In this topic
Encoder-Decoder
Labels are expensive, but an input can serve as its own target. The encoder maps an input with features to a code with features; the decoder maps back to a reconstruction with features. Training minimizes a reconstruction loss such as mean squared error, or binary cross-entropy when pixels lie in [0, 1]. With linear layers and squared error, the best autoencoder spans the same subspace as PCA's top components, so nonlinear layers are what let it follow curved data manifolds. The code becomes a feature vector for later tasks such as clustering or classification.
An autoencoder minimizes reconstruction loss where , with encoder and decoder (). For linear activations, the optimal encoder learns the PCA subspace — the columns of the encoder weight matrix span the same space as the top eigenvectors of the data covariance matrix. Non-linear activations enable capturing non-linear manifolds that PCA cannot represent.
MNIST image (784 pixels) → latent code (32 dims) → reconstruction. What loss function? What's the compression ratio?
Bottleneck
Reconstruction alone teaches nothing if copying is allowed. When the code dimension is at least the input dimension , the network can learn the identity map: perfect reconstruction with features that are no more useful than raw pixels. A bottleneck with much smaller than forces the encoder to keep only the information that best explains the data and to discard noise and redundancy. Choosing is a trade-off: too small loses real structure and blurs reconstructions; too large drifts back toward copying. Reconstruction error always falls as grows, so pick by downstream usefulness, not by reconstruction loss alone.
The bottleneck dimension controls the information bottleneck: the encoder must discard dimensions of information. By the rate-distortion theorem, there exists a minimum distortion achievable at any given rate (bottleneck size). Setting too small loses important structure (underfitting); setting too large allows the identity mapping (no compression). Cross-validation on reconstruction error helps select , but the optimal value depends on the intrinsic dimensionality of the data manifold.
Autoencoder with latent dim = input dim. What happens? Why is this bad?
Denoising Autoencoder
A denoising autoencoder blocks copying without a narrow code. Each training input is corrupted to , for example by zeroing a random 30 percent of pixels or adding Gaussian noise, and the network must output the clean . Copying now gives a poor loss, so the network must learn how pixels depend on one another in order to fill in the gaps. Vincent and colleagues showed that stacked denoising features pre-trained deep networks well. The noise level is a hyperparameter: too little allows near-copying, while too much destroys the information needed. Later, denoising score matching connected this idea to diffusion models.
A denoising autoencoder minimizes where is a corruption process such as Gaussian noise, masking, or salt-and-pepper noise. Vincent (2011) showed that with small Gaussian noise of variance , this objective is equivalent to denoising score matching: the optimal reconstruction satisfies , so the network points toward regions of higher data density. That link underlies score-based diffusion models. Because the input is corrupted, copying is no longer optimal, so the network learns useful features even when .
Training: input MNIST with 30% pixels set to 0 (dropout noise). Target: original clean image. Why does this help?
Sparse Autoencoder
A sparse autoencoder can use a code wider than the input, called overcomplete, yet still avoid copying, because a penalty keeps most code units inactive for any given input. One form adds times the sum over units of , where is unit 's average activation over the data and is a small target such as 0.05; another adds an L1 penalty on activations. Each unit then specializes in a pattern that appears in only a few inputs, which tends to make units more interpretable. The penalty weight trades reconstruction for sparsity; set too high, units go permanently dead.
Sparse autoencoders add a sparsity penalty where is the average activation of neuron and is the target sparsity. The KL divergence penalty drives most neurons to be inactive for most inputs, encouraging each neuron to specialize in a specific feature. This produces overcomplete representations () that are still useful because only a sparse subset activates for any given input.
Latent dim=100, but add penalty: average activation per neuron should be ~0.05. What does the network learn?
Theory Exercise
Problem:
A linear autoencoder (no non-linear activations) is trained to minimize squared reconstruction error with a bottleneck of dimension d. What does its learned representation correspond to, and what does this imply about when a non-linear autoencoder is actually worth it?
Hints:
- Think about what subspace minimizes reconstruction error for a linear projection
- Recall the optimal rank-d approximation of a centered data matrix
- Consider whether the data lies on a flat or curved manifold
Coding Exercise
Problem:
Build a small fully-connected autoencoder in PyTorch on synthetic data, train it to reconstruct the input through a tight bottleneck, and report how reconstruction error drops over training.
Hints:
- Encoder shrinks input_dim -> hidden -> latent; decoder mirrors it back to input_dim
- Use nn.MSELoss() as the reconstruction objective and Adam as the optimizer
- Track the loss each epoch to confirm the bottleneck still learns a useful compression
Related Problems on PixelBank
The previous topic, Autoencoders, compressed each input to a single code and reconstructed it. That code is good for compression but poor for generation. If you pick a random point in a plain autoencoder's latent space and decode it, you usually get garbage, because the encoder only ever placed training images at scattered points and nothing taught the decoder what lies between them. To generate new data we need a latent space with no holes and a known distribution we can sample from.
This topic, Variational Autoencoders, gets there by making the encoder probabilistic. It starts with probabilistic encoding, where each input maps to a Gaussian distribution rather than a point. The ELBO loss then explains the training objective: reconstruction plus a KL divergence that pulls every encoded distribution toward a standard normal prior. The reparameterization trick shows how to backpropagate through random sampling. Finally, generation covers sampling from the prior and interpolating between codes, and explains why VAE samples tend to look blurrier than the GAN samples of the next topic.
Definition
A variational autoencoder is a latent-variable generative model with a prior , a decoder , and an encoder that approximates the true posterior. The encoder and decoder are trained jointly to maximize the evidence lower bound (ELBO) on , using the reparameterization trick to obtain low-variance gradients.
In this topic
Probabilistic Encoding
A plain encoder maps each input to one point, so the decoder never learns what nearby points mean. A VAE encoder instead outputs the parameters of a Gaussian, : a mean vector and a standard deviation vector , one entry per latent dimension. In practice it outputs for numerical stability. During training the decoder sees a different sample on each pass, so it must decode a whole neighborhood around to the same input. That smears each input over a region, filling the gaps between training points. When collapses toward 0, the VAE degenerates into an ordinary autoencoder.
The VAE maximizes the evidence lower bound: . The first term is reconstruction quality; the KL term regularizes the approximate posterior to be close to the prior . The reparameterization trick with enables backpropagation through the sampling step since the randomness is externalized.
Regular autoencoder: image → z = [2.3, -1.5]. VAE: image → μ=[2.3, -1.5], σ=[0.1, 0.2]. What's different?
ELBO Loss
We want to maximize the data likelihood , but it requires integrating over every , which is intractable. The evidence lower bound replaces it: . The first term rewards accurate reconstruction of from sampled codes. The second penalizes each encoded Gaussian for straying from the prior ; for Gaussians it has the closed form one half times the sum of . The gap between ELBO and true likelihood is the KL to the true posterior. If KL dominates, the encoder ignores , called posterior collapse; beta-VAE reweights the KL deliberately.
For Gaussian and prior , the KL divergence has a closed form: . This penalizes latent dimensions that deviate from the standard normal — pushing means toward 0 and variances toward 1. The 'posterior collapse' problem occurs when the KL dominates and the encoder ignores the input entirely, outputting for all inputs.
Encoder outputs μ=[3.0, 0], σ=[0.5, 0.5]. Prior is N(0,1). KL divergence high or low? What happens during training?
Reparameterization Trick
Training a VAE needs the gradient of the reconstruction loss with respect to and , but is drawn at random, and you cannot differentiate through the act of sampling. The reparameterization trick rewrites the sample as with drawn independently of the parameters. Now is a deterministic, differentiable function of and , and the randomness enters like an extra input. Gradients flow as usual: and . The alternative, the score-function or REINFORCE estimator, is unbiased but has far higher variance, which makes training slow.
The VAE latent space is continuous and smooth: nearby points in decode to similar outputs because the KL regularization prevents holes in the latent space. Linear interpolation between two encoded points produces a smooth transition in output space. The prior means we can generate new samples by simply drawing and decoding — the KL term ensures the decoder sees this distribution during training.
Without trick: z ~ N(μ,σ²). Can't backprop through sampling! With trick: z = μ + σ·ε. How does gradient flow?
Generation
Once trained, the encoder is no longer needed to create data. Sample , pass it through the decoder, and you get a new example; this works because the KL term made the encoded distributions cover the prior. Interpolating between two codes, for from 0 to 1, yields a smooth sequence of plausible outputs, and adding attribute directions, such as the average code of smiling faces minus that of neutral faces, can edit images. VAE samples are often blurry, because a Gaussian or squared-error decoder averages over the plausible outputs for each code.
A conditional VAE (CVAE) conditions both encoder and decoder on auxiliary information : and . The ELBO becomes . This enables controlled generation: for image generation conditioned on class labels, the CVAE learns to separate style () from content (), allowing generation of new images for any specified class.
Face VAE: z₁ encodes 'smiling woman', z₂ encodes 'frowning man'. Generate face at z = 0.5·z₁ + 0.5·z₂. Result?
Theory Exercise
Problem:
In a VAE, why is the reparameterization trick z = mu + sigma * epsilon necessary? What goes wrong if you instead sample z directly from N(mu, sigma^2) inside the forward pass?
Hints:
- Think about what backpropagation needs to compute gradients of the loss with respect to mu and sigma
- A sampling operation is not a differentiable function of its parameters
- Where does the randomness live in the reparameterized form?
Coding Exercise
Problem:
Implement the closed-form Gaussian KL term of the VAE loss and the reparameterization trick in PyTorch, then confirm that the KL is zero when the encoder outputs exactly match the standard-normal prior.
Hints:
- The encoder produces mu and logvar (log of variance) for numerical stability
- KL = -0.5 * sum(1 + logvar - mu^2 - exp(logvar)) per the closed form for diagonal Gaussians
- Reparameterize with z = mu + exp(0.5 * logvar) * epsilon, epsilon ~ N(0, I)
Related Problems on PixelBank
The previous topic, Variational Autoencoders, generated data by sampling a latent code and decoding it, but its samples tend to be blurry. The cause is the loss: a squared-error or Gaussian likelihood rewards the decoder for predicting the average of all plausible outputs, and an average of many sharp faces is a soft face. What if, instead of scoring samples against a fixed formula, we trained a second network to judge whether a sample looks real, and let it find whatever flaws remain?
This topic, Generative Adversarial Networks, builds that idea. A generator turns random noise into samples while a discriminator learns to tell real data from generated data, and each improves by exploiting the other's weaknesses. It starts with the minimax game that defines the objective and what its equilibrium means. Training process covers the alternating updates and why the generator's original loss saturates. Mode collapse explains the most common failure, a generator that produces only a few kinds of output. Popular variants surveys DCGAN, WGAN, and StyleGAN. The last topic of the chapter then moves from models to production.
Definition
A generative adversarial network trains two networks together: a generator that maps noise to samples , and a discriminator that outputs the probability that its input came from the real data. maximizes and minimizes the objective , whose optimum has .
In this topic
Minimax Game
A GAN replaces a hand-written loss with a learned critic. The discriminator outputs the probability that its input is real; the generator maps noise to a sample. maximizes , which is ordinary binary cross-entropy for real versus fake, while minimizes the same quantity. For a fixed , the best discriminator is , and plugging it in shows that minimizes the Jensen-Shannon divergence between the two distributions. At the global optimum , outputs 0.5 everywhere, and the objective equals .
The GAN objective has the global optimum . At equilibrium, and everywhere. The minimax value at optimality equals , and the divergence being minimized is the Jensen-Shannon divergence .
D outputs P(real). For real image, D(x)=0.9. For G's fake, D(G(z))=0.3. What's each term? Who's winning?
Training Process
GAN training alternates two gradient steps on each minibatch. First update on a batch of real samples labeled 1 and generated samples labeled 0. Then update , holding fixed, to make score its samples as real. The original paper allows discriminator steps per generator step and used . A trap appears early: when fakes are poor, is near 0, and the gradient of with respect to 's logit is just , nearly zero. The fix is the non-saturating loss: maximizes , which has the same fixed point but strong gradients when is losing.
GAN training alternates between steps of discriminator optimization and 1 step of generator optimization. The discriminator gradient is , maximizing classification accuracy. The generator gradient uses the modified objective instead of because the latter has vanishing gradients when is confident — the modified version provides stronger learning signal early in training.
D gets too strong (always outputs 0 for fakes). What happens to G's gradient?
Mode Collapse
Mode collapse is the most common GAN failure: the generator maps many different noise vectors to the same few outputs, covering only part of the data distribution. It happens because is updated against the current only; if one kind of sample fools today, every moves toward it. then learns to reject that kind, jumps to another mode, and the pair can cycle without converging. Each sample may look sharp, so per-sample quality hides the problem. Fixes include minibatch discrimination, where sees statistics across a whole batch, feature matching, unrolled GANs, and Wasserstein losses.
Mode collapse occurs when maps many different values to the same output region. Formally, the support of is much smaller than the support of . This is a Nash equilibrium failure: G finds a mode that maximally fools D, and D cannot distinguish it from real data in that mode. Minibatch discrimination adds a feature that measures diversity within a batch, penalizing generators that produce identical outputs.
MNIST GAN: G only outputs '1' digits (gets D(G(z))=0.5 for all). Why? How to detect?
Popular Variants
Variants fix specific weaknesses of the original GAN. DCGAN set the convolutional recipe: strided convolutions instead of pooling, batch normalization, ReLU in the generator, and LeakyReLU in the discriminator. WGAN replaces Jensen-Shannon divergence with the Wasserstein, or earth mover's, distance, estimated by a critic that must be 1-Lipschitz, which was first enforced by weight clipping and later by a gradient penalty in WGAN-GP. Its loss tracks sample quality, so it doubles as a training curve. StyleGAN feeds a mapped latent code into every layer, giving control over coarse and fine features; BigGAN scales batch size and class conditioning. Diffusion models have since overtaken GANs for most image generation.
WGAN replaces the JS divergence with the Wasserstein-1 distance , estimated via the Kantorovich-Rubinstein dual: . The Lipschitz constraint is enforced via weight clipping or gradient penalty . Wasserstein distance provides gradients everywhere, even when distributions have non-overlapping support — the key advantage over JS divergence.
WGAN uses Wasserstein distance instead of JS divergence. Why does this help training stability?
Theory Exercise
Problem:
The original GAN generator loss log(1 - D(G(z))) suffers from vanishing gradients early in training. Explain why, and describe the 'non-saturating' fix and why it provides a stronger learning signal.
Hints:
- Early on, the discriminator easily rejects fakes, so D(G(z)) is near 0
- Examine the gradient of log(1 - D(G(z))) as D(G(z)) -> 0
- Consider the alternative objective of maximizing log(D(G(z)))
Coding Exercise
Problem:
Demonstrate why the minimax GAN generator loss saturates by numerically comparing the gradient magnitude of log(1 - D(G(z))) versus the non-saturating -log(D(G(z))) as a function of the discriminator's output on fakes.
Hints:
- Treat D(G(z)) as a single scalar d in (0, 1) and use autograd to get gradients
- Compute both the saturating loss log(1 - d) and the non-saturating loss -log(d)
- Evaluate at small d (fakes easily detected) to see which loss still has a usable gradient
Related Problems on PixelBank
The previous topic, Generative Adversarial Networks, ended the modelling half of this chapter, and every model in this plan so far has been judged on a held-out test set. A test score is where the work starts, not where it ends. The model must be reproduced months later, served to real users under a latency budget, rolled back when a release goes wrong, and kept accurate as the world drifts away from the training data. Google engineers argued in 2015 that the model code is a small fraction of a real ML system and that most of the cost hides in the surrounding data, configuration, and serving machinery. How do we manage that machinery deliberately rather than by accident?
This topic, MLOps and Production, closes both the chapter and the plan. It covers four practices. Experiment tracking records every run's parameters, data, and metrics so results can be reproduced. A model registry versions trained models with their lineage and controls promotion. Model serving turns a model into a fast, scalable service. Monitoring and drift detection tell you when the deployed model is quietly getting worse.
Definition
MLOps is the set of practices that apply DevOps principles of versioning, automation, testing, and monitoring to machine learning systems across their full lifecycle: data preparation, training, evaluation, deployment, and maintenance. Unlike ordinary software, an ML system can fail when its data changes without any code change, so data and models are versioned and monitored alongside code.
In this topic
Experiment Tracking
A training run is defined by far more than its code: hyperparameters, data version, random seed, library versions, and hardware all change the result. Experiment tracking logs each run as a record of parameters, metrics over time, artifacts such as weights and plots, the code commit, and a hash of the training data. Tools such as MLflow, Weights and Biases, and Neptune store these records in a queryable database with dashboards for comparing runs. Without tracking, a good result found in week two often cannot be reproduced in week ten. Log automatically from the training script; manual notes in notebooks are incomplete exactly when they matter.
Experiment tracking logs the mapping where is the hyperparameters, the data version hash, the metrics, and the artifacts such as weights and configs. Tracking matters because search spaces are large: a grid of values for each of hyperparameters needs runs, 625 for 5 values of 4 hyperparameters. Random search instead lands in the top fraction of configurations with probability after trials, so 59 trials hit the top 5 percent with 95 percent probability, regardless of dimension. Bayesian optimization fits a Gaussian-process surrogate, costing for past runs, to choose each next trial.
Run 50 experiments varying lr, batch_size, architecture. 3 months later, which config gave best F1? How do you reproduce it?
Model Registry
A model registry is the version-controlled catalogue of trained models that are candidates for deployment. Each version records its lineage: the training run, code commit, data version, evaluation metrics, and who approved it. Versions move through stages such as staging, production, and archived, and serving systems load whatever version currently holds the production alias rather than a hard-coded file. That makes promotion and rollback a metadata change instead of a redeploy. Lineage also supports debugging and audits, since a bad prediction can be traced to the exact data and code behind it. Without a registry, teams lose track of which file is actually serving.
A model registry maintains a DAG (directed acyclic graph) of model lineage: . Versioning enables reproducibility: any model can be re-created from its lineage. The staging workflow (development -> staging -> production -> archived) ensures only validated models serve traffic. Canary deployments route of traffic to the new model and monitor metrics before full rollout.
Production model v2.3 has bug. Need to rollback. What metadata should registry have?
Model Serving
Serving turns a trained model into a dependable prediction service. Online serving answers single requests behind an API under a latency budget; batch serving scores large datasets on a schedule; edge serving runs on devices. Model servers such as TorchServe, Triton, TensorFlow Serving, and BentoML add dynamic batching, which waits a few milliseconds to group requests so the GPU works efficiently, as well as versioning and health checks. A load balancer spreads traffic across replicas and can split it between models for A/B tests or canary releases. Measure tail latency, p95 or p99, not the average, because the slowest requests are what users notice.
Serving latency decomposes as . For a transformer with layers, dimensions, and input length : . Dynamic batching groups requests arriving within a time window to amortize GPU kernel launch overhead — optimal batch size balances throughput () against latency ( for the first request in the batch). Quantization (INT8) reduces model size by and inference latency by - with accuracy loss.
Model needs: <50ms latency, 1000 req/s, A/B testing between 2 models. What serving architecture?
Monitoring & Drift
A deployed model degrades silently, because nothing crashes when the world changes. Data drift is a change in the input distribution , such as new merchant categories. Concept drift is a change in : the same features now mean a different outcome, as when fraudsters change tactics. Data drift can be measured on unlabeled traffic with statistics such as the Population Stability Index or a Kolmogorov-Smirnov test per feature; concept drift needs labels, which often arrive weeks late. Monitor input statistics, the prediction distribution, and delayed accuracy, alert on thresholds, and retrain on fresh data when drift is persistent rather than a one-day spike.
Data drift is detected by comparing feature distributions: Population Stability Index measures divergence between training () and serving () distributions, with indicating significant drift. Concept drift occurs when changes even if is stable — detectable only with labeled data. Monitoring tracks prediction distribution entropy: a sudden increase suggests the model is uncertain about new data patterns. Automated retraining triggers when drift exceeds thresholds for consecutive measurement windows.
Credit scoring model trained on 2023 data. In 2024, fraud patterns changed (new attack vectors). How do you detect this?
Theory Exercise
Problem:
Your deployed model's accuracy dropped from 95% to 85% over 3 months. What could cause this, and how would you detect and fix it?
Hints:
- Think about what changes over time
- Consider data distribution shifts
- What monitoring would help?
Coding Exercise
Problem:
Implement a lightweight data-drift check using the Population Stability Index (PSI). Compare a 'training' feature distribution against a shifted 'serving' distribution and flag drift when PSI crosses the standard 0.2 threshold.
Hints:
- Bin both distributions using the SAME bin edges (derived from the reference/training data)
- PSI = sum over bins of (p_serve - p_train) * ln(p_serve / p_train); add a tiny epsilon to avoid log(0)
- PSI < 0.1 = stable, 0.1-0.2 = moderate shift, > 0.2 = significant drift