Chapter 18: Diffusion models
ver. 1.0.0, chapter_18_diffusion_models
Practical and theoretical introduction to diffusion generative models: forward noising, learned reverse, ELBO derivation, noise-prediction loss, implementation for images, fast sampling, and conditioning/cascades
This unit teaches diffusion-based generative modeling: how a fixed forward noising chain creates Gaussian latents, how to collapse it into a closed-form diffusion kernel, how to learn a reverse denoising chain supervised by tractable conditionals, and how the ELBO reduces to per-step KLs that, with a reparameterization, become a simple noise-prediction squared-error loss. It covers the training and ancestral sampling loops, implementing image models with a time-conditioned U-Net, accelerating generation with deterministic/ODE/step-skipping techniques, and conditioning (classifier/CLIP guidance and classifier-free) plus cascaded super-resolution.
This unit develops diffusion models from first principles to practical image generation. It starts from the idea of a fixed forward Markov encoder that gradually corrupts data with Gaussian noise under a variance (noise) schedule, then shows how composing those steps yields a closed-form diffusion kernel q(z_t|x) that reaches an isotropic Gaussian at large t. Building a learned reverse chain, the unit identifies the tractable Gaussian conditional q(z_{t-1}|z_t,x) that can supervise each decoder step and frames learning as matching the model backwards transitions to these conditionals.
The ELBO for the full multi-step latent sequence is derived and simplified into a sum of per-timestep KL divergences (plus a reconstruction term). Reparameterizing the target and absorbing arithmetic into the network yields the common noise-prediction objective: a sum of squared errors between the network output and the Gaussian noise added at a sampled timestep. From this form you get the straightforward stochastic training loop (sample datapoint, timestep, noise, form z_t, take a gradient step) and the ancestral sampling loop that iteratively denoises from z_T ~ N(0,I) back to data space.
Turning theory into practice for images, the unit shows how a single U-Net shared across timesteps is time-conditioned (via sinusoidal embeddings) to implement the noise-predictor g_t. It then confronts the computational cost of hundreds of sampling steps and explains acceleration strategies: DDIM-style deterministic non-Markovian updates, probability-flow ODE formulations, and step-skipping methods that cut sampling to tens of steps while preserving sample quality.
Conditioning techniques are covered so generation can be steered by class labels or text: classifier guidance and classifier-free guidance are explained and compared. Scaling to high resolution is handled with cascaded pipelines that generate a low-res sample and then super-resolve it. The unit ties these pieces together and points toward using learned text encoders (e.g., CLIP) to provide the semantic conditioning needed for text-to-image models.
By the end, a learner understands the forward noising process and variance schedule, can use the closed-form diffusion kernel to jump to arbitrary noise levels, can derive and simplify the diffusion ELBO into per-step KLs, and can reparameterize to obtain the noise-prediction loss. They can implement training and ancestral sampling, build an image diffusion model with a time-conditioned U-Net, apply fast deterministic or ODE-based samplers, and condition and cascade models for high-quality, high-resolution controlled generation.
Materials
Source document
- Understanding Deep Learning, Simon Prince, 2026 — Link — Page 363-387