Chapter 17: Variational autoencoders

Keywords

ver. 1.0.0, chapter_17_variational_autoencoders

How to build, train and use variational autoencoders: ELBO, encoder/decoder, reparameterization, and practical issues.

A compact, practical introduction to variational autoencoders (VAEs): latent-variable modelling by marginalising joint distributions; a continuous nonlinear latent model defined by a decoder network; why maximum likelihood is intractable and how the evidence lower bound (ELBO) is derived and interpreted; amortised variational inference with an encoder; the reparameterization trick for gradient-based learning; the VAE training algorithm and loss; using VAEs to generate, estimate probabilities, and manipulate data; and common extensions and failure modes (disentanglement, blurriness, posterior collapse, holes).

This unit teaches how to model complex data densities with latent-variable models and trains them as variational autoencoders (VAEs). You will learn to view a high-dimensional data distribution as the marginal of a joint distribution over observed data x and latent variables z, first seeing the framework in the finite, tractable case of a mixture of Gaussians and then moving to a continuous, nonlinear latent-variable model where a neural decoder maps z to the parameters of p(x|z).

Because the marginal likelihood p(x) = ∫ p(x,z) dz is intractable for such models, you will derive the evidence lower bound (ELBO) on the log marginal using Jensen’s inequality, understand that the ELBO equals log p(x) minus the KL divergence between an auxiliary q(z|x) and the true posterior, and read the ELBO as a reconstruction term minus a KL regulariser to the prior. Restricting q to a Gaussian family and letting a single encoder network predict its parameters yields an amortised variational approximation that can be optimised across a dataset.

You will assemble encoder and decoder into a VAE: compute encoder outputs μ(x), Σ(x), sample z, decode to reconstruct x, and optimise a loss that combines a Monte Carlo estimate of expected reconstruction and a closed-form Gaussian KL. To backpropagate through the stochastic latent layer you will apply the reparameterization trick, rewriting sampling as z = μ + Σ^{1/2} ε with ε ∼ N(0,I), moving randomness into an auxiliary input and enabling efficient gradient-based learning.

After training, you will use the VAE to generate samples by drawing z from the prior and decoding, to estimate data likelihoods via importance sampling, and to resynthesise or interpolate data by manipulating latent codes (including spherical interpolation). Finally, you will explore loss modifications to encourage disentangled latents (for example β-VAE and total-correlation penalties) and recognise common failure modes of the basic VAE such as blurry outputs, posterior collapse, and holes in latent space.

By the end, you will be able to: explain how latent-variable models represent densities by marginalisation; define and sample from a continuous nonlinear decoder-based latent model; derive and interpret the ELBO and know when it is tight; justify and implement amortised Gaussian variational inference with an encoder; implement the reparameterization trick; write down the VAE training algorithm and loss; use a trained VAE for generation, likelihood estimation and latent editing; and apply or criticise extensions designed for disentanglement and robustness.

Materials

Source document

  • Understanding Deep Learning, Simon Prince, 2026 — Link — Page 341-362