2022 Classifier Free Diffusion Guidance

Keywords

ver. 1.0.0, 2022_classifier_free_diffusion_guidance

How to remove external classifiers and use classifier-free guidance to trade diversity for fidelity in diffusion models

This unit teaches how to express diffusion sampling in score form, reinterpret classifier guidance as an additive score term, and eliminate the need for an external classifier by training a single model that can be used both conditionally and unconditionally. Learners will apply classifier-free guidance at sampling time, sweep the guidance weight to observe the fidelity–diversity tradeoff on ImageNet, and pick training and sampling hyperparameters (unconditional training probability and step count). The unit also explains why the gains are generative rather than adversarial and lists the method’s costs.

The unit shows how diffusion generation is controlled through score estimates rather than raw noise predictions, using continuous time indexed by log signal-to-noise ratio (log-SNR). It derives the relationship between the usual noise-prediction parameterization and the score ∇_{z_λ} log p(z_λ), and rewrites classifier guidance as an additive modification to the model’s score: guided sampling can be written as ε̂θ = ε_θ(z_λ, c) − w·σ_λ·∇{z_λ} log p_φ(c | z_λ), i.e., as sampling from a density proportional to p(z_λ | c)·p_φ(c | z_λ)^w.

Rather than relying on a separately trained classifier, the unit teaches classifier-free guidance: a single network is trained to output both conditional and unconditional score estimates by randomly replacing the conditioning with a null token with probability p_uncond. At sampling time the two outputs are combined via extrapolation, e.g. _θ = (1+w)·ε_θ(z_λ,c) − w·ε_θ(z_λ), producing a controllable interpolation/extrapolation between unconditional and conditional generations.

Practical experimentation is emphasized: sweeping the guidance strength w on ImageNet (64×64 and 128×128) demonstrates the typical fidelity–diversity curve — Inception Score tends to increase monotonically with w while FID improves up to a small nonzero w and then degrades — so tuning w lets you trade sample fidelity for diversity. The unit also fixes remaining knobs: p_uncond ≈ 0.1–0.2 is empirically effective, and the number of sampling steps interacts with guidance strength in predictable ways, so both should be chosen together for best results.

Finally, the unit addresses interpretive concerns and costs. It argues that classifier-free guidance produces genuine generative quality improvements rather than merely attacking evaluation metrics: the method is not equivalent to gradient ascent on a classifier because an unconstrained model’s score estimate need not form a conservative vector field. Costs are also listed: increased sampling bias toward modes, reduced diversity, the need to tune w and p_uncond, and potential sensitivity to sampling schedule and step count. The unit situates classifier-free guidance in the broader generative-models arc, connecting score-based diffusion sampling back to earlier topics (e.g., VAEs) and forward to controllable text-to-image generation.

Materials

Source document

  • Ho, Jonathan, and Tim Salimans. “Classifier-free diffusion guidance.” arXiv preprint arXiv:2207.12598 (2022).