CLIP Guided Image Diffusion
ver. 1.0.0, clip_guided_image_diffusion
Guiding diffusion models with classifiers and CLIP for high‑quality text‑to‑image generation, measuring and improving sample quality, enabling inpainting/editing, and weighing safety and release trade‑offs.
This unit explains how to steer diffusion-based image generators with gradients from classifiers and CLIP to produce class‑ and text‑conditioned images. It covers the standard sample‑quality metrics, U‑Net architectural improvements that raise unconditional quality, the mathematical derivation of classifier guidance (and its DDIM deterministic form), and how scaling the guidance trades diversity for fidelity. The unit then describes GLIDE’s text conditioning, noise‑aware CLIP guidance, comparisons with classifier‑free guidance and GANs, text‑guided inpainting fine‑tuning, and the safety considerations around releasing text‑to‑image models.
This unit builds a complete, practical picture of guided diffusion for conditional image synthesis. It begins with the evaluation tools used to judge generation: FID (and sFID), Inception Score, and precision/recall, and shows how U‑Net architectural ablations can substantially improve unconditional diffusion quality before any guidance is applied.
It then derives classifier guidance: conditioning the reverse diffusion transition on a label shifts the Gaussian mean by an amount proportional to the classifier gradient, and the same intuition yields a DDIM deterministic counterpart. Introducing an explicit scalar on that gradient makes the fidelity–diversity tradeoff controllable: multiplying the gradient by s is equivalent to sharpening the conditional p(c|x) to concentrate samples on higher‑confidence modes.
With guidance in hand, the unit compares diffusion results to GANs (BigGAN‑deep) across resolutions and shows how guided diffusion combined with DDIM sampling and upsampling stacks pushed diffusion ahead on image quality—at the cost of larger models, slower sampling, and reliance on auxiliary guidance models. Moving from class labels to open text, the unit presents GLIDE: a text‑conditional design that feeds a transformer text encoder into a large diffusion U‑Net via AdaGN‑style modulation and cross‑attention at each layer, enabling free‑form caption conditioning.
To generalize guidance from closed label sets to arbitrary captions, CLIP replaces the classifier: guidance uses the image embedding gradient against a text embedding. Because CLIP was trained on clean images, it must be made noise‑aware (retrained or adapted on noisy images) for effective guidance. Head‑to‑head comparisons show a surprising result: optimizing the CLIP score directly does not guarantee the best human judgments—classifier‑free guidance (which removes the need for an external classifier) often wins on both photorealism and caption similarity, revealing that the CLIP objective can be misleading.
The unit also shows how to turn a generator into an editor: fine‑tuning a text‑conditional diffusion model with mask and masked‑image channels produces effective text‑guided inpainting and localized editing, outperforming naively replacing pixels during sampling. Finally, it examines the safety and release decision for text‑to‑image models—how filtering and retraining affect biases and abuse risks, what GLIDE (filtered) achieves, and which failure modes remain—and closes by consolidating the techniques (architecture, classifier guidance, gradient scaling, text conditioning, CLIP guidance, inpainting, safety) and pointing to the next development: classifier‑free diffusion guidance that eliminates the external guidance model.
Materials
Source documents
- Dhariwal, P., & Nichol, A. (2021). Diffusion models beat gans on image synthesis. Advances in neural information processing systems, 34, 8780-8794.
- Nichol, Alex, et al. “Glide: Towards photorealistic image generation and editing with text-guided diffusion models.” arXiv preprint arXiv:2112.10741 (2021).