2021 Clip

Keywords

ver. 1.0.0, 2021_clip

Learners will understand how CLIP trains joint image–text embeddings from web captions to enable efficient contrastive learning, zero‑shot classification, robustness under shift, and its practical limitations and risks.

This unit explains how training on raw image captions at web scale produces versatile image–text embeddings (CLIP). You will learn why captions are a stronger supervision signal than fixed labels, how the WebImageText (400M pairs) dataset was constructed, and how CLIP’s symmetric contrastive objective and dual encoders (image and text projected into a shared space) are trained efficiently. The unit shows how to build zero‑shot classifiers from class names using prompt engineering and ensembling, how to evaluate embeddings with linear probes and few‑shot transfer, why zero‑shot CLIP is more robust to distribution shifts, where CLIP fails compared to humans, and what biases and deployment risks arise from uncurated web data. Finally, it connects CLIP to guided image diffusion.

This unit teaches the principles, practice, and limits of CLIP-style contrastive pretraining so you can both build and critically evaluate image–text models.

Why raw captions? You will be able to explain why free-form captions from the web provide a far richer and more scalable supervision signal than training on a closed set of class labels: captions capture nuanced, open‑vocabulary semantics and remove the bottleneck of fixed taxonomies, enabling models to generalize beyond the classes seen during training.

Dataset scale and construction. You will understand how the WebImageText corpus was assembled (≈400 million image–text pairs gathered via roughly 500k search queries with per‑query balancing) and why that scale and diversity matter for learning broadly useful multimodal features.

The training objective and efficiency. You will be able to state and interpret CLIP’s symmetric contrastive objective: for a batch of N image–text pairs, compute similarity scores for all N×N image–text combinations (via cosine similarity) and maximize the score of each correct pairing against the N^2−N negatives on both image→text and text→image directions. You will also be able to explain why this contrastive task is much cheaper than full caption prediction and why it yields strong cross‑modal alignment.

Model architecture and scaling. You will be able to describe the dual‑encoder architecture: a vision encoder (modified ResNet with attention pooling or a Vision Transformer) and a Transformer text encoder, each followed by a linear projection into a shared multimodal embedding space. You will understand how both sides are scaled and trained to produce comparable embeddings.

Zero‑shot classifiers via prompts. You will be able to construct zero‑shot classifiers by embedding class name prompts and comparing image embeddings via cosine similarity, and you will know how prompt engineering and ensembling of multiple textual templates reliably increase zero‑shot accuracy with no extra image supervision.

Evaluating representations. You will be able to evaluate CLIP as a representation learner using linear probes (freeze the encoder, train a linear classifier) and few‑shot transfer, and read comparative results against supervised and self‑supervised baselines across many datasets.

Robustness to distribution shift. You will be able to explain effective robustness: how not training on a specific dataset can reduce overfitting to its idiosyncrasies, and why zero‑shot CLIP recovers a large fraction of the robustness gap on natural distribution‑shift benchmarks.

Failure modes and sample efficiency. You will be able to identify CLIP’s failure cases (tasks where it performs near random), assess overlap between pretraining and evaluation data, and compare CLIP’s sample efficiency to human performance on specific tasks (e.g., where humans learn from far fewer examples).

Biases and deployment risks. You will be able to assess the kinds of societal biases CLIP inherits from uncurated web data (demographic misclassification, gendered associations) and articulate deployment risks introduced by an open‑ended zero‑shot interface (e.g., misuse in surveillance or biased decision pipelines).

Connections and next steps. Finally, you will be able to summarize what contrastive multimodal pretraining provides—scalable alignment, strong zero‑shot transfer, improved robustness—and how CLIP’s shared embedding space enables downstream techniques like CLIP‑guided image diffusion for text‑conditioned generation.

By the end of the unit you will both know how to train and use CLIP‑style models and be equipped to evaluate their performance, robustness, limitations, and ethical risks in practice.

Materials

Source document

  • Radford, Alec, et al. “Learning transferable visual models from natural language supervision.” International conference on machine learning. PmLR, 2021.