Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, Ilya Sutskever
State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories. This restricted form of supervision limits their generality and usability since additional labeled data is needed to specify any other visual concept. Learning directly from raw text about images is a promising alternative which leverages a much broader source of supervision. We demonstrate that the simple pre-training task of predicting which caption goes with which image is an efficient and scalable way to learn SOTA image representations from scratch on a dataset of 400 million (image, text) pairs collected from the internet. After pre-training, natural language is used to reference learned visual concepts (or describe new ones) enabling zero-shot transfer of the model to downstream tasks. We study the performance of this approach by benchmarking on over 30 different existing computer vision datasets, spanning tasks such as OCR, action recognition in videos, geo-localization, and many types of fine-grained object classification. The model transfers non-trivially to most tasks and is often competitive with a fully supervised baseline without the need for any dataset specific training. For instance, we match the accuracy of the original ResNet-50 on ImageNet zero-shot without needing to use any of the 1.28 million training examples it was trained on. We release our code and pre-trained model weights at https://github.com/OpenAI/CLIP.
State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories. This res...
This abstract introduces a fundamentally different approach to training computer vision systems. Instead of the traditional pipeline—gather labeled images, define fixed categories, train a classifier—this paper proposes learning from natural language supervision. This is revolutionary because it sidesteps the bottleneck of needing manually labeled datasets for every new task.
Think of it this way: Traditional computer vision is like a student who memorizes answers to a specific exam. CLIP is like a student who learns to understand concepts deeply, so they can answer any exam question they encounter.
Standard computer vision systems (pre-2021) followed this pattern:
The limitation: If you want to recognize a new concept not in your original categories, you need to:
Instead of predicting fixed categories, learn from image-caption pairs. The key observation:
Images naturally occur alongside text on the internet. Why not use that free supervision?
Dataset scale: 400 million (image, text) pairs—orders of magnitude larger than traditional labeled vision datasets.
The pre-training task is deceptively simple: Predict which caption matches which image
More formally, given a batch of image-caption pairs , the model learns to:
This is a contrastive learning objective, though the abstract doesn't explicitly name it.
Once trained, the model has learned visual concepts connected to language. Here's the magic:
Traditional approach for a new task:
CLIP approach:
Why this works: The model has learned that visual patterns in dog photos align with the linguistic concept "dog." It never saw a labeled dog image during pre-training, but it learned the connection through internet-scale text-image correlation.
The authors evaluate on 30+ vision datasets covering diverse tasks:
| Task Category | Examples |
|---|---|
| Text recognition | OCR (Optical Character Recognition) |
| Video understanding | Action recognition in videos |
| Geographic reasoning | Geo-localization |
| Fine-grained classification | Distinguishing dog breeds, car models, etc. |
The abstract highlights one striking result:
"We match the accuracy of the original ResNet-50 on ImageNet zero-shot"
Let's unpack what this means:
ResNet-50 (traditional approach):
CLIP (zero-shot approach):
This is remarkable because CLIP achieved equivalent performance without task-specific training data.
While the abstract doesn't provide equations, here's the underlying mathematical framework:
The pre-training minimizes a contrastive loss. For a batch of pairs:
Where:
Interpretation: The loss is minimized when:
For a new task with class names , predict the class of image as:
No gradient updates, no fine-tuning—just compare embeddings.
| Advantage | Explanation |
|---|---|
| Generality | Works for any visual concept describable in natural language |
| Scalability | Leverages 400M freely available internet image-text pairs |
| Zero-shot capability | Performs new tasks without retraining or task-specific data |
| Language interface | Humans can describe tasks naturally without technical expertise |
| Compositionality | Can combine concepts: "a red dog" or "a car in the snow" |
The CLIP abstract describes a seismic shift: moving from task-specific supervised learning (memorize answers to one exam) to universal visual understanding via language (understand concepts deeply). By training on 400M image-caption pairs to predict which caption matches which image, CLIP learns visual representations so semantically rich that they match purpose-built classifiers—without ever seeing labeled training data for the target task.
This is why the paper is foundational: it showed that natural language supervision at scale is more effective than traditional labeled datasets for learning transferable visual knowledge.
Pre-training methods which learn directly from raw text have revolutionized NLP over the last few years. Task-agnostic o...
This introduction section is making a crucial argument: Why should we learn image representations from raw text instead of using traditional labeled datasets? The authors draw inspiration from a major revolution that happened in Natural Language Processing (NLP) and ask: can we do the same thing in computer vision?
Think of it like this:
The introduction builds a case for why the new way might be better. Let's break down how.
The authors start by noting that NLP was transformed by pre-training methods that learn directly from raw text. Key examples:
The authors make an important observation:
"These results suggest that the aggregate supervision accessible to modern pre-training methods within web-scale collections of text surpasses that of high-quality crowd-labeled NLP datasets."
What does this mean? The implicit structure in billions of web documents provides more useful supervision signal than millions of carefully hand-labeled examples. This is counterintuitive but powerful:
GPT-3 and similar models use a standardized approach: treat everything as a text input → text output problem. This means:
The authors ask: Can we replicate this NLP success in computer vision?
Currently, the standard practice is still to pre-train on ImageNet—a crowd-labeled dataset with ~1.28 million images in ~1,000 categories. But why not learn from web-scale text instead, like NLP does?
The obvious answer: "Yes, prior work tried this. It didn't work well."
The authors review 20+ years of work trying to learn image representations from text. Here's the conceptual progression:
Early work realized you could train image models by predicting text associated with images:
The idea is simple: if you can predict which words go with which images, you've learned something about what images represent.
Rather than isolated words, use richer text:
N-grams are important because "dog running" contains richer information than just "dog" alone. Mathematically, if we're predicting words in a caption, an n-gram model predicts sequences rather than individual words.
**The key limitation**: Li et al. achieved only **11.5% accuracy on ImageNet zero-shot**. Remember: ResNet-50 (fully supervised) gets ~76% accuracy. So this approach was ~6.5× worse! #### **Stage 3: Transformer-based approaches (2020)** More recent work uses modern architectures: - **VirTex (Desai & Johnson, 2020)**: Transformer language modeling on images - **ConVIRT (Zhang et al., 2020)**: Contrastive learning from image-text pairs - **ICMLM (Bulent Sariyildiz et al., 2020)**: Masked language modeling approach These showed promise as "proofs of concept," but **performance still lagged significantly**. ### The problem: SCALE The crucial insight is stated explicitly: > "A crucial difference between these weakly supervised models and recent explorations of learning image representations directly from natural language is scale." All the previous approaches used relatively small datasets. You can't learn good representations from millions of image-text pairs when you're competing against models trained on billions of web documents. --- ## PART 4: THE SOLUTION - CLIP ### What the authors did differently 1. **Collected a massive dataset**: 400 million (image, text) pairs from the internet 2. **Used a proven architecture**: Simplified version of ConVIRT 3. **Called it CLIP**: **C**ontrastive **L**anguage-**I**mage **P**re-training ### How CLIP differs from traditional vision models **Traditional approach**: - Input: Image $\mathbf{x}$ - Model learns: Image encoder that extracts features, plus a linear classifier - Training objective: Predict a fixed set of $K$ predefined categories - Output: Probability distribution over $K$ classes **CLIP's approach**: - Input: A batch of $N$ (image, text) pairs: $\{(\mathbf{x}_i, \mathbf{t}_i)\}_{i=1}^{N}$ - Model learns: Image encoder AND text encoder (both as neural networks) - Training objective: Match images with their correct captions (a contrastive task) - Output: At test time, use the text encoder to create classifiers from class descriptions ### The contrastive learning setup Let me explain what "contrastive" means mathematically. Given a batch of $N$ paired examples: \text{Batch} = \{(\mathbf{x}_1, \mathbf{t}_1), (\mathbf{x}_2, \mathbf{t}_2), \ldots, (\mathbf{x}_N, \mathbf{t}_N)\}Where:
The model learns two encoding functions:
The learning objective works like this:
This is "contrastive" because the model learns by contrasting correct pairs against incorrect ones.
When you have 400 million examples instead of millions:
Here's the magic: at test time, you don't need labeled training data for new tasks.
Example: Want to classify "dogs" vs "cats"?
No retraining needed. No labeled examples needed. This is zero-shot transfer.
In the paper's results, they show this works remarkably well:
The figure shows the conceptual difference:
Left side (traditional):
Right side (CLIP):
At test time, you synthesize a classifier by encoding class descriptions in text.
This introduction makes the case: Scale + natural language supervision = a better path forward than hand-labeled datasets.
The authors acknowledge that small-scale attempts failed, but they hypothesize that the bottleneck wasn't the approach—it was the data scale. By using 400 million image-text pairs, they can finally make this work as well as (or better than) traditional supervised learning.
At the core of our approach is the idea of learning perception from supervision contained in natural language. We emphas...
This section is the conceptual foundation of the entire CLIP paper. While previous sections explained what CLIP does (learn from image-text pairs), this section explains why this approach is fundamentally better than existing alternatives. The authors are making a philosophical argument: natural language is a better supervisory signal than the traditional labels used in computer vision.
Think of it this way: Traditional computer vision models learn from labels like "cat," "dog," "car"—a fixed set of categories decided in advance. But CLIP learns from the rich, diverse language that naturally appears alongside images on the internet. This is a paradigm shift in how we think about supervision.
The authors define their approach by emphasizing that all their techniques are unified by one principle: using natural language text as the training signal rather than hand-annotated categorical labels.
Let's formalize this conceptually:
Traditional supervised learning: We have a dataset where each image has a single label from some discrete set of categories.
Natural language supervised learning: We have a dataset where each image is paired with a text string (a caption or description containing arbitrary words and phrases).
The authors acknowledge this isn't entirely new. Going back 20+ years, researchers have explored learning from captions and associated text. However, there's a crucial evolution:
The key pattern: As our tools for understanding language improved (from n-grams → word embeddings → contextual representations), so did the ability to extract supervision from natural language.
The authors present two major advantages. Let me break these down:
The Problem with Traditional Labels:
In classical computer vision (ImageNet, CIFAR-10, etc.), you need:
This is expensive and doesn't scale beyond the categories you pre-defined.
Why Natural Language Supervision is Different:
The internet contains billions of images with naturally associated text (captions, descriptions, alt-text, surrounding text, etc.). This text wasn't created for machine learning—it was created by humans describing images for other humans. The authors call this "passive supervision" because:
You don't need to ask annotators to fit descriptions into "machine learning compatible format"—they're already there in natural form. A caption like "a golden retriever playing fetch in the park" is more informative than a single label "dog."
This is the crucial insight that differentiates natural language supervision from self-supervised learning (like contrastive learning on images alone).
The Problem with Representation Learning Without Language:
If you learn image representations using only image pairs (e.g., "this image and its augmentation are similar"), you get good representations but they're disconnected from language. To use them for a new task, you still need labeled examples to train a classifier.
Why Language Connection Matters:
When you learn from image-text pairs, your representation space becomes grounded in language. This means:
Mathematically, this is possible because both images and text are embedded in a shared space. If we denote:
Then for a new category described as text , we can compute:
where denotes cosine similarity. No training required—just the pre-trained encoders!
Early work on learning from language struggled because they used weak language representations:
Example: The n-gram approach couldn't distinguish between "dog" and "cat" beyond frequency patterns.
The rise of Transformers and contextual language models (like BERT, GPT) changed everything. These models understand:
This means natural language can now carry rich supervisory information that translates into visual concepts automatically.
[Figure 2 in the paper shows a comparison of three approaches:]
This empirically validates a key theoretical point: not all ways of using language supervision are equally efficient. The specific method matters enormously—but the core advantage of using language at all remains clear.
| Argument | Why It Matters |
|---|---|
| Scale | We can learn from internet-scale data without expensive annotation |
| Language Grounding | Our learned representations connect to language, enabling zero-shot transfer |
| Timing | Modern deep language models finally give us the tools to extract rich supervision from text |
The section sets up why CLIP's approach—learning image representations from natural language supervision at scale—is both feasible and powerful.
Existing work has mainly used three datasets, MS-COCO (Lin et al., 2014), Visual Genome (Krishna et al., 2017), and YFCC...
This section addresses a fundamental bottleneck in the CLIP research: data availability. The paper claims to learn from natural language supervision at scale, but existing image-caption datasets are tiny by modern deep learning standards. This section explains why the authors couldn't use existing datasets and how they solved the problem by building their own massive dataset called WIT (WebImageText).
Think of it this way: imagine you're trying to prove that learning from internet text is powerful—but you only test on tiny, carefully curated datasets. You'd never know if the approach actually scales. This section ensures the authors have enough data to properly test their hypothesis.
Let me break down the landscape of available datasets at the time:
| Dataset | Size | Quality Issue |
|---|---|---|
| MS-COCO | ~100K photos | Small by modern standards |
| Visual Genome | ~100K photos | Small by modern standards |
| YFCC100M | 100M photos | Sparse, low-quality metadata |
The key insight about YFCC100M is particularly important. Let me explain the filtering problem mathematically:
Original dataset size: 100 million images
After filtering: 15 million images
The reduction factor is
This means roughly 83% of the images had to be discarded because their accompanying text was either non-English, missing, or too sparse to be useful. This is a dramatic loss of data.
The authors make a crucial observation: even after filtering, YFCC100M (15M images) is only comparable to ImageNet (~1.28M labeled training images mentioned in the abstract, plus additional unlabeled images). This defeats the purpose of using web-scale supervision.
Rather than be limited by existing datasets, the authors constructed a new dataset from scratch. Here's how:
Scale target: 400 million pairs
The construction involves several key components:
Let me formalize what they did:
Let's denote:
The constraint they applied was:
The total dataset WIT is constructed as:
The theoretical maximum size would be:
However, they report collecting 400 million pairs, suggesting:
This means they're sampling well below the theoretical maximum. This could occur because:
The key innovation is using diverse, broad queries rather than carefully curated annotations. This is qualitatively different from ImageNet, which has a fixed set of predetermined categories.
The authors make an important comparison to validate their dataset:
"The resulting dataset has a similar total word count as the WebText dataset used to train GPT-2."
This is a clever validation metric. Let me explain why it matters:
GPT-2's training dataset (WebText):
WIT's properties:
The intuition is: if GPT-2 learned powerful language representations from this amount of text, then CLIP should have access to equally rich linguistic supervision when paired with images.
This ties back to the motivation in Section 1. The paper argues that:
Without a dataset like WIT (400M pairs), this question cannot be properly answered. Testing on MS-COCO or Visual Genome would be like asking "does scaling language models help?" but only testing on small supervised datasets—you'd get a misleading answer.
| Aspect | Existing Datasets | WIT |
|---|---|---|
| Largest usable size | 15M (after filtering) | 400M |
| Scale increase | 1× baseline | ~27× larger |
| Data source | Crowd-labeled | Diverse internet sources |
| Supervision type | Fixed categories | Natural language |
| Text quality | Minimal | Web-scale text diversity |
The bottom line: To properly test whether natural language supervision can match or beat traditional supervised learning, you need data at a comparable scale. WIT provides that scale, enabling the zero-shot transfer results discussed in later sections and the abstract.
State-of-the-art computer vision systems use very large amounts of compute. In the course of our efforts, we found train...
This section is crucial because it explains how CLIP actually works and why the authors chose this particular approach over alternatives. The paper's title promises "learning transferable visual models from natural language supervision," but there are multiple ways to do that. This section answers: Given that we have 400 million image-text pairs, what's the most efficient way to learn from them?
The key insight is that efficiency matters at scale. When you're training on 400 million examples, even small improvements in learning speed multiply out to massive computational savings. The authors found that a simple change in the training objective—from trying to predict exact words to just matching images with their text descriptions—made the model learn 12 times faster (3x from switching to bag-of-words, then 4x more from switching to contrastive learning).
The first approach was straightforward:
Why does this fail? Consider what "text accompanying an image" actually looks like on the internet:
The wide variety of valid descriptions makes exact word prediction fundamentally hard. The model has to learn language modeling as a side effect of learning vision, which is inefficient.
Recent work in computer vision had discovered something important: contrastive objectives learn better representations than predictive objectives.
What does "contrastive" mean? Instead of asking "what exact words go with this image?", ask a simpler question: "which of these text samples actually goes with this image?"
This is easier because:
Given a batch of (image, text) pairs, let's define:
Both encoders project their outputs into a shared multi-modal embedding space (imagine a space where both images and text are represented as vectors). Let's denote:
Both and are vectors (typically 512 or 1024 dimensional).
CLIP computes the cosine similarity between all possible pairs of image and text embeddings within a batch. Define the similarity between image and text as:
where denotes the dot product and denotes the L2 norm (Euclidean length).
This creates an similarity matrix :
Ideally:
Before computing the loss, CLIP applies a temperature scaling parameter (tau). The scaled similarities are:
The temperature parameter controls how "sharp" the distinctions are:
In CLIP, is learned during training as a log-parameterized multiplicative scalar (they optimize directly, which ensures stays positive).
CLIP optimizes a symmetric cross-entropy loss. This has two components:
Image-to-text matching loss: For each image, predict which of the texts in the batch is the correct one:
This is a standard cross-entropy loss where:
The log and fraction together create a softmax that outputs a probability distribution over texts.
Text-to-image matching loss: Similarly, for each text, predict which of the images is correct:
Note that this loss looks at rows instead of columns of the similarity matrix.
Total loss: Average both directions:
The "symmetric" in "symmetric cross-entropy" refers to this bidirectional matching.
Think of it as a matching game:
This is much easier than generating language (predictive) because:
The authors state: "We train CLIP from scratch without initializing the image encoder with ImageNet weights or the text encoder with pre-trained weights."
This is significant because:
Both encoders use only a linear projection to map to the shared embedding space:
where and are learned linear transformation matrices.
Why linear instead of something more complex? Simplicity and efficiency. The complex feature learning happens in the CNN and Transformer; the embedding space just needs to be a common coordinate system.
Only "a random square crop from resized images" is used. This is surprisingly minimal compared to typical vision training. Why?
At scale (400 million examples), the diversity of natural images already provides regularization. You don't need aggressive augmentation; instead, the model naturally sees countless variations of each visual concept.
The pseudocode in Figure 3 shows the essential algorithm. The key steps are:
The paper reports dramatic efficiency improvements:
| Approach | Relative Learning Speed |
|---|---|
| Transformer language model (predicts words) | 1× (baseline) |
| Bag-of-words (predicts words) | 3× |
| CLIP (contrastive objective) | 12× (3× × 4×) |
This 12× speedup means what previously took 12 hours of training now takes 1 hour. At the scale of their 400 million example dataset, this is enormous in terms of computational cost and environmental impact.
This section sets up why CLIP can do zero-shot transfer so effectively. Because:
This will be explored more in later sections, but the training objective here already ensures that the text encoder can handle arbitrary text descriptions, which is what makes zero-shot transfer possible.
Key takeaways:
We consider two different architectures for the image encoder. For the first, we use ResNet-50 as the base architecture ...
Before training CLIP on 400 million image-text pairs, the authors need to decide:
This is crucial because the choice of architecture directly impacts:
Think of this like choosing the blueprint for a building before construction—the design determines both cost and functionality.
The authors start with ResNet-50, a well-established convolutional neural network (CNN) architecture that was introduced in 2015 and has become a standard baseline in computer vision.
Key modifications they make:
Why replace global average pooling with attention pooling?
In standard ResNet:
With attention pooling:
where are learned attention weights that sum to 1 (usually computed via softmax), and represents the feature vector at spatial position
The authors also experiment with Vision Transformer (ViT), a newer architecture that processes images as sequences of patches rather than using convolutional operations.
How ViT differs from ResNet:
Minor modification in CLIP:
The motivation for using ViT is that transformers have shown superior scaling properties (they improve more efficiently with additional compute) compared to CNNs like ResNet.
The text encoder is a Transformer with the following specifications:
| Specification | Value | Explanation |
|---|---|---|
| Parameter count | 63 million | Total number of learnable weights |
| Depth | 12 layers | Number of transformer blocks stacked sequentially |
| Width | 512 | The hidden dimension (size of internal representations) |
| Attention heads | 8 | Number of parallel attention mechanisms in each transformer layer |
Text preprocessing pipeline:
Tokenization: The text is converted to a byte pair encoding (BPE) representation
Sequence length cap: Maximum sequence length is capped at 76 tokens
Special tokens: The text sequence is bracketed with:
[SOS] (Start of Sequence) token at the beginning[EOS] (End of Sequence) token at the endFeature extraction: The output representation is extracted from the [EOS] token
[EOS] position are usedWhy use the [EOS] token representation?
In transformer-based models, the special token at the end of a sequence typically accumulates contextual information from the entire input. It functions as a learned summary of the whole sequence, making it an ideal choice for the overall text representation.
This is where the authors address a critical practical question: If you have more compute available, how should you allocate it?
The authors reference EfficientNet (Tan & Le, 2019), which discovered that naive approaches to scaling don't work well.
Naive scaling approaches (less effective):
Why is this suboptimal? Each dimension affects the model differently:
Compound scaling (more effective):
For ResNet-based image encoders, CLIP adopts this compound scaling strategy:
This means when creating larger versions of CLIP, they build bigger ResNet encoders that are deeper, wider, and accept higher-resolution images.
Here's a crucial empirical finding from the authors:
"CLIP's performance is less sensitive to the capacity of the text encoder"
This means the bottleneck is the image encoder, not the text encoder. Therefore:
Why this asymmetry?
The text modality is inherently simpler than vision:
A relatively modest text encoder (63M parameters) is sufficient to represent the semantic concepts that the image encoder needs to match, so adding more layers provides diminishing returns. The 63M parameter size is kept constant across scaling experiments, while the image encoder grows.
| Aspect | Choice | Why It Matters |
|---|---|---|
| Image Encoder | ResNet-50 (modified) OR Vision Transformer | ResNet is proven; ViT is theoretically better at scaling |
| Attention Pooling | Learned weighted average of spatial features | Focuses on informative image regions |
| Text Encoder | 12-layer, 512-width Transformer | Balanced: large enough but not overly deep |
| Tokenization | BPE with 49,152 vocabulary | Subword units balance flexibility with efficiency |
| Scaling Strategy | Compound scaling for images; width-only for text | Maximizes efficiency; exploits asymmetry in modalities |
The key insight is that vision is the harder problem, so most scaling effort goes there, while the text encoder is sized sufficiently but not excessively.
We train a series of 5 ResNets and 3 Vision Transformers. For the ResNets we train a ResNet-50, a ResNet-101, and then 3...
This section describes the practical implementation details of how CLIP is actually trained. Think of it as the "recipe" section of a cookbook—previous sections told us what we're making and why we're making it, but this section tells us exactly how to make it: which models to train, what hyperparameters to use, and how long it takes.
Why does this matter? Because training is where theory meets reality. A great algorithm idea can fail in practice due to poor hyperparameter choices, numerical instability, or insufficient compute. This section shows how the authors overcame these challenges to successfully train CLIP at scale.
The authors don't train just one model—they train 8 different models total to explore how performance scales with model size:
ResNet-based models (5 total):
Vision Transformer models (3 total):
Why multiple models? This is a standard practice in ML research to show how performance scales with compute. They want to demonstrate that CLIP's benefits work across different model sizes, not just one particular configuration.
When the authors write "RN50x4", they mean:
So RN50x64 is about 64 times more computationally expensive than RN50 but should be proportionally more powerful. This follows the EfficientNet scaling philosophy mentioned in section 2.4, which allocates extra compute across width, depth, and resolution simultaneously rather than just one dimension.
An epoch is one complete pass through the entire dataset. With 400 million (image, text) pairs in the WIT dataset (from section 2.2), this is a substantial amount of training. At 32 epochs, each image is seen roughly 32 times during training.
The authors use Adam optimizer with a modification called "decoupled weight decay regularization."
What's the standard Adam optimizer?
Adam is an adaptive learning rate optimization algorithm. Without diving into full mathematics, the key idea is: Adam adjusts the learning rate individually for each parameter based on the history of gradients. Standard Adam can be written as:
where:
What does "decoupled weight decay" mean?
Weight decay is a regularization technique that penalizes large parameter values to prevent overfitting. In standard implementations, weight decay is folded into the gradient computation. Decoupled weight decay applies it separately to the parameters themselves after the Adam update, not through the gradient:
where is the weight decay coefficient.
Why does this matter? When Adam already adapts learning rates per parameter, mixing weight decay into gradients can interfere with this adaptation. Decoupling them keeps the adaptive learning rate benefits clean and has been shown to work better in modern large-scale training.
Important detail: Weight decay is applied "to all weights that are not gains or biases." This means:
Why exclude biases and gains? These parameters are typically less prone to overfitting and don't need as much regularization pressure.
The learning rate is decayed using a cosine schedule. This means the learning rate decreases over training following a cosine curve:
where:
Intuition: Early in training, we want a large learning rate to make rapid progress. As we approach convergence, we want a small learning rate to make fine adjustments without overshooting the optimal solution. A cosine curve provides a smooth interpolation between these two regimes.
From section 2.3, recall that CLIP uses a contrastive loss based on cosine similarity between image and text embeddings. The similarity scores are scaled by a temperature parameter before computing the loss.
The temperature parameter controls how "sharp" or "soft" the probability distributions are. Specifically, if the raw cosine similarity between embeddings is , it gets scaled to:
before feeding into a softmax or cross-entropy loss function.
Small (e.g., 0.01): Makes the similarities sharper. If one pair has similarity 0.9 and another has 0.8, scaling by small creates very large logits, which pushes one probability close to 1 and others close to 0. This can be good for hard negatives but risky.
Large (e.g., 1.0): Makes similarities smoother. The same similarities scaled by large produce logits closer to the original values, creating softer probability distributions.
The authors:
Why clip? If becomes too small, the logits become extremely large, causing numerical instability. Specifically, when you compute for very large , floating-point numbers overflow. By capping the maximum logit at 100, they prevent this instability while still allowing to adapt during training.
The authors use an exceptionally large minibatch size of 32,768 images.
Why so large?
Remember from section 2.3: CLIP trains to predict which of the possible (image, text) pairings across a batch actually occurred. Larger batches mean:
For a batch of size , there are:
This massive number of negative pairs per batch helps the model learn discriminative embeddings.
Mixed precision means using different numerical precisions for different parts of computation:
Why do this? It's a engineering trick that provides two practical benefits:
The downside is potential numerical instability, but with careful implementation and dynamic loss scaling, these issues can be managed. This is now standard practice in large-scale deep learning.
The training required massive computational resources:
For RN50x64 (largest ResNet):
For ViT-L/14 (largest Vision Transformer):
For context: A V100 GPU is a high-end professional graphics processor (~$8,000-10,000). This represents millions of dollars in compute.
Interestingly, the Vision Transformer (ViT-L/14) is trained on fewer GPUs in less wall-clock time than the largest ResNet (RN50x64). This might seem counterintuitive, but consider:
When the paper says "all results reported in this paper as 'CLIP' use this model," they're identifying which specific model represents the main contribution. This is important because researchers often compare different variants.
The notation "ViT-L/14@336px" means:
The authors apply an additional optimization step:
This technique is inspired by FixRes, which discovered that using higher resolution for the final epochs can boost performance significantly. Higher resolution provides more pixel-level detail for the model to learn from, which is especially beneficial in the later stages of training when the model is polishing its learned representations.
Why only 1 epoch at higher resolution? Computational cost grows with resolution. One additional epoch at 336×336 pixels is a reasonable trade-off: it's expensive enough to matter but cheap enough to add minimal total training time.
| Aspect | Choice | Reason |
|---|---|---|
| Models | 8 variants (ResNets + ViTs) | Show scaling across architectures |
| Epochs | 32 | Sufficient data exposure with 400M images |
| Optimizer | Adam + decoupled weight decay | Modern, stable, works well at scale |
| LR Schedule | Cosine annealing | Smooth decay from high to low learning rates |
| Temperature | Learned, initialized to 0.07, clipped at 100 | Adaptive but numerically stable |
| Batch size | 32,768 | Massive negatives pool per batch; efficient GPU utilization |
| Precision | Mixed (float16 + float32) | Speed and memory savings |
| Final model | ViT-L/14@336px | Best empirical performance |
The fundamental tension in training large models is:
The authors' choices navigate this balance:
These aren't arbitrary—they represent accumulated knowledge from years of large-scale deep learning research.
In computer vision, zero-shot learning usually refers to the study of generalizing to unseen object categories in image ...
This section is fundamentally about redefining what "zero-shot learning" means in computer vision, and more importantly, establishing it as the right way to measure whether a model can learn tasks (not just learn better representations).
Think of it this way: The CLIP paper has spent the previous sections describing how to train a model on 400 million image-text pairs. But how do you actually measure whether this model is useful? The authors argue that instead of just testing it on one task (like ImageNet), we should test it on many different tasks it's never seen before. If it can do well on unseen tasks, that proves the model learned something general about how to perform tasks — not just how to recognize ImageNet objects.
This distinction is crucial and somewhat subtle. Let me explain why.
Representation Learning: A model learns to extract useful features from images (e.g., "images of dogs should look similar to other images of dogs in the learned embedding space"). This helps with any downstream task using those features.
Task Learning: A model learns to understand and perform specific tasks without ever seeing examples of that task. For instance, a language model might learn what a "translation task" is just by seeing text, then be able to translate to a new language it encountered in training text.
The key insight: You can have great representations but still fail at new tasks. And conversely, task learning requires a deeper kind of understanding.
Most computer vision research focuses on representation learning—they measure it by checking "does this model extract good features?" But CLIP's language supervision enables something different: the model learns what tasks are, because text can describe tasks abstractly.
For example:
Historically in computer vision, "zero-shot learning" meant:
This is a narrow definition focused on generalization within image classification.
The paper explicitly states: "We instead use the term in a broader sense and study generalization to unseen datasets."
What does this mean mathematically? Let's define it precisely:
Traditional zero-shot setting (in mathematical notation):
CLIP's zero-shot setting:
Examples of the breadth:
Because it forces the model to not just memorize "what objects look like" but actually learn how to approach new problems. If CLIP can match accuracy on 30+ different datasets without any task-specific training, it's not just learning good features—it's learning how to be a general visual model.
The paper frames generalization to unseen datasets as "a proxy for performing unseen tasks."
Here's the formal intuition:
Imagine there exists a distribution over possible computer vision tasks:
A task consists of:
Traditional supervised learning:
where is the learned function with parameters .
Zero-shot transfer (CLIP's version):
The hypothesis: If is low across many different tasks, then has learned something fundamental about task structure—not just memorized one task.
The paper credits this as "first" work studying zero-shot transfer to standard image classification datasets. What made it significant?
Instead of training on datasets like ImageNet and testing on ImageNet, Visual N-Grams did:
CLIP extends this dramatically in scope (testing on 30+ datasets instead of 1-2).
This is a pivotal insight. The paper references Liu et al. (2018) discovering that:
This observation led to:
Why is this relevant to CLIP?
The authors are arguing: "Just like GPT-2 discovered that language models can learn to perform tasks, we want to show that vision models can too."
The mechanism is similar:
When the paper says it measures "task-learning capabilities," here's what it means mathematically:
Definition (Informal): A model has learned task structure if, when presented with a new task description (e.g., "classify these categories: cat, dog, bird"), it can perform that task without examples.
How CLIP does this:
For example, given the text prompt "a photo of a cat," CLIP:
This is task learning because the model never saw "cat" as a specific ImageNet class—it just understood from natural language what "a cat" means.
This section justifies the experimental methodology for the rest of the paper:
This is the evaluation framework that makes the claims in the abstract credible. If CLIP only worked well on ImageNet with zero-shot transfer, we'd just say "okay, it learned ImageNet concepts." But if it works on OCR, action recognition, geo-localization, and more—without any task-specific training—then we've really demonstrated task learning.
This section reframes the fundamental question from:
To:
And it argues that zero-shot transfer across many unseen datasets is the right way to measure task learning, inspired by how NLP models like GPT-2 demonstrated task learning through zero-shot transfer in language.
CLIP is pre-trained to predict if an image and a text snippet are paired together in its dataset. To perform zero-shot c...
This section explains the clever trick that makes CLIP work for zero-shot classification. During pre-training, CLIP learned to match images with text captions. Now the authors show how to repurpose this matching capability to classify images into categories the model has never explicitly seen during training.
The key insight: Instead of training a separate classifier for each new task, we can use the text encoder as a "classifier generator" that creates classifiers on-the-fly from natural language descriptions of the classes.
During pre-training (from Section 2.5), CLIP learned by seeing 400 million (image, text) pairs. The fundamental task was: "Given an image and a caption, are they paired together?"
For zero-shot transfer, the authors reuse this exact capability but redirect it toward classification.
Here's how it works step-by-step:
Step 1: Prepare class names as text
Step 2: Encode the image and all class texts
Step 3: Compute similarity scores
Step 4: Convert similarities to probabilities
Let me walk through the math rigorously.
Let:
Both embeddings are L2-normalized, meaning:
The notation is the L2 norm (Euclidean length). This normalization ensures all vectors lie on the unit hypersphere in .
The cosine similarity between two L2-normalized vectors is simply their dot product:
where denotes the transpose (treating vectors as row vectors for multiplication).
Why cosine similarity? For L2-normalized vectors, cosine similarity ranges from to and measures angular distance. It's invariant to the magnitude of the vectors (which is always 1 here), focusing purely on direction.
Scale the similarities by temperature:
where is the learnable temperature parameter. Think of as a "sharpness" control:
Then normalize into a probability distribution using softmax:
The softmax function is standard in machine learning for converting raw scores into valid probabilities (all non-negative, sum to 1).
The predicted class is:
We choose the class with the highest probability.
The section makes a subtle but important observation. Let me unpack it:
A typical linear classifier computes:
where:
CLIP's zero-shot classifier can be rewritten as:
This is equivalent to a linear classifier where:
Traditionally, you train the classifier weights on a dataset. CLIP takes a different approach:
The text encoder acts as a hypernetwork — a neural network that generates the weights of another neural network. Specifically:
This is powerful because you don't need labeled training data to create classifiers for new categories — you just give natural language descriptions, and the text encoder generates appropriate classifier weights.
The section makes another key observation:
During pre-training, CLIP sees 400 million (image, text) pairs. The authors note this is equivalent to:
Every step of CLIP pre-training can be viewed as optimizing the performance of a randomly created proxy to a computer vision dataset which contains 1 example per class and has 32,768 total classes.
What does this mean?
So: Pre-training inherently teaches the model to perform zero-shot classification on-the-fly, with 32,768 "classes" (text descriptions) per batch.
The last point is practical but important:
For zero-shot evaluation, we cache the zero-shot classifier once it has been computed by the text encoder and reuse it for all subsequent predictions.
Once you have the class names for a dataset, you can:
This is much faster than re-encoding the class names for every single image, especially when you have thousands or millions of test images.
| Aspect | Why It Matters |
|---|---|
| Reuses pre-training | CLIP was trained on image-text matching; we repurpose this for classification |
| No task-specific training | No labeled data needed for new tasks—just class names |
| Text as weights | Natural language descriptions become classifier weights via the hypernetwork (text encoder) |
| Implicit training | Pre-training on 32K-way classification teaches zero-shot skills automatically |
| Computational efficiency | Caching text embeddings avoids redundant computation |
This elegant mechanism is why CLIP can generalize to new tasks zero-shot: the model has already learned to map diverse visual concepts to diverse text descriptions, and classification is just a rephrasing of the matching task it was trained on.
In Table 1 we compare Visual N-Grams to CLIP. The best CLIP model improves accuracy on ImageNet from a proof of concept ...
This section addresses a critical practical problem: CLIP's zero-shot performance depends heavily on how you phrase the class labels to the text encoder. The authors show that:
This matters because zero-shot transfer is only useful if it works reliably—and these findings show that the quality of the natural language description of classes is surprisingly important.
| Dataset | Visual N-Grams | CLIP |
|---|---|---|
| ImageNet | 11.5% | 76.2% |
| aYahoo | Low (baseline) | 95% error reduction |
| SUN | Low (baseline) | 2x accuracy improvement |
Key insight on ImageNet: CLIP matches ResNet-50's accuracy (trained on 1.28 million labeled examples) using zero labeled examples. The fact that CLIP achieves 95% top-5 accuracy (matching Inception-V4) is notable because it shows the model learns robust hierarchical concepts—if the top prediction is wrong, the correct answer is usually in the top 5 choices.
The core issue is polysemy (words with multiple meanings). Consider the word "bank":
Recall from Section 3.1.2: During zero-shot classification, CLIP computes embeddings. The text encoder is treating the class name as a hypernetwork that generates classifier weights. A single-word label provides insufficient information to disambiguate meaning.
Additionally, the pre-training data distribution doesn't match test conditions:
The authors found that wrapping class labels in a template dramatically helps:
Why this works:
Quantitative impact on ImageNet: +1.3 percentage points (76.2% → 77.5%)
This seemingly small gain reveals a fundamental principle: representation quality depends on how you linguistically specify concepts.
Rather than using a single prompt, the authors ensemble predictions from 80 different context prompts. Here's how it works mathematically:
Step 1: Compute text embeddings for each prompt
For prompt , the text encoder generates embedding , where is the embedding dimension (e.g., 512).
Step 2: Average embeddings in embedding space (not probability space)
Step 3: Compute final cosine similarity
Given an image embedding , the final similarity score is:
(Recall from Section 3.1.2 that embeddings are L2-normalized, so this simplifies to a dot product.)
This is a crucial design decision. They could have averaged in probability space:
But this would require computing the softmax 80 times per prediction—expensive at inference time.
Instead, by averaging embeddings before computing similarities, they get:
This is amortizable: compute the average embedding once and cache it for all images. The ensemble then costs the same as a single classifier once the embeddings are cached.
Performance gain on ImageNet: +3.5 percentage points (77.5% → 81.0%)
Prompt engineering (template) + ensembling = +4.8 percentage points on average across 36 datasets
This is remarkable: without any additional data or model retraining, they gain ~5 points—equivalent to training with 4× more compute but "free" at inference time once embeddings are cached.
[Figure 4 shown in the original text] visualizes this improvement pattern across 36 datasets:
The caption notes: "This improvement is similar to the gain from using 4 times more compute with the baseline zero-shot method but is 'free' when amortized over many predictions."
This is because the computational cost of caching one averaged embedding vector (dimension , typically ~512 floats) is negligible compared to storing multiple separate embeddings.
Recall that embeddings are directions in high-dimensional space (they're L2-normalized, lying on a unit sphere). When you average multiple embeddings:
You're finding the centroid direction that's "closest" to all the individual directions. Different prompts capture different aspects of the semantic meaning (e.g., "A photo of X", "Artwork of X", "X", "An X"), and their average in embedding space produces a more robust representation that's resistant to prompt-specific biases.
This is geometrically similar to finding the mean direction of a set of vectors—robust to individual perturbations.
| Concept | Key Insight |
|---|---|
| Polysemy | Single words are ambiguous; context matters |
| Distribution mismatch | Pre-training ≠ test; bridge with templates |
| Prompt engineering | Wrapping labels in natural templates: +1-2% |
| Embedding-space ensembling | Average embeddings, not probabilities, for efficiency |
| Combined effect | +5% gain without retraining—practical and elegant |
The section demonstrates that zero-shot transfer quality isn't just about model architecture—linguistic specification of concepts matters enormously. This insight has proven influential in later work on prompt engineering for large language models and vision-language models.
Since task-agnostic zero-shot classifiers for computer vision have been understudied, CLIP provides a promising opportun...
This section is CRUCIAL because it answers a fundamental question: Does CLIP's zero-shot learning actually work?
Up to this point in the paper, we've seen impressive numbers (76.2% on ImageNet), but the researchers need to rigorously characterize when and why the model succeeds or fails. This section systematically compares CLIP against multiple baselines and analyzes patterns in its performance across diverse tasks. Think of this as the researchers saying: "Let's stress-test our model and understand its limitations."
The section conducts three major comparative analyses:
Let's break each down.
The researchers create a reference baseline: they take ResNet-50 (a standard computer vision model pre-trained on ImageNet) and fit a simple logistic regression classifier on top of its learned features.
Key concept: Both CLIP and this baseline use the same underlying feature representations, but in different ways:
For the fully supervised baseline, suppose we have training examples. We fit a logistic regression classifier:
Where:
For CLIP's zero-shot classifier (from section 3.1.2), recall it uses:
Where:
Result: Zero-shot CLIP wins on 16 out of 27 datasets—slightly better than random (13.5/27), but impressive because it uses zero labeled examples while the baseline uses potentially thousands.
The researchers observe extreme variability on fine-grained classification tasks:
Why? Fine-grained classification (distinguishing between car models or bird species) typically requires learning subtle visual distinctions. For example:
Critical weaknesses emerge on:
These failures make sense: the CLIP pre-training dataset isn't designed to represent these narrow domains.
Few-shot learning (e.g., "4-shot" means 4 labeled examples per class) is the natural middle ground between zero-shot and fully supervised. If zero-shot CLIP matches or exceeds few-shot performance, it suggests that language descriptions encode information as efficiently as direct visual examples.
For -shot learning, you have labeled examples per class. A linear classifier trained on these examples solves:
Where:
Striking result: Zero-shot CLIP matches the performance of a 4-shot linear classifier and nearly matches 16-shot performance.
Think about what this means:
The researchers provide crucial intuition:
"CLIP's zero-shot classifier is generated via natural language which allows for visual concepts to be directly specified, whereas normal supervised learning must infer concepts indirectly from training examples."
In other words:
The researchers compute a crucial metric: How many labeled examples would a supervised classifier need to match CLIP's zero-shot performance?
For each dataset, they fit fully supervised linear classifiers using labeled examples per class (where ). For dataset and number of shots , define:
Then, for each dataset , they find the threshold where:
This tells us: "This dataset needs about labeled examples per class to match zero-shot CLIP."
Across all datasets, there's a very strong positive correlation (, ) between zero-shot performance and fully supervised performance.
What does this mean?
The correlation coefficient means that if you rank datasets by supervised performance, zero-shot performance roughly follows the same ranking. Mathematically:
The researchers fit a linear regression model:
Where:
Interpretation of the slope: For every 1% improvement in supervised accuracy, zero-shot accuracy improves by 1.28%.
This is interesting because : zero-shot CLIP's performance scales super-linearly with task difficulty. As tasks get harder (lower supervised baseline), zero-shot CLIP actually maintains a larger performance gap, suggesting its language-based approach has some built-in advantages on challenging, specialized tasks.
However, Figure 8 reveals the broader picture: zero-shot performance is typically 10-25 percentage points lower than fully supervised performance, showing room for improvement.
This is a critical result that the paper emphasizes. The researchers train 5 different CLIP models with varying computational budgets, spanning a 44x range of compute.
Across 39 evaluations on 36 datasets, they observe that average zero-shot error follows a log-log linear relationship:
Equivalently:
Where:
Why is this important?
Power-law scaling suggests predictable, systematic improvements as you invest more computational resources. This is remarkable because:
This connects to broader findings in the literature (Kaplan et al., "Scaling Laws for Neural Language Models") suggesting that scaling laws are universal across different architectures and domains.
| Comparison | Finding | Implication |
|---|---|---|
| vs. Fully Supervised | Wins on 16/27 datasets | Zero-shot is competitive with strong baselines |
| vs. Few-Shot | Matches 4-shot, nearly matches 16-shot | Language descriptions are data-efficient |
| Correlation | with supervised baseline | Consistent, interpretable performance |
| Scaling Laws | Power-law with 44x compute range | Predictable, systematic improvements |
This section transforms CLIP from "impressively high numbers on ImageNet" to "a rigorously characterized system with predictable strengths and weaknesses."
The analysis shows:
This honest assessment—showing both strengths and limitations—is what makes the paper scientifically credible and practically useful.
While we have extensively analyzed the task-learning capabilities of CLIP through zero-shot transfer in the previous sec...
After spending significant time analyzing CLIP's zero-shot transfer capabilities (the ability to classify new tasks without any training examples), the authors now want to step back and ask a more fundamental question: How good are the visual representations that CLIP learns?
This is crucial because a model might get lucky on specific tasks, but the real test of pre-training quality is whether the learned features transfer well across many different scenarios. The authors use a standard benchmark in representation learning: fitting a linear classifier on top of frozen features. This approach reveals the underlying quality of the representations without the model being able to "cheat" by fine-tuning its core visual understanding.
The section opens with an important methodological choice that deserves careful explanation.
The Problem with Fine-tuning:
The Solution: Linear Classifiers on Frozen Features
Think of it like this: if you buy a tool that claims to be versatile, you don't test it by completely rebuilding it for each job. Instead, you test whether the tool's inherent design works well across different tasks without modification.
A linear probe is conceptually straightforward. Given:
We train a linear classifier by optimizing:
where comes from a softmax over linear logits:
Breaking this down:
Why this works as a representation quality metric: If CLIP learned good, general-purpose features, a simple linear classifier should achieve strong accuracy because the classes are already well-separated in the learned feature space.
The authors first evaluate on a conservative benchmark established by Kornblith et al. (2019) with 12 datasets. The key observation:
Small CLIP models (e.g., ResNet-50 backbone): Outperform standard ImageNet-trained ResNets, but underperform ResNets trained on ImageNet-21K (a much larger dataset with 14 million images and 21,000 classes)
Large CLIP models: Scale much better. The largest model tested (ResNet-50x64, where "x64" means 64 times more channels in the ResNet) slightly outperforms the previous state-of-the-art
CLIP Vision Transformers: Achieve approximately 3x better compute efficiency than CLIP ResNets. This means ViT-based CLIP models match ResNet performance while requiring only ~1/3 the computational cost.
Best model: A Vision Transformer (ViT-L/14, with L = Large, 14 = patch size of 14 pixels) fine-tuned at higher resolution (336 pixels instead of typical 224 pixels) for one additional epoch, outperforming prior state-of-the-art by 2.6% on average
When tested on a much wider variety of visual tasks (Figure 11), the improvements become more dramatic:
Where CLIP shows the biggest advantages:
| Task Category | Why CLIP Excels |
|---|---|
| OCR and Text (SST2, HatefulMemes) | Language supervision directly provides text understanding |
| Geo-localization & Scene Recognition (Country211, SUN397) | Descriptions help with geographic and environmental context |
| Activity Recognition in Videos (Kinetics700, UCF101) | Text descriptions of actions align well with temporal visual patterns |
The expansion from 12 to 27 datasets is methodologically important. Notice in Figure 10's right panel versus left panel:
The fact that benefits increase (2.6% → 5%) on the broader evaluation suggests that CLIP's representation learning strength comes from its training on diverse web text and images, making it particularly good at general visual understanding across varied domains.
Recall from the previous section on zero-shot transfer: CLIP matched 4-shot learning performance. Now the representation learning section shows that:
This creates a powerful narrative: CLIP doesn't just solve specific downstream tasks well—it learns fundamentally better visual representations for general computer vision.
Section 3.2 demonstrates that CLIP's success isn't limited to zero-shot transfer (covered in 3.1). Instead, the visual representations themselves are of exceptionally high quality, achieving state-of-the-art performance when simple linear classifiers are trained on top of them. This validates that the pre-training approach—matching images with natural language descriptions from 400 million web examples—is an excellent way to learn general-purpose visual features, not just task-specific classifiers.
In 2015, it was announced that a deep learning model exceeded human performance on the ImageNet test set. However, resea...
Before diving into the technical details, let's understand what this section is really about. The fundamental question is: Can CLIP maintain its strong performance when tested on data that looks different from what it was trained on?
This is crucial because real-world performance doesn't happen in a vacuum. When you deploy a model to production, it encounters images under different conditions, from different sources, with different visual characteristics than the training data. The fact that a model does well on ImageNet doesn't guarantee it will do well on photos from YouTube or sketches or adversarial examples.
This section shows that CLIP has a remarkable property: it's much more robust to these distribution shifts than traditional supervised models, even though traditional models achieve higher accuracy on the original ImageNet dataset. This is a surprising and important finding.
Deep neural networks are extraordinarily good at one thing: finding correlations in data. If you give a model labeled images, it will find patterns that predict the labels. But here's the catch:
Many of these patterns are spurious correlations—they only work because of quirks in how the training data was collected, not because they represent genuine, universal visual concepts.
Analogy: Imagine training a model to classify "outdoor scenes" using photos all taken in sunny weather. The model might learn to recognize "outdoor = mostly bright pixels" rather than learning actual outdoor features. When you test it on photos taken on cloudy days, performance collapses.
The section references a 2015 milestone: deep learning models exceeded human performance on ImageNet. This sounds like complete victory, but researchers discovered something humbling—these same models that beat humans on ImageNet fail badly on slightly different image distributions:
This pattern repeats across many domains, suggesting the problem is fundamental to how supervised learning works.
The section cites Taori et al. (2020), which proposes a crucial distinction. To understand this, let's define their framework mathematically.
Let's denote:
For a given model, there's an empirically observed relationship between these quantities across many models. We can write this as:
where is some function that describes the typical relationship. For example, if a model achieves 80% on ImageNet, we'd expect (based on historical data) roughly 75% on ImageNetV2. The function captures this typical relationship.
Definition: This measures the absolute improvement in out-of-distribution accuracy.
This is straightforward: "How well does my model do on the new distribution?" Higher is better.
Definition: This measures how much better you do compared to what you'd predict from the in-distribution to out-of-distribution relationship.
where is the predicted out-of-distribution accuracy based on the documented relationship.
Intuition: Imagine two models both achieving 80% on ImageNet. If the typical relationship predicts 75% on the new distribution:
Model B is more effectively robust—it doesn't just have good performance; it has better-than-typical performance given its ImageNet score.
Here's the key insight:
Zero-shot models are never trained on any specific distribution.
Think about how CLIP works (from earlier sections): it learned by matching image-text pairs from internet data. It developed a general understanding of visual concepts through language supervision, not by optimizing for any particular image classification task.
When a model is not trained on a distribution, it cannot possibly have learned spurious correlations that are specific to that distribution. It has no opportunity to overfit to distribution-specific quirks.
In contrast, an ImageNet model (trained via supervised learning on 1.28M labeled images) has had extensive opportunity to learn patterns that work on ImageNet specifically but fail elsewhere.
The section mentions evaluating on 7 different distribution shifts:
Each represents a different way the distribution can shift while keeping the underlying task (classification) the same.
From Figure 13's left panel, we can visualize the problem:
The "robustness gap" is the vertical distance between:
Quantitatively:
The ideal (impossible) case would be a horizontal line (Robustness Gap = 0).
The paper shows that zero-shot CLIP:
The section states: "reduce the size of the gap between ImageNet accuracy and accuracy under distribution shift by up to 75%."
This means if a standard ImageNet model has a robustness gap of, say, 15 percentage points, CLIP reduces it to roughly 3.75 percentage points.
The mechanism: Zero-shot CLIP's text descriptions provide explicit semantic specification of the task.
When you use CLIP zero-shot, you're saying: "Here's what I mean by 'dog': the text description 'a photo of a dog'." This explicit definition:
Here's something counterintuitive the section reveals:
When you take zero-shot CLIP and fine-tune it on ImageNet training data (supervised adaptation), accuracy improves by 9.2%, reaching 85.4% on ImageNet. This matches 2018 state-of-the-art.
But the robustness barely improves—it actually slightly decreases!
Let's denote:
We observe:
This means:
The gap actually widens when you fine-tune.
Fine-tuning the model on ImageNet training data allows it to learn distribution-specific correlations again. Even though you start with a robust representation, you end up overfitting to ImageNet's particular visual characteristics. The 9.2% accuracy gain comes because the model is learning these ImageNet-specific patterns—which helps in-distribution but hurts out-of-distribution.
This suggests a fundamental trade-off:
Figure 15 shows an interesting insight: few-shot adaptation (where you only use a few labeled examples per class) falls between zero-shot and fully supervised:
We can express this as a parametrized family:
where represents the "amount of zero-shot-ness" (0 = fully supervised, 1 = fully zero-shot).
As decreases (more fine-tuning), relative robustness increases but effective robustness decreases. This reflects a Pareto trade-off: you can't optimize both simultaneously.
| Aspect | Implication |
|---|---|
| Zero-shot robustness | CLIP reduces the robustness gap by up to 75% compared to ImageNet models |
| Mechanism | Not training on a distribution prevents learning spurious correlations |
| Effective vs. Relative | Zero-shot excels at effective robustness; supervised excels at relative robustness |
| Adaptation paradox | Improving in-distribution performance can harm out-of-distribution robustness |
| Few-shot middle ground | Offers a trade-off between robustness and specialization |
This section is critical because it shows CLIP isn't just a "different approach that trades one thing for another"—it actually achieves something genuinely better (robustness) while remaining competitive on traditional metrics (accuracy). This makes CLIP genuinely useful for deployment in real-world scenarios where data distribution is unpredictable.
How does CLIP compare to human performance and human learning? To get a better understanding of how well humans perform ...
This section addresses a fundamental question: How does CLIP compare to human intelligence? This matters because:
Benchmark for general intelligence: Humans are the gold standard for visual understanding. If CLIP matches human performance on zero-shot tasks, it suggests we've achieved something meaningful.
Understanding learning efficiency: Humans learn remarkably well from few examples. By comparing how humans and CLIP learn from 0, 1, and 2 examples per class, we can understand what separates human learning from current machine learning methods.
Identifying CLIP's limitations: The paper shows CLIP is competitive with supervised models on many tasks, but how does it compare when we strip away the computational advantage? This provides context for what we've actually achieved.
The authors conducted a controlled experiment on the Oxford IIT Pets dataset - a benchmark containing images of cat and dog breeds.
Experimental parameters:
This mirrors exactly how CLIP is evaluated in earlier sections (see Section 3.1.5), enabling a direct comparison.
The authors report:
Let me represent this more formally. If we denote human accuracy as where is the number of examples per class:
where is a small positive value (negligible improvement).
This gives us a marginal learning curve:
The key observation is:
In other words, the first example per class is vastly more informative than the second example.
This finding has deep implications for understanding human learning. The authors propose that humans have well-calibrated uncertainty estimates. Let me formalize this intuition:
Consider a human's internal belief about which breed an image belongs to, represented as a probability distribution .
Before seeing examples:
When given one example:
Mathematically, this follows from Bayes' theorem. If represents the variance in human beliefs:
\text{Expected improvement} \propto \text{Var}(P_{\text{human}}(\text{breed} | \text{image}))
The images with **highest variance** benefit most from new information. After one example, the human has substantially reduced uncertainty on those difficult cases, so the second example yields diminishing returns. --- ## Comparison with CLIP's Learning Here's where things get interesting. The paper contrasts two learning paradigms: ### Human Learning Pattern: - Rapid improvement on first example (22 percentage points) - Plateau after one example - Driven by uncertainty reduction on difficult cases ### CLIP's Few-Shot Learning (from earlier sections): - More gradual improvement as examples increase - Continues improving through 4-shot and beyond (see Figure 6) - Doesn't show the same steep-then-plateau pattern This suggests that **CLIP's few-shot learning mechanism is fundamentally different** from how humans learn. The paper notes: > "there is a large difference between how humans learn from a few examples and the few-shot methods in this paper" --- ## The Difficulty Alignment Finding **Figure 16** presents a crucial observation: The paper found a correlation between images that are hard for CLIP and images that are hard for humans.Let's denote:
This is significant because it suggests that CLIP's failure modes aren't arbitrary or due to quirks of the training process - rather, CLIP struggles with genuinely ambiguous categories, just like humans do. This indicates that CLIP has learned to identify the intrinsic difficulty of visual discrimination tasks.
Zero-shot CLIP (54%) is substantially below human zero-shot performance on this task, suggesting CLIP hasn't matched human prior knowledge about visual categories.
The human learning plateau is different from CLIP's curve, suggesting current few-shot learning methods don't capture the uncertainty-driven updating that humans perform.
Shared difficulty structure indicates CLIP is learning meaningful visual representations, not gaming the task with spurious correlations.
The paper is careful to note:
Here's a quick reference comparing the different conditions:
| Condition | Human Accuracy | CLIP Zero-Shot (typical) | Key Insight |
|---|---|---|---|
| 0 examples per class | 54% | Varies by task (10-80%) | Humans have good prior knowledge of animal breeds |
| 1 example per class | 76% | Not directly evaluated | Large jump; humans leverage uncertainty reduction |
| 2 examples per class | ~76% | Not directly evaluated | Minimal improvement; diminishing returns |
The main takeaway: Humans and CLIP learn differently. This section reveals we still have much to understand about few-shot learning and how to make AI systems learn as efficiently as humans do.
A concern with pre-training on a very large internet dataset is unintentional overlap with downstream evals. This is imp...
Imagine you're taking a test to prove you've learned something new, but unbeknownst to you, you accidentally studied the exact answer key. Your high score wouldn't prove you learned the material—it would just prove you memorized answers you'd already seen. This section addresses that exact problem for CLIP.
The core issue: CLIP was trained on 400 million image-caption pairs from the internet—an enormous, messy dataset. The researchers then tested CLIP on 35 different benchmark datasets to see how well it generalizes. The worry is: What if some of those benchmark images were accidentally included in the training data? If so, CLIP's reported performance wouldn't reflect true generalization ability; it would be inflated by familiarity.
This section rigorously quantifies:
Let me walk through their procedure:
For each evaluation dataset, they used a duplicate detector to identify which images appeared in both the training data and the test set. This partitions each evaluation dataset into:
Think of it like sorting a deck of cards: you're separating the cards you've definitely seen before from the cards that are completely new.
For each dataset, they computed zero-shot CLIP accuracy on all three subsets. The key metric is:
What this measures: The performance boost attributable to the overlapping images.
Computing the difference isn't enough—they needed to verify it's statistically significant. Here's where they applied a binomial significance test.
The null hypothesis: Overlap provides no benefit, so the true accuracy on overlapping examples equals the true accuracy on clean examples.
Under this hypothesis, the observed accuracy difference follows a binomial distribution. If you have:
Then the number of correct predictions on overlapping examples follows a binomial distribution: .
They also applied Bonferroni correction, which adjusts the significance threshold when running multiple tests. If you're testing 35 datasets simultaneously, the probability of finding at least one false positive by random chance is very high. Bonferroni correction combats this by dividing the significance threshold (typically ) by the number of tests:
where is the number of datasets (35 in this case). This gives:
Or equivalently, at 99.5% confidence (as mentioned in the figure caption), they require much stronger evidence before declaring a result significant.
Let me break down the quantitative findings:
| Metric | Value |
|---|---|
| Datasets with zero overlap | 9 out of 35 |
| Median overlap | 2.2% |
| Average overlap | 3.2% |
This tells us the problem, while present, isn't catastrophic. Most datasets have only a few percent overlap.
The researchers reported:
The counterintuitive finding: Even the dataset with the highest overlap (21.5% on Country211) showed minimal performance boost (0.2%). This suggests CLIP's learned representations are so robust that even seeing some training examples doesn't dramatically inflate performance.
Here's the crucial finding:
In plain language: Even when accounting for overlap, the performance gains are either tiny or within expected statistical noise.
The figure shows two panels worth examining:
Left panel: Shows confidence intervals for accuracy differences. Each horizontal line represents a dataset, and the width of the interval represents uncertainty in the estimated accuracy difference.
Right panel: A bar chart directly showing accuracy gains due to overlap for each dataset.
Why is overlap impact so small? Consider the structure of what CLIP learns:
CLIP is trained to predict which caption matches which image via contrastive learning. What it learns is a feature representation—essentially, a mathematical embedding space where semantically similar images cluster together.
If an image appears in the training data, CLIP learns associations between its visual content and its caption. But the zero-shot evaluation doesn't ask "do you remember this image?" Instead, it asks "given a new image, does this label's text description match?"
The key insight: Even if CLIP "sees" an image during training, it still must generalize the visual concept that image represents. Seeing one banana during training doesn't mean CLIP automatically recognizes all bananas better—it only helps if that banana's visual features transfer to other bananas.
Mathematically, if we denote:
Then:
where is small because the representation space is built for generalization, not memorization.
This section provides an important validity check on the paper's claims:
✓ CLIP's impressive zero-shot performance is not artificially inflated by data leakage
✓ The benchmark evaluations remain meaningful tests of generalization
✓ Any performance improvements from seeing test data are small and often not statistically significant
This strengthens the paper's central conclusion: CLIP genuinely learned transferable visual concepts from internet data, not just memorized benchmark examples.
There are still many limitations to CLIP. On datasets with training splits, the performance of zero-shot CLIP is on aver...
Before diving into the details, let's understand what this section accomplishes. The previous sections of the paper have presented CLIP as a remarkably successful model—it achieves competitive zero-shot performance on 30+ datasets, demonstrates robustness to distribution shift, and can transfer across diverse visual tasks without task-specific training.
This section is crucial because it honestly documents where CLIP fails. This matters for several reasons:
Now let's break down each limitation.
"On datasets with training splits, the performance of zero-shot CLIP is on average competitive with the simple supervised baseline of a linear classifier on top of ResNet-50 features. On most of these datasets, the performance of this baseline is now well below the overall state of the art."
Let's unpack what's being compared here:
Zero-shot CLIP: A model that makes predictions for classes it has never seen labeled examples of during training.
Supervised Baseline: A two-step approach:
The finding: CLIP's zero-shot accuracy ≈ this supervised baseline's accuracy.
While this sounds good, there's a critical caveat: the ResNet-50 baseline is now outdated. State-of-the-art models on most benchmarks substantially exceed this baseline.
Consider the performance gap. Let's denote:
Empirically, the paper finds:
The researchers then estimate that reaching requires approximately 1000× more compute.
This is a scaling estimate. In deep learning, performance often improves logarithmically with compute. If we denote performance as a function of computational budget :
where is the asymptotic performance ceiling and is a scaling constant. The 1000× estimate suggests the model is in a regime where each order-of-magnitude increase in compute yields meaningful but marginal performance gains.
The paper identifies several categories of tasks where CLIP struggles significantly:
Examples given:
Why is this hard?
Fine-grained classification requires distinguishing between visually similar categories that differ in subtle details. Mathematically, we're asking the model to partition a high-dimensional feature space (where visual features live) using decision boundaries that discriminate between classes that are very close together.
For fine-grained tasks, the margin between class decision boundaries must be very small. If we denote the classification score for class as:
where is the weight vector for class and is the extracted feature vector, then fine-grained classification requires:
CLIP learns from internet image-text pairs where fine-grained distinctions are rarely explicitly described in captions. The training signal doesn't emphasize these subtle visual differences.
Example: Counting objects in an image
This is fundamentally different from recognition. While CLIP learns visual patterns from image-text correlation, counting requires:
The paper notes CLIP fails catastrophically here. This reveals a core limitation: CLIP learns associations, not logical rules.
The most striking failure is MNIST (handwritten digit recognition). Here, the paper reports:
"CLIP's zero-shot MNIST performance is notably poor - an embarrassingly simple baseline of logistic regression on raw pixels outperforms zero-shot CLIP."
This is genuinely surprising because:
MNIST is arguably the simplest possible image classification task. Logistic regression is perhaps the most basic machine learning algorithm. The formula for logistic regression on raw pixels is:
where is the sigmoid function: , and is the raw pixel vector.
That logistic regression beats CLIP suggests CLIP's zero-shot approach has a fundamental mismatch with MNIST. Why?
"Although CLIP can flexibly generate zero-shot classifiers for a wide variety of tasks and datasets, CLIP is still limited to choosing from only those concepts in a given zero-shot classifier. This is a significant restriction compared to a truly flexible approach like image captioning which could generate novel outputs."
Let's be precise. CLIP's classification works as follows:
For a test image, CLIP computes a score for each class in a predefined set :
where:
The constraint: The model can only output scores for classes in . It cannot produce novel outputs or predictions outside this set.
In contrast, image captioning can generate arbitrary sequences. The output is not restricted to a predefined vocabulary of responses (though it is limited to learned vocabulary). Captioning generates:
where each comes from an action space that's exponentially large in the vocabulary size (for sequences of length ). This offers far greater flexibility.
CLIP is fundamentally a closed-set classifier. Domains requiring open-set prediction (novel categories) or fine-grained generation cannot be addressed without retraining or architectural changes.
"CLIP does not address the poor data efficiency of deep learning. Instead CLIP compensates by using a source of supervision that can be scaled to hundreds of millions of training examples."
Data efficiency measures how many labeled examples a model needs to achieve good performance. Mathematically, it's asking: what is the minimum dataset size such that:
where is accuracy trained on examples and is a target accuracy threshold.
A data-efficient learner requires small . A data-inefficient learner requires large .
Traditional deep learning requires massive labeled datasets due to the curse of high dimensionality. CLIP doesn't solve this problem; it circumvents it by:
Mathematically, CLIP trades the hardness of the supervised learning problem for the easier problem of finding naturally co-occurring image-text pairs.
Instead of requiring labeled examples where each example requires manual annotation, CLIP uses automatic pairs harvested from the internet, where but each pair requires no human effort.
"CLIP is trained on text paired with images on the internet. These image-text pairs are unfiltered and uncurated and result in CLIP models learning many social biases."
A bias in this context is a systematic association learned from training data that reflects societal inequities. For instance:
If the training data contains many images of "nurse" labeled with text describing women and "doctor" with text describing men, then:
CLIP will learn this association because its training objective—matching images to text—directly encodes these patterns.
The paper identifies the source: unfiltered internet data.
The 400 million image-text pairs are collected with minimal curation. They reflect:
These biases are embedded in the learned representations. When CLIP learns to match images of people to text descriptions, it simultaneously learns the gender, race, and other demographic associations present in the raw data.
The bias is not a bug in the algorithm; it's a feature of training on real-world data. Mathematically, the model is simply performing:
where is the (biased) internet data distribution and is the contrastive loss. If encodes biases, then the learned parameters will too.
Fixing this requires either:
Let's synthesize these limitations into a framework:
| Limitation | Current State | Target State | Required Effort |
|---|---|---|---|
| Overall accuracy | ~ResNet-50 baseline | SOTA | 1000× compute |
| Fine-grained classification | Poor | Competitive | Task-specific training |
| Abstract reasoning | Fails (e.g., counting) | Succeeds | Fundamental architecture change |
| Output flexibility | Closed-set | Open-set | Generative component |
| Data efficiency | Poor | Good | Algorithmic innovation |
| Social biases | Present | Absent | Data curation + debiasing |
This limitations section is brutally honest. It establishes that while CLIP is a breakthrough in training efficiency (learning from internet-scale supervision) and transfer capability (zero-shot generalization), it is not a breakthrough in:
Understanding these tradeoffs is essential for practitioners deciding whether CLIP fits their use case.
CLIP has a wide range of capabilities due to its ability to carry out arbitrary image classification tasks. CLIP also in...
The previous sections of this paper focused on what CLIP can do technically—how it learns from image-text pairs, how well it performs on various benchmarks, and what its limitations are. Section 7 shifts perspective to ask a more important question: what does it mean when we deploy this technology in the real world?
This is crucial because CLIP is fundamentally different from traditional computer vision systems in a way that amplifies ethical concerns. Let me explain why this matters:
Traditional supervised models are constrained—they can only predict the specific categories they were trained on. If you want to detect a new category, you need labeled data and retraining.
CLIP breaks this constraint. Because it learned from natural language supervision, it can classify any concept you can describe in words without retraining. This flexibility is technically impressive, but it creates new risks.
The paper makes a crucial observation: CLIP's zero-shot capability—the ability to classify arbitrary categories by just writing text prompts—is both a strength and a vulnerability. Here's the logical structure:
Problem with traditional ML: If a system has built-in biases, those biases are fixed to the predetermined categories. The harm is limited in scope.
Problem with CLIP: Any person can define their own classification task. If the underlying model has learned societal biases from internet text, those biases can now be weaponized in arbitrary ways by arbitrary users.
Precedent concern: The paper cites GPT-3, a large language model. GPT-3 can generate text on almost any topic without retraining. Researchers discovered it had harmful capabilities only after testing it extensively. CLIP has the same property—we don't know all the ways it could be misused until people start trying.
This is formally similar to the concept of emergent capabilities in large models: properties that appear only when models reach sufficient scale and are not explicitly programmed or labeled.
Think of it this way. If a model is trained on data with certain statistical regularities, those regularities don't disappear—they scale with the model:
While this isn't a precise equation in the paper, it captures the intuition: larger models trained on larger datasets can express learned biases more fluidly across more domains.
With a traditional supervised model, the creator decides what categories exist. This creates a checkpoint where bias can (theoretically) be caught.
With CLIP, end-users define categories. For example:
The model itself doesn't know it's being asked for something harmful—it's just matching images to text. This is the capability without consent problem.
The paper explicitly notes: "CLIP is trained on text paired with images on the internet. These image-text pairs are unfiltered and uncurated and result in CLIP models learning many social biases."
What does "learning biases" mean mathematically? When CLIP's training objective is:
where the text comes from captions on the internet, the model is learning to match images to language patterns that exist in internet data. If those patterns contain stereotypes (e.g., "nurse" appears more often with images of women than men in the training data), CLIP learns those associations.
This is technically a compression of the statistics in the training data:
The model doesn't know these associations are harmful—it's learning statistical regularities, but those regularities encode human biases.
The paper describes their bias evaluation strategy:
FairFace benchmark: A dataset specifically designed to test whether models exhibit demographic bias (e.g., predicting age, gender, race)
Exploratory bias probes: Custom tests to examine CLIP's behavior in scenarios suspected to contain bias
Downstream task analysis: Studying a specific high-stakes domain—surveillance—to understand real-world harms
The authors are honest about limitations:
"Our bias tests represent our initial efforts to probe aspects of how the model responds in different scenarios, and are by nature limited in scope."
This is a crucial admission. Consider the mathematical perspective:
Let be the theoretical space of all possible tasks where CLIP could be applied, and let be the subset they actually tested.
|\mathcal{T}_{\text{tested}}| \ll |\mathcal{T}| Because CLIP can describe *any concept in natural language*, the space of possible tasks is essentially infinite. You cannot comprehensively evaluate an infinitely-parametrizable system. This is why the authors compare to GPT-3: large generative models have "capability discovery" dynamics where harmful uses are found reactively, not proactively. --- ## The Surveillance Case Study: Concrete Example Why did they focus on surveillance specifically? Surveillance is high-stakes because: 1. **Scale**: Once deployed, impacts millions of people systematically 2. **Permanence**: Creates records and precedents 3. **Power asymmetry**: People don't consent to being analyzed 4. **Difficult correction**: Errors compound over time If CLIP inherited biases about who "looks suspicious," and those biases correlate with protected characteristics (race, gender, etc.), the system could systematically harm certain populations. Mathematically, this is about the **harm amplification factor**: \text{Total Harm} = \text{Model Bias} \times \text{Number of Decisions} \times \text{Consequence per Decision} In surveillance: - "Model Bias" = whatever stereotypes CLIP learned - "Number of Decisions" = every frame of every camera = billions daily - "Consequence per Decision" = potential wrongful arrest, harassment, discrimination --- ## Why This Section Matters for the Broader ML Community This section is important for two reasons: ### 1. **Capability ≠ Alignment** CLIP is *capable* of zero-shot transfer to arbitrary tasks. But capability doesn't mean the model will behave fairly or safely. This is a general principle: \text{Technical Success} \not\Rightarrow \text{Ethical Deployment}The shift from "predetermined categories" to "any category via natural language" is qualitative, not just quantitative. It changes:
| Dimension | Traditional Model | CLIP |
|---|---|---|
| Categories defined by | Model creators | End-users (via text) |
| Bias scope | Limited to predetermined categories | Any category can inherit biases |
| Testing strategy | Comprehensive testing of known tasks | Reactive testing (discovering uses after deployment) |
| Governance | Easier—centralized decision point | Harder—distributed across users |
CLIP's flexibility is a feature and a bug: The same capability that makes it powerful makes it dangerous, because any user can instantiate arbitrary tasks.
Bias is statistical, not intentional: CLIP learned associations from internet data; these associations encode societal biases but aren't deliberately programmed.
Evaluation is incomplete by necessity: You cannot test an infinitely-parametrizable system comprehensively. The authors acknowledge this honestly.
Real-world deployment requires governance: Unlike the technical sections of the paper, this section argues we need social and institutional structures, not just better algorithms, to deploy CLIP safely.
The broader message: as ML systems become more capable and more flexible, the burden shifts from technical evaluation to governance and responsible deployment.
Algorithmic decisions, training data, and choices about how classes are defined and taxonomized can all contribute to an...
Before diving into the technical details, let's understand what this section is addressing. CLIP is a powerful tool that can classify images based on any text description a developer provides—no retraining needed. This flexibility is powerful, but it's also dangerous: if the underlying model has absorbed social biases from its training data, then developers can easily create classification systems that perpetuate discrimination, often without realizing it.
The section investigates whether CLIP exhibits measurable social biases in how it classifies images, particularly images of people. The authors probe for two types of bias:
This is crucial because algorithmic bias isn't just a fairness issue—it can cause real harm when systems like CLIP are deployed in real-world applications.
The authors use the FairFace dataset as their initial bias probe. FairFace is a carefully curated benchmark designed to evaluate algorithmic bias. The dataset includes face images annotated with demographic attributes like race, gender, and age.
They evaluate two versions of CLIP:
The logistic regression model can be mathematically expressed as:
Where:
Key finding: LR-CLIP actually outperformed both ResNext-101 (trained on Instagram) and FairFace's own specialized model on most tasks. This tells us that CLIP's features are high-quality and encode demographic information useful for classification—which is necessary for fair performance but also means the model has absorbed demographic information that could be misused.
The authors then performed a more concerning experiment: they probed CLIP using classification terms designed to reveal biases. They created a set of classification labels that could cause representational harm—the harm that comes from associating people with non-human categories.
The specific labels tested were:
'animal', 'chimpanzee', 'gorilla', 'orangutan'For a demographic group and this set of non-human labels , they computed:
This is simply the proportion of images from group that CLIP classified into one of the non-human categories.
Overall baseline: 4.9% of all images were misclassified as non-human.
Broken down by race (approximate rates):
This is a stark difference. The misclassification rate for the Black demographic group is roughly 2.8× higher than the overall average. This is classic representational harm—a model systematically treating one demographic group as less than human.
Mathematically, we can quantify this disparity as a disparity ratio:
A ratio of 1.0 would indicate no disparity; ratios significantly above 1.0 reveal which groups experience worse outcomes.
The authors then probed for allocation harm—negative treatment in the allocation of resources or opportunities. They tested whether certain demographic groups are disproportionately associated with crime-related language.
The crime-related labels tested were:
'thief', 'suspicious person', 'criminal'For each gender group , they computed the proportion of images misclassified into these crime categories:
For crime-related terms:
The disparity ratio here is:
Males are classified as criminals at roughly 1.68× the rate of females. The authors also note "significant disparities in classifications across races for crime related terms," though specific numbers aren't provided for this dimension.
The authors made an important discovery: when they added an explicit 'child' category to the classification options, the misclassification rates into both crime categories and non-human animal categories dropped dramatically.
This reveals something profound about how CLIP works. When forced to choose between "person," "criminal," "gorilla," and now "child," the model does a better job recognizing when an image is of a young person rather than forcing it into an adult crime category or animal category.
Mathematically, this relates to how the softmax function allocates probability mass:
Where is the model's score (logit) for class , and is the total number of classes. When you add a class that CLIP recognizes better (like "child"), it can shift probability away from harmful categories.
Key insight: This shows that class design is not neutral. The labels developers choose to include (or exclude) fundamentally shape what biases the model exhibits. A seemingly technical choice about which categories to include is actually a deeply consequential design decision.
CLIP's core innovation is that anyone can define arbitrary classes without retraining. But this creates a problem:
The misclassifications documented here likely stem from:
This is related to the problem of proxy variables. Even if you never directly tell CLIP "Black = criminal," if the training data has statistical correlations between certain visual features and crime-related language, CLIP will learn those spurious correlations. Mathematically:
Where extracts CLIP's features. Even if isn't explicitly trained to encode race, it might incidentally do so because race is correlated with many other visual features.
The authors are appropriately humble about their methodology. They note their bias tests are:
| Bias Type | Finding | Disparity Ratio |
|---|---|---|
| Representational (animals) | Black people misclassified as animals at 14% vs 4.9% overall | ~2.86× |
| Allocation (crime, gender) | Males misclassified as criminals at 16.5% vs females 9.8% | ~1.68× |
| Intervention (adding "child" class) | Dramatically reduced both harmful categories | Variable |
This section reveals a fundamental challenge in deploying powerful ML systems: we cannot simply optimize for accuracy. A model that achieves high average accuracy may still harbor serious biases that systematically harm specific demographic groups. CLIP's flexibility—allowing any developer to define classes—actually amplifies this risk because developers might unknowingly create harmful classification systems.
The most important finding may be the class design intervention: it shows that biases aren't fixed properties of CLIP but rather emerge from the interaction between the model and how developers choose to deploy it. This puts responsibility on developers to be thoughtful about what categories they create.
We next sought to characterize model performance in relation to a downstream task for which there is significant societa...
This section evaluates CLIP's performance on a particularly sensitive real-world application: surveillance and identity detection. The authors are asking: "How well does CLIP perform on tasks that matter for surveillance systems?" This matters because:
The key finding: CLIP performs reasonably well on some surveillance tasks, but not well enough to be practically superior to existing supervised models, and very poorly on fine-grained detection tasks.
What they tested: Can CLIP classify images from surveillance cameras (CCTV)?
The results show high variability:
"The model had a top-1 accuracy of 91.8% on the CCTV images for the initial evaluation. The accuracy dropped significantly to 51.1% for the second evaluation..."
What do these numbers mean?
Why the dramatic drop? The authors provide a clue: "the model incorrectly choosing the 'close' answer 40.7% of the time."
This suggests the evaluation setup changed. One possibility:
Critical observation: A 40.7% misclassification rate into a specific wrong category indicates the model isn't just randomly failing—it's systematically confused, making the same mistake repeatedly.
"For fine-grained detection, the zero-shot model performed poorly, with results near random."
What's "fine-grained detection"? This means distinguishing between visually similar objects within the same category—for example:
This is much harder than basic classification because the visual differences are subtle. The mathematical challenge: fine-grained detection requires the model's learned representation to be sensitive to small, nuanced visual differences.
If we think of CLIP's learned visual space as high-dimensional vectors, fine-grained tasks require that similar-looking objects (e.g., two faces) have vectors that are far apart, while coarse categories have vectors that are close. CLIP wasn't trained on such fine-grained distinctions, so it fails here.
The authors tested celebrity identification with varying numbers of possible identities:
Scenario 1: 100 possible celebrities
Scenario 2: 1,000 possible celebrities
Why does accuracy drop as we add more classes?
This relates to a fundamental concept: class imbalance and disambiguation. When CLIP must choose from 100 classes, it can afford to be relatively uncertain—there's a higher prior probability (baseline) that any given prediction is correct just by chance.
For 100 classes: random guessing = 1%
For 1,000 classes: random guessing = 0.1%
CLIP's 59.2% (vs. 1% baseline) is much better than random but still not practical.
More importantly, as the classification set grows, CLIP must make finer distinctions. With 1,000 celebrity names, many will be visually similar people, and CLIP struggles because:
The section ends with a crucial insight:
"CLIP offers significant benefit for tasks that have relatively little data given its zero-shot capabilities. However, large datasets and high performing supervised models exist for many in-demand surveillance tasks such as facial recognition. As a result, CLIP's comparative appeal for such uses is low."
What does this mean?
For zero-shot learning performance, CLIP shines when you have:
But for surveillance specifically:
Mathematical perspective: If you have labeled examples for your surveillance task, supervised learning typically achieves:
\text{Accuracy}_{\text{supervised}} = f(m, \text{architecture}) \approx f(m >> 0, \text{ResNet}) $While zero-shot CLIP achieves:
\text{Accuracy}_{\text{CLIP zero-shot}} = f(\text{internet knowledge}, \text{natural language}) $
For large , the supervised approach wins because it can specialize to your specific task and distribution. CLIP's advantage only kicks in when is very small or zero.
This honest assessment of limitations is important because:
The surveillance section essentially says: "Yes, CLIP works reasonably well on this, but here's why practitioners shouldn't worry about replacing their specialized systems with CLIP."
This preliminary analysis is intended to illustrate some of the challenges that general purpose computer vision models p...
This brief but important section serves as a bridge between the critical analysis the authors have just completed (examining CLIP's biases and surveillance capabilities) and the broader research community. The authors are essentially saying: "We've identified some concerning patterns in our model. Rather than leaving it at that, here's what we think needs to happen next."
The core message is: General-purpose vision models like CLIP are powerful but potentially problematic, and we need systematic community effort to understand them fully before deploying them widely.
"This preliminary analysis is intended to illustrate some of the challenges that general purpose computer vision models pose and to give a glimpse into their biases and impacts."
What this means mathematically/conceptually:
The authors are being explicit about the scope of their investigation. They've tested CLIP on:
However, this is not exhaustive. The model has been trained on 400 million image-text pairs from the internet (as stated in the abstract), meaning there are countless possible downstream tasks and scenarios they haven't tested.
The mathematical point: They've sampled a small subset of possible tasks from a much larger space of potential applications. If we denote:
Then (the tested set is much smaller than the full set).
"We hope that this work motivates future research on the characterization of the capabilities, shortcomings, and biases of such models."
What does "characterization" mean here?
In scientific and engineering contexts, characterization means systematically determining:
For CLIP specifically, this means the community should study:
| Category | Examples from the Paper |
|---|---|
| Capabilities | Matches ResNet-50 on ImageNet zero-shot; strong performance on face classification (FairFace dataset); surveillance with 91.8% accuracy on one CCTV evaluation |
| Shortcomings | Performance drops to 51.1% on second CCTV evaluation; only 43.3% accuracy on 1k-class celebrity identification; poor fine-grained detection |
| Biases | 14% misclassification of Black faces into animal categories (vs. <8% for other races); 16.5% of male faces into crime categories (vs. 9.8% for females) |
"We believe one good step forward is community exploration to further characterize the capabilities of models like CLIP..."
What are they asking for?
The authors suggest a distributed, collaborative approach to understanding CLIP. Rather than one team testing everything, they're inviting researchers worldwide to:
The mathematical framework they're implying:
For each domain and task , we can define:
And we can decompose problematic cases as:
Where the bias component might look like:
This measures whether error rates differ across demographic groups—exactly what they observed in Section 7.1.
The authors outline four key activities:
Find applications where CLIP genuinely helps. The paper shows this includes:
Counter-example: Surveillance. While CLIP achieves 91.8% on one evaluation, large-scale supervised systems already exist and outperform it. So CLIP is not beneficial here.
Identify domains where errors have serious consequences. Examples:
For these tasks, the "cost of being wrong" is high, so even moderate bias becomes problematic.
This requires systematic testing. They found biases through:
The statistical idea: For a given task and demographic group , compute:
where = observed frequency of label for group , and = expected frequency. If this exceeds a threshold (they used 0.5%), the difference is "significant."
This is about understanding why failures occur. For instance:
In Figure 18: Male Members of Congress were assigned gendered labels more often than female members. Why? Perhaps because the training data contains biased associations between certain professions/characteristics and gender.
In the animal classification error: Why do Black faces get misclassified into animal categories at 14% vs. <8% for other races? Possible reasons:
Throughout this paper, the authors have made CLIP extremely general-purpose. As they note in Section 7:
"CLIP makes it possible to easily create your own classes for categorization without a need for re-training."
This is powerful but dangerous. Consider:
Each of these could amplify societal biases, especially if the specific bias hasn't been studied.
The authors are essentially proposing a research roadmap:
| Phase | Activity | Goal |
|---|---|---|
| Current | Initial bias probes and capability assessments | Identify known issues |
| Near-term | Invite community testing across domains | Discover unknown issues |
| Medium-term | Systematic characterization of each domain | Build a map of where CLIP is safe/unsafe |
| Long-term | Improve model design to reduce bias | Build better models with CLIP's benefits and fewer harms |
The mathematical intuition: They're asking the community to help fill in the vast space —to systematically explore what they couldn't possibly test themselves.
This is a honest acknowledgment that a single paper, no matter how thorough, cannot fully characterize a general-purpose model. That's an inherently community-level task.
Any model that leverages written, spoken, signed or any other form of human language as part of its training signal is a...
We have investigated whether it is possible to transfer the success of task-agnostic web-scale pre-training in NLP to an...