Skip to content
Go to insights Blog

CLIP - The AI that connects images and language

CLIP - The AI that connects images and language
Written by
Data Science Lab
Published on
25 February 2025

Image recognition has a problem

Traditional computer vision models work well for specific tasks, but struggle to generalise. A model that recognises cats and dogs cannot automatically recognise people. Even ImageNET, a dataset with 1,000 categories, has its limitations. It does not recognise people, for example, let alone distinguish between a child and an adult. But now there is CLIP.

CLIP: Generalisation in image recognition

Contrastive Language–Image Pre-training (CLIP), developed by OpenAI, addresses this problem. CLIP is designed as a general-purpose model that understands images without specific training for each category. It recognises not only cats and dogs, but also abstract concepts such as "an astronaut riding a horse in space." This makes it a powerful computer vision model: instead of training multiple models for different tasks, CLIP offers a one-size-fits-all solution. A similar architecture is also used in GPT-4 for image processing.


CLIP an austronaut riding a horse in space

Image 1. The description that CLIP finds matches the image best

How CLIP works

OpenAI describes CLIP as

[“…a neural network trained on a variety of (image, text) pairs. It can be instructed in natural language to predict the most relevant text snippet, given an image, without directly optimizing for the task, similarly to the zero-shot capabilities of GPT-2 and 3”]

Those zero-shot capabilities mean it can recognise concepts it has never seen before. For example, the model has never seen an image of an "astronaut riding a horse in space", but it does know what an astronaut, a horse and space are. By combining these concepts, it can assign the correct classification. To understand how CLIP achieves this generalisation and matches images with text, we need to look more closely at the technology.

Under the hood

CLIP links text and images through embeddings—numerical representations of data in a high-dimensional space. In this space, matching images and text are close together. This makes it possible to perform tasks such as image captioning, object detection and classification. To classify an image, CLIP compares its image embedding with text descriptions such as “A photo of a dog” and “A photo of a cat”. The model chooses the description that best matches the image.

CLIP was trained on 400 million image-text pairs collected from the internet. This gives it an advantage over traditional datasets with simple labels such as “cat”. CLIP’s dataset contains richer descriptions, such as “a sleeping cat on a chair” or “a cat playing with a ball”. As a result, the model learns not only what a cat is, but also what a chair or a ball is.

Converting these image-text pairs into useful embeddings happens in a step called contrastive pre-training . In this step, the model is trained to match genuine image-text pairs within a batch of N randomly selected pairs from the dataset. A text encoder and an image encoder each transform their input into an embedding. The aim is for the image and text embeddings of pair 1 to match, and likewise for every pair in the batch. The training process tries to maximise the cosine similarity between the embeddings of genuine pairs and minimise the similarity between the embeddings of incorrect pairs.

Diagram of contrastive pre-training with a text encoder and an image encoder.

Figure 2: Contrastive pre-training.

By repeating this process multiple times across the entire dataset, CLIP learns representations of many concepts. Once trained, the learned embeddings can be used for different purposes. CLIP can be used for the image classification mentioned earlier, but there are other interesting use cases that make use of its capabilities.

Use cases: What is CLIP used for?

CLIP was originally developed for image classification, but is now used for many more applications:

  • Multimodality in LLMs: The ability to combine text and images also makes CLIP models useful in LLMs. They can be used for visual question answering, where questions about images are answered, or for multimodal tasks, such as describing an image provided by the user.
  • Semantic search system: Image search engines often rely on manually added tags and metadata, which takes a lot of work and is prone to errors. CLIP matches text queries directly to relevant images without the need for manual tagging.
  • Content moderation: CLIP’s versatility also makes it a useful tool for keeping content in online communities safe. CLIP can detect and automatically remove unwanted content such as nudity, weapons or drugs. Its flexibility means you can easily add categories without retraining the model.
  • Accessibility: Blind and visually impaired people rely on screen readers to tell them what is on a web page. For screen readers, images need alt text describing what they show. Unfortunately, alt text is not always present. In these cases, CLIP could be used to generate an image description automatically.
  • Recommendation system: CLIP can match images without text. For example, a customer of an online shop could upload a photo of a sofa, and the system would recommend similar products. 
  • Generative models: Generative models such as [DALL-E] and [Stable Diffusion] use CLIP behind the scenes to turn a user’s prompt into numerical input a generative model can use.

The future of CLIP

CLIP is a versatile solution to the challenge of generalisation in computer vision. Its ability to understand images and text opens the door to a range of applications, from content moderation to improved accessibility, efficient recommendation systems and many more.

Its integration into modern LLMs such as GPT-4 also shows that CLIP goes beyond computer vision alone. The technology makes AI even more flexible and capable. And given how quickly AI is developing, this is just the beginning.

Blog

You may also find this interesting,

Sign up for our newsletter.

Want to be the first to hear about a new blog post?

Enter a valid email address.