Three lessons from our AI projects
Over two years, we supported dozens of AI projects, from initial exploration to a model that runs every night. Some went exactly as planned; others…
Traditional computer vision models work well for specific tasks, but struggle to generalise. A model that recognises cats and dogs cannot automatically recognise people. Even ImageNET, a dataset with 1,000 categories, has its limitations. It does not recognise people, for example, let alone distinguish between a child and an adult. But now there is CLIP.
Contrastive Language–Image Pre-training (CLIP), developed by OpenAI, addresses this problem. CLIP is designed as a general-purpose model that understands images without specific training for each category. It recognises not only cats and dogs, but also abstract concepts such as "an astronaut riding a horse in space." This makes it a powerful computer vision model: instead of training multiple models for different tasks, CLIP offers a one-size-fits-all solution. A similar architecture is also used in GPT-4 for image processing.
Image 1. The description that CLIP finds matches the image best
OpenAI describes CLIP as
Those zero-shot capabilities mean it can recognise concepts it has never seen before. For example, the model has never seen an image of an "astronaut riding a horse in space", but it does know what an astronaut, a horse and space are. By combining these concepts, it can assign the correct classification. To understand how CLIP achieves this generalisation and matches images with text, we need to look more closely at the technology.
CLIP links text and images through embeddings—numerical representations of data in a high-dimensional space. In this space, matching images and text are close together. This makes it possible to perform tasks such as image captioning, object detection and classification. To classify an image, CLIP compares its image embedding with text descriptions such as “A photo of a dog” and “A photo of a cat”. The model chooses the description that best matches the image.
CLIP was trained on 400 million image-text pairs collected from the internet. This gives it an advantage over traditional datasets with simple labels such as “cat”. CLIP’s dataset contains richer descriptions, such as “a sleeping cat on a chair” or “a cat playing with a ball”. As a result, the model learns not only what a cat is, but also what a chair or a ball is.
Converting these image-text pairs into useful embeddings happens in a step called contrastive pre-training . In this step, the model is trained to match genuine image-text pairs within a batch of N randomly selected pairs from the dataset. A text encoder and an image encoder each transform their input into an embedding. The aim is for the image and text embeddings of pair 1 to match, and likewise for every pair in the batch. The training process tries to maximise the cosine similarity between the embeddings of genuine pairs and minimise the similarity between the embeddings of incorrect pairs.
Figure 2: Contrastive pre-training.
By repeating this process multiple times across the entire dataset, CLIP learns representations of many concepts. Once trained, the learned embeddings can be used for different purposes. CLIP can be used for the image classification mentioned earlier, but there are other interesting use cases that make use of its capabilities.
CLIP was originally developed for image classification, but is now used for many more applications:
CLIP is a versatile solution to the challenge of generalisation in computer vision. Its ability to understand images and text opens the door to a range of applications, from content moderation to improved accessibility, efficient recommendation systems and many more.
Its integration into modern LLMs such as GPT-4 also shows that CLIP goes beyond computer vision alone. The technology makes AI even more flexible and capable. And given how quickly AI is developing, this is just the beginning.
Over two years, we supported dozens of AI projects, from initial exploration to a model that runs every night. Some went exactly as planned; others…
Many organisations have now run an AI pilot. The model works, the demo gets applause, and then nothing else happens. In our experience, most projects…
Artificial Intelligence is developing rapidly. New models appear almost every week, and more and more organisations are experimenting with AI. At the…
Want to be the first to hear about a new blog post?
Thanks for signing up!