CLIP
CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
About this project
CLIP [[Blog]](https://openai.com/blog/clip/) [[Paper]](https://arxiv.org/abs/2103.00020) [[Model Card]](model-card.md) [[Colab]](https://colab.research.google.com/github/openai/clip/blob/master/notebooks/InteractingwithCLIP.ipynb) CLIP (Contrastive Language-Image Pre-Training) is a neural network trained on a variety of (image, text) pairs. It can be instructed in natural language to predict the most relevant text snippet, given an image, without directly optimizing for the task, similarly to the zero-shot capabilities of GPT-2 and 3. We found CLIP matches the performance of the original ResNet50 on ImageNet “zero-shot” without using any of the original 1.28M labeled examples, overcoming…
Technologies
Project health
GitHub
Reviews
Built by
Maintain openai/CLIP? Claiming verifies admin access through your GitHub account and gives you control of this listing.
