
Description
Searching images by text or classifying without training usually means a new vision model per task. CLIP is OpenAI's Contrastive Language-Image Pretraining model judging how well images match text.
It enables zero-shot classification and image-text retrieval, a building block of Stable Diffusion and many multimodal models.
Zero-shot:No training.
Retrieval:Text to image.
Embeddings:General purpose.
Ecosystem:Widely used.
It enables zero-shot classification and image-text retrieval, a building block of Stable Diffusion and many multimodal models.
Features
Zero-shot:No training.
Retrieval:Text to image.
Embeddings:General purpose.
Ecosystem:Widely used.
