CLIP

CLIP

Contrastive Language-Image Pretraining

Description

Searching images by text or classifying without training usually means a new vision model per task. CLIP is OpenAI's Contrastive Language-Image Pretraining model judging how well images match text.

It enables zero-shot classification and image-text retrieval, a building block of Stable Diffusion and many multimodal models.

Features



Zero-shot:No training.

Retrieval:Text to image.

Embeddings:General purpose.

Ecosystem:Widely used.