/
LATENT REFERENCES / TAG1

CLIP

This reference note belongs to Tag1 in Latent References, an archive curated by Keigo Yoshida. Its archive region is Groups · Coachella. The note preserves its source text and links so that readers can trace the material behind the 3D map.

Collection
Tag1
Archive region
Groups · Coachella

Archived reference note

English translation of the archived note. JP shows the original text. Source links and literal code are retained; the translation does not update or independently verify the source claims.

CLIP is a multimodal language-and-image model released by OpenAI in 2021/2. Training the model on a dataset of 4 billion image–text pairs collected from the internet made it possible to improve zero-shot performance on many downstream tasks. https://trail.t.u-tokyo.ac.jp/ja/blog/22-12-02-clip/ https://scrapbox.io/files/6403792a722730001bd4d0d2.png #clip

Source updated 2024-08-04 · Snapshot 2026-10-08

Source links and calculated neighbors

Cosine values measure shared lexical features, not truth, agreement or identical meaning. Original reference links are labeled separately.