LATENT REFERENCES / TAG1
CLIP
Original title: CLIP
This reference note belongs to Tag1 in Latent References, an archive curated by Keigo Yoshida. Its archive region is Groups · Coachella. The note preserves its source text and links so that readers can trace the material behind the 3D map.
- Collection
- Tag1
- Archive region
- Groups · Coachella
Archived reference note
English translation of the archived note. JP shows the original text. Source links and literal code are retained; the translation does not update or independently verify the source claims.
CLIP is a multimodal language-and-image model released by OpenAI in 2021/2.
Training the model on a dataset of 4 billion image–text pairs collected from the internet made it possible to improve zero-shot performance on many downstream tasks.
https://trail.t.u-tokyo.ac.jp/ja/blog/22-12-02-clip/
https://scrapbox.io/files/6403792a722730001bd4d0d2.png
#clip
CLIPは,2021年2月にOpenAIによって公開された,言語と画像のマルチモーダルモデルです.
インターネットから集めた画像とテキストの40億ペアからなるデータセットからモデルを学習することで,多くの下流タスクに対するゼロショット性能を高めることが可能になりました.
https://trail.t.u-tokyo.ac.jp/ja/blog/22-12-02-clip/
https://scrapbox.io/files/6403792a722730001bd4d0d2.png
#clip
Source updated 2024-08-04 · Snapshot 2026-10-08
Source links and calculated neighbors
Cosine values measure shared lexical features, not truth, agreement or identical meaning. Original reference links are labeled separately.
- research-fieldSource reference link in the archive
- High-Quality Data Is Disappearing from the InternetComputed lexical cosine similarity 0.181 · shared title, text, tags and references
- Zero-Shot PredictionComputed lexical cosine similarity 0.137 · shared title, text, tags and references