LATENT REFERENCES / TAG1
ESM Metagenomic Atlas: The first view of the ‘dark matter’ of the protein universe
Original title: ESM Metagenomic Atlas: The first view of the ‘dark matter’ of the protein universe
This reference note belongs to Tag1 in Latent References, an archive curated by Keigo Yoshida. Its archive region is Media. The note preserves its source text and links so that readers can trace the material behind the 3D map.
- Collection
- Tag1
- Archive region
- Media
Archived reference note
English translation of the archived note. JP shows the original text. Source links and literal code are retained; the translation does not update or independently verify the source claims.
https://ai.meta.com/blog/protein-folding-esmfold-metagenomics/
Meta AI created the first database revealing structures of the metagenomic world on the scale of hundreds of millions of proteins. These proteins exist in soil microbes, the deep ocean, and even our bodies, far outnumbering the microbes constituting animals and plants. Yet they are the least understood proteins on Earth.
Deciphering metagenomic structures can resolve longstanding evolutionary-history mysteries and discover proteins potentially useful for curing diseases, cleaning environments, and producing cleaner energy.
Predicting structures at this scale requires a breakthrough in protein folding speed. We trained a large language model to learn evolutionary patterns directly from protein sequences and generate accurate structure predictions end to end. Predictions are up to 60 times faster than the current state of the art while maintaining accuracy, and our approach can scale to much larger databases.
We are now sharing our model, research paper, a database of more than 600 million metagenomic structures, and an API allowing scientists to easily retrieve specific protein structures relevant to their research.
Here, ESM metagenomic
#list
https://ai.meta.com/blog/protein-folding-esmfold-metagenomics/
Meta AIは、数億ものタンパク質規模でメタゲノム世界の構造を明らかにする初のデータベースを作成しました。これらのタンパク質は、土壌中の微生物や海の深部、さらには私たちの体内にも存在し、動植物を構成する微生物をはるかに上回っています。しかし、彼らは地球上で最も理解されていないタンパク質です。
メタゲノム構造を解読することで、進化史の長年にわたる謎を解き明かし、疾患の治癒や環境の浄化、よりクリーンなエネルギーの生成に役立つ可能性のあるタンパク質を発見することができます。
このスケールで構造予測を行うためには、タンパク質の折りたたみ速度における画期的な進歩が必要です。私たちは、大規模な言語モデルを訓練し、タンパク質の配列から直接、進化パターンを学習し、端から端まで正確な構造予測を生成しました。予測は、現在の最先端技術より最大60倍高速でありながら、精度を維持し、当社のアプローチははるかに大規模なデータベースへ拡張可能です。
私たちは現在、当社のモデル、研究論文、そして6億を超えるメタゲノム構造のデータベース、さらに科学者が研究に関連する特定のタンパク質構造を容易に取得できるAPIを共有しています。
こちらでESMメタゲノミック
#list
Source updated 2026-09-11 · Snapshot 2026-10-08
Source links and calculated neighbors
Cosine values measure shared lexical features, not truth, agreement or identical meaning. Original reference links are labeled separately.
- AlphaFold: Installing Protein Structure PredictionComputed lexical cosine similarity 0.150 · shared title, text, tags and references
- Protein Language ModelsComputed lexical cosine similarity 0.140 · shared title, text, tags and references
- World Atlas of Language StructuresComputed lexical cosine similarity 0.106 · shared title, text, tags and references
- CPPN: Compositional Pattern Producing NetworkComputed lexical cosine similarity 0.092 · shared title, text, tags and references
- ISMIR MIDI DatabaseComputed lexical cosine similarity 0.089 · shared title, text, tags and references