LATENT REFERENCES / TAG1
Protein Language Models
Original title: タンパク質 言語モデル
This reference note belongs to Tag1 in Latent References, an archive curated by Keigo Yoshida. Its archive region is Media. The note preserves its source text and links so that readers can trace the material behind the 3D map.
- Collection
- Tag1
- Archive region
- Media
Archived reference note
English translation of the archived note. JP shows the original text. Source links and literal code are retained; the translation does not update or independently verify the source claims.
A protein language model (Protein Language Model: pLM) is an AI technology that treats amino acid sequences as natural-language “sentences” and extracts the grammar and regularities behind them through large-scale deep learning. This technology makes it possible to predict and design protein three-dimensional structures and functions instantaneously. How protein language models work: This technology applies natural-language processing models (LLMs), processing amino acids as follows: Characters (tokens): The 20 types of amino acids are treated as the alphabet or words of natural language. Grammar (context): By reading large quantities of protein sequences, the model learns evolutionarily conserved sequence rules and motifs (word sequences) with particular functions. Representation (embedding): From the order of amino acids, features describing how the protein folds (its 3-dimensional structure) and what it does are mapped into a vector space. Major representative models: Just as natural language has GPT and BERT, powerful models have been developed and released in the protein field. ESM (Evolutionary Scale Modeling): A representative model developed by Meta (formerly Facebook) and others. It learns large quantities of evolutionary data and accurately predicts how amino acid mutations affect protein function and stability. ProtTrans: A group of models applying the architectures of many large language models (such as T5 and BERT) to protein sequences. AlphaFold: Synonymous with protein structure prediction, it incorporates deep learning and language-model-like approaches to generate highly accurate three-dimensional structures directly from amino acid sequences. Applications: AI drug discovery: Designing new proteins (drug candidates) that bind to molecules associated with target diseases from scratch. Enzyme design: Artificially producing new industrially useful functions, such as increasing the activity of enzymes that break down plastics. Mutation effect prediction: Rapidly assessing diseases caused by genetic mutations and mechanisms of drug resistance.
#list
タンパク質の言語モデル(Protein Language Model: pLM)とは、アミノ酸配列を自然言語の「文章」と見なし、大規模な深層学習によってその背後にある文法や規則性を抽出するAI技術です。この技術により、タンパク質の立体構造や機能を瞬時に予測・設計することが可能になっています。タンパク質言語モデルの仕組み自然言語の処理モデル(LLM)を応用した技術で、以下のようにアミノ酸を処理します:文字(トークン): 20種類のアミノ酸を自然言語におけるアルファベットや単語として扱います。文法(コンテクスト): 大量のタンパク質配列を読み込ませることで、進化的に保存されている配列のルールや、特定の機能を持つモチーフ(単語の並び)を学習します。表現(埋め込み): アミノ酸の並び順から、そのタンパク質がどう折りたたまれるか(3次元構造)や、どんな働きをするかの特徴量をベクトル空間に落とし込みます。主な代表的モデル自然言語でいうGPTやBERTのように、タンパク質分野でも強力なモデルが開発・公開されています。ESM (Evolutionary Scale Modeling): Meta(旧Facebook)などが開発した代表的なモデルです。大量の進化データを学習し、アミノ酸の変異がタンパク質の機能や安定性に与える影響を高精度で予測します。ProtTrans: 多数の大規模言語モデル(T5やBERTなど)のアーキテクチャをタンパク質配列に応用したモデル群です。AlphaFold: タンパク質構造予測の代名詞ですが、深層学習や言語モデル的なアプローチを組み込むことで、アミノ酸配列から直接立体構造を高精度で生成します。何に応用されているかAI創薬: 標的となる病気の分子に結合する新しいタンパク質(医薬品候補)をゼロから設計します。酵素デザイン: プラスチックを分解する酵素の活性を高めるなど、工業的に役立つ新しい機能を人工的に生み出します。変異効果予測: 遺伝子の変異が引き起こす病気や、薬剤耐性のメカニズムを素早く評価します。
#list
Source updated 2026-09-11 · Snapshot 2026-10-08
Source links and calculated neighbors
Cosine values measure shared lexical features, not truth, agreement or identical meaning. Original reference links are labeled separately.
- AlphaFold: Installing Protein Structure PredictionComputed lexical cosine similarity 0.270 · shared title, text, tags and references
- ESM Metagenomic Atlas: The first view of the ‘dark matter’ of the protein universeComputed lexical cosine similarity 0.140 · shared title, text, tags and references
- Directed EvolutionComputed lexical cosine similarity 0.125 · shared title, text, tags and references
- Matsuo Lab LLM CourseComputed lexical cosine similarity 0.095 · shared title, text, tags and references
- Orphan ReceptorsComputed lexical cosine similarity 0.092 · shared title, text, tags and references