/
LATENT REFERENCES / TAG1

Protein Language Models

This reference note belongs to Tag1 in Latent References, an archive curated by Keigo Yoshida. Its archive region is Media. The note preserves its source text and links so that readers can trace the material behind the 3D map.

Collection
Tag1
Archive region
Media

Archived reference note

English translation of the archived note. JP shows the original text. Source links and literal code are retained; the translation does not update or independently verify the source claims.

A protein language model (Protein Language Model: pLM) is an AI technology that treats amino acid sequences as natural-language “sentences” and extracts the grammar and regularities behind them through large-scale deep learning. This technology makes it possible to predict and design protein three-dimensional structures and functions instantaneously. How protein language models work: This technology applies natural-language processing models (LLMs), processing amino acids as follows: Characters (tokens): The 20 types of amino acids are treated as the alphabet or words of natural language. Grammar (context): By reading large quantities of protein sequences, the model learns evolutionarily conserved sequence rules and motifs (word sequences) with particular functions. Representation (embedding): From the order of amino acids, features describing how the protein folds (its 3-dimensional structure) and what it does are mapped into a vector space. Major representative models: Just as natural language has GPT and BERT, powerful models have been developed and released in the protein field. ESM (Evolutionary Scale Modeling): A representative model developed by Meta (formerly Facebook) and others. It learns large quantities of evolutionary data and accurately predicts how amino acid mutations affect protein function and stability. ProtTrans: A group of models applying the architectures of many large language models (such as T5 and BERT) to protein sequences. AlphaFold: Synonymous with protein structure prediction, it incorporates deep learning and language-model-like approaches to generate highly accurate three-dimensional structures directly from amino acid sequences. Applications: AI drug discovery: Designing new proteins (drug candidates) that bind to molecules associated with target diseases from scratch. Enzyme design: Artificially producing new industrially useful functions, such as increasing the activity of enzymes that break down plastics. Mutation effect prediction: Rapidly assessing diseases caused by genetic mutations and mechanisms of drug resistance. #list

Source updated 2026-09-11 · Snapshot 2026-10-08

Source links and calculated neighbors

Cosine values measure shared lexical features, not truth, agreement or identical meaning. Original reference links are labeled separately.