ESMC (EvolutionaryScale Cambrian)
Named after the Cambrian period, when the most rapid evolutionary explosion in human history occurred, ESMC is a next-generation foundation model announced in May 2026 by a joint research team from the Chan Zuckerberg Biohub and EvolutionaryScale. This model was trained on approximately 2.8 billion amino acid sequences collected from living organisms on Earth in order to understand the blueprint of the vast protein universe, much like the billions of pages of text written by humans.
ESMC, named after the Cambrian period, during which the most rapid evolutionary explosion in human history occurred, is a next-generation foundation model for biology, released in May 2026 by a joint research team from the Chan Zuckerberg Biohub and EvolutionaryScale. This model was trained on approximately 2.8 billion amino acid sequences collected from living organisms on Earth to understand the blueprint of the vast protein universe. Similar to how a large language model (LLM) learns the grammar of a language by reading billions of pages of text written by humans, ESMC adopts a masked language modeling architecture that allows it to independently learn the physical and chemical folding rules and complex interaction mechanisms of proteins through the arrangement of protein amino acid sequences.
Existing bioinformatics analysis tools primarily relied on multiple sequence alignment methods, which focused on static and local analysis by comparing sequences conserved between species, limiting their ability to predict the dynamic and three-dimensional interactions created by new amino acid combinations. ESMC goes beyond this static comparative analysis by mapping the entire biological sequence into a multi-dimensional continuous embedding space, organically capturing the evolutionary context and structural potential of proteins. This is similar to how ESMC understands the multi-layered nuances of a word at once by grasping the context in numerous novels, whereas traditional comparative biology methods were like looking up the meaning of a word one by one in an old dictionary. This tool is particularly useful because it is provided in an open-weight format, allowing researchers around the world to flexibly deploy a model of appropriate size, from 300M to 6B parameters, in their local environment, tailored to their computational resources and research purposes, to maximize protein language representation performance.
Researchers in the field of biotechnology can place ESMC as a key predictive engine in the process of discovering artificial protein binders that strongly bind to new drug target proteins. For example, based on the three-dimensional binding interface information of a specific cancer cell surface receptor, candidate amino acid sequences that can specifically bind to that receptor can be designed and screened using the language model evaluation score (likelihood score) of ESMC, significantly narrowing the scope of candidates for actual synthesis experiments. Furthermore, by integrating a structure prediction pipeline with the ESMFold2 model for structural biology research, it can predict a precise 3D structure at the all-atom resolution within one second after inputting sequence information, allowing for rapid preliminary validation in a virtual space of the crystal structure analysis process, which previously took several months in traditional laboratories. This integrated analysis technique is usefully applied in the large-scale variant-induced library screening stage and ultimately becomes a key framework that improves the efficiency of precision medicine and next-generation bio-pharmaceutical design by more than tenfold, such as controlling the activity of target proteins and elucidating resistance mechanisms.
๐ป System Requirements
Minimum NVIDIA GPU VRAM 8GB (for ESMC-300M inference), recommended 16GB or higher (for ESMC-6B FP16 inference). When running on CPU alone, it takes 1 to 10~30 seconds per amino acid sequence, resulting in slow performance.
When downloading model weights, ensure sufficient disk space is available from 1GB (300M) to 15GB (6B).
โก Installation
4-1. Quick Start
pip install esm@git+https://github.com/Biohub/esm.git@main
4-2. Detailed installation
# Example of model loading and inference through Hugging Face Transformers
import torch
from transformers import AutoModelForMaskedLM, AutoTokenizer
# 1. Load the model and tokenizer
model_id = "biohub/ESMC-6B" # or "biohub/ESMC-300M"
model = AutoModelForMaskedLM.from_pretrained(model_id, device_map="auto").eval()
tokenizer = AutoTokenizer.from_pretrained(model_id)
# 2. Define the amino acid sequence of the protein to be analyzed.
sequences = ["MSKGEELFTGVVPILVELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTLVTTFSYGVQCFSRYPDHMKQHDFFKSAMPEGYVQERTIFFKDDGNYKTRAEVKFEGDTLVNRIELKGIDFKEDGNILGHKLEYNYNSHNVYIMADKQKNGIKVNFKIRHNIEDGSVQLADHYQQNTPIGDGPVLLPDNHYLSTQSALSKDPNEKRDHMVLLEFVTAAGITHGMDELYK"]
# 3. Input data preprocessing and device tensor transfer
inputs = tokenizer(sequences, return_tensors="pt", padding=True)
inputs = {k: v.to(model.device) for k, v in inputs.items()}
# 4. Run inference mode and output protein embedding/prediction logs
with torch.inference_mode():
output = model(inputs)
print(output.logits)
FAQ
What is ESMC (EvolutionaryScale Cambrian)?
ESMC, named after the Cambrian period, during which the most rapid evolutionary explosion in human history occurred, is a next-generation foundation model for biology, released in May 2026 by a joint research team from the Chan Zuckerberg Biohub and EvolutionaryScale. This model was trained on approximately 2.8 billion amino acid sequences collected from living organisms on Earth to understand the blueprint of the vast protein universe. Similar to how a large language model (LLM) learns the grammar of a language by reading billions of pages of text written by humans, ESMC adopts a masked language modeling architecture that allows it to independently learn the physical and chemical folding rules and complex interaction mechanisms of proteins through the arrangement of protein amino acid sequences. Existing bioinformatics analysis tools primarily relied on multiple sequence alignment methods, which focused on static and local analysis by comparing sequences conserved between species, limiting their ability to predict the dynamic and three-dimensional interactions created by new amino acid combinations. ESMC goes beyond this static comparative analysis by mapping the entire biological sequence into a multi-dimensional continuous embedding space, organically capturing the evolutionary context and structural potential of proteins. This is similar to how ESMC understands the multi-layered nuances of a word at once by grasping the context in numerous novels, whereas traditional comparative biology methods were like looking up the meaning of a word one by one in an old dictionary. This tool is particularly useful because it is provided in an open-weight format, allowing researchers around the world to flexibly deploy a model of appropriate size, from 300M to 6B parameters, in their local environment, tailored to their computational resources and research purposes, to maximize protein language representation performance. Researchers in the field of biotechnology can place ESMC as a key predictive engine in the process of discovering artificial protein binders that strongly bind to new drug target proteins. For example, based on the three-dimensional binding interface information of a specific cancer cell surface receptor, candidate amino acid sequences that can specifically bind to that receptor can be designed and screened using the language model evaluation score (likelihood score) of ESMC, significantly narrowing the scope of candidates for actual synthesis experiments. Furthermore, by integrating a structure prediction pipeline with the ESMFold2 model for structural biology research, it can predict a precise 3D structure at the all-atom resolution within one second after inputting sequence information, allowing for rapid preliminary validation in a virtual space of the crystal structure analysis process, which previously took several months in traditional laboratories. This integrated analysis technique is usefully applied in the large-scale variant-induced library screening stage and ultimately becomes a key framework that improves the efficiency of precision medicine and next-generation bio-pharmaceutical design by more than tenfold, such as controlling the activity of target proteins and elucidating resistance mechanisms.
When should I use ESMC (EvolutionaryScale Cambrian)?
Named after the Cambrian period, when the most rapid evolutionary explosion in human history occurred, ESMC is a next-generation foundation model announced in May 2026 by a joint research team from the Chan Zuckerberg Biohub and EvolutionaryScale. This model was trained on approximately 2.8 billion amino acid sequences collected from living organisms on Earth in order to understand the blueprint of the vast protein universe, much like the billions of pages of text written by humans.
๐ Update Notes
No update notes yet.
๐งช Related Code of Life
No related Code of Life posts yet.