ESMC (EvolutionaryScale Cambrian)
Named after the Cambrian period, when the most rapid evolutionary explosion in human history occurred, ESMC is a next-generation foundation model announced in May 2026 by a joint research team from the Chan Zuckerberg Biohub and EvolutionaryScale. This model was trained on approximately 2.8 billion amino acid sequences collected from living organisms on Earth in order to understand the blueprint of the vast protein universe, much like the billions of pages of text written by humans.
ESMC, named after the Cambrian period, during which the most rapid evolutionary explosion in human history occurred, is a next-generation foundation model for biology, released in May 2026 by a joint research team from the Chan Zuckerberg Biohub and EvolutionaryScale. This model was trained on approximately 2.8 billion amino acid sequences collected from living organisms on Earth to understand the blueprint of the vast protein universe. Similar to how a large language model (LLM) learns the grammar of a language by reading billions of pages of text written by humans, ESMC adopts a masked language modeling architecture that allows it to independently learn the physical and chemical folding rules and complex interaction mechanisms of proteins through the arrangement of protein amino acid sequences.
Existing bioinformatics analysis tools primarily relied on multiple sequence alignment methods, which focused on static and local analysis by comparing sequences conserved between species, limiting their ability to predict the dynamic and three-dimensional interactions created by new amino acid combinations. ESMC goes beyond this static comparative analysis by mapping the entire biological sequence into a multi-dimensional continuous embedding space, organically capturing the evolutionary context and structural potential of proteins. This is similar to how ESMC understands the multi-layered nuances of a word at once by grasping the context in numerous novels, whereas traditional comparative biology methods were like looking up the meaning of a word one by one in an old dictionary. This tool is particularly useful because it is provided in an open-weight format, allowing researchers around the world to flexibly deploy a model of appropriate size, from 300M to 6B parameters, in their local environment, tailored to their computational resources and research purposes, to maximize protein language representation performance.
Researchers in the field of biotechnology can place ESMC as a key predictive engine in the process of discovering artificial protein binders that strongly bind to new drug target proteins. For example, based on the three-dimensional binding interface information of a specific cancer cell surface receptor, candidate amino acid sequences that can specifically bind to that receptor can be designed and screened using the language model evaluation score (likelihood score) of ESMC, significantly narrowing the scope of candidates for actual synthesis experiments. Furthermore, by integrating a structure prediction pipeline with the ESMFold2 model for structural biology research, it can predict a precise 3D structure at the all-atom resolution within one second after inputting sequence information, allowing for rapid preliminary validation in a virtual space of the crystal structure analysis process, which previously took several months in traditional laboratories. This integrated analysis technique is usefully applied in the large-scale variant-induced library screening stage and ultimately becomes a key framework that improves the efficiency of precision medicine and next-generation bio-pharmaceutical design by more than tenfold, such as controlling the activity of target proteins and elucidating resistance mechanisms.
💻 System Requirements
Minimum NVIDIA GPU VRAM 8GB (for ESMC-300M inference), recommended 16GB or higher (for ESMC-6B FP16 inference). When running on CPU alone, it takes 1 to 10~30 seconds per amino acid sequence, resulting in slow performance.
When downloading model weights, ensure sufficient disk space is available from 1GB (300M) to 15GB (6B).
⚡ Installation
4-1. Quick Start
pip install esm@git+https://github.com/Biohub/esm.git@main
4-2. Detailed installation
# Example of model loading and inference through Hugging Face Transformers
import torch
from transformers import AutoModelForMaskedLM, AutoTokenizer
# 1. Load the model and tokenizer
model_id = "biohub/ESMC-6B" # or "biohub/ESMC-300M"
model = AutoModelForMaskedLM.from_pretrained(model_id, device_map="auto").eval()
tokenizer = AutoTokenizer.from_pretrained(model_id)
# 2. Define the amino acid sequence of the protein to be analyzed.
sequences = ["MSKGEELFTGVVPILVELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTLVTTFSYGVQCFSRYPDHMKQHDFFKSAMPEGYVQERTIFFKDDGNYKTRAEVKFEGDTLVNRIELKGIDFKEDGNILGHKLEYNYNSHNVYIMADKQKNGIKVNFKIRHNIEDGSVQLADHYQQNTPIGDGPVLLPDNHYLSTQSALSKDPNEKRDHMVLLEFVTAAGITHGMDELYK"]
# 3. Input data preprocessing and device tensor transfer
inputs = tokenizer(sequences, return_tensors="pt", padding=True)
inputs = {k: v.to(model.device) for k, v in inputs.items()}
# 4. Run inference mode and output protein embedding/prediction logs
with torch.inference_mode():
output = model(inputs)
print(output.logits)
FAQ
What is ESMC (EvolutionaryScale Cambrian)?
ESMC, named after the Cambrian period, during which the most rapid evolutionary explosion in human history occurred, is a next-generation foundation model for biology, released in May 2026 by a joint research team from the Chan Zuckerberg Biohub and EvolutionaryScale. This model was trained on approximately 2.8 billion amino acid sequences collected from living organisms on Earth to understand the blueprint of the vast protein universe. Similar to how a large language model (LLM) learns the grammar of a language by reading billions of pages of text written by humans, ESMC adopts a masked language modeling architecture that allows it to independently learn the physical and chemical folding rules and complex interaction mechanisms of proteins through the arrangement of protein amino acid sequences. Existing bioinformatics analysis tools primarily relied on multiple sequence alignment methods, which focused on static and local analysis by comparing sequences conserved between species, limiting their ability to predict the dynamic and three-dimensional interactions created by new amino acid combinations. ESMC goes beyond this static comparative analysis by mapping the entire biological sequence into a multi-dimensional continuous embedding space, organically capturing the evolutionary context and structural potential of proteins. This is similar to how ESMC understands the multi-layered nuances of a word at once by grasping the context in numerous novels, whereas traditional comparative biology methods were like looking up the meaning of a word one by one in an old dictionary. This tool is particularly useful because it is provided in an open-weight format, allowing researchers around the world to flexibly deploy a model of appropriate size, from 300M to 6B parameters, in their local environment, tailored to their computational resources and research purposes, to maximize protein language representation performance. Researchers in the field of biotechnology can place ESMC as a key predictive engine in the process of discovering artificial protein binders that strongly bind to new drug target proteins. For example, based on the three-dimensional binding interface information of a specific cancer cell surface receptor, candidate amino acid sequences that can specifically bind to that receptor can be designed and screened using the language model evaluation score (likelihood score) of ESMC, significantly narrowing the scope of candidates for actual synthesis experiments. Furthermore, by integrating a structure prediction pipeline with the ESMFold2 model for structural biology research, it can predict a precise 3D structure at the all-atom resolution within one second after inputting sequence information, allowing for rapid preliminary validation in a virtual space of the crystal structure analysis process, which previously took several months in traditional laboratories. This integrated analysis technique is usefully applied in the large-scale variant-induced library screening stage and ultimately becomes a key framework that improves the efficiency of precision medicine and next-generation bio-pharmaceutical design by more than tenfold, such as controlling the activity of target proteins and elucidating resistance mechanisms.
When should I use ESMC (EvolutionaryScale Cambrian)?
Named after the Cambrian period, when the most rapid evolutionary explosion in human history occurred, ESMC is a next-generation foundation model announced in May 2026 by a joint research team from the Chan Zuckerberg Biohub and EvolutionaryScale. This model was trained on approximately 2.8 billion amino acid sequences collected from living organisms on Earth in order to understand the blueprint of the vast protein universe, much like the billions of pages of text written by humans.
📝 Update Notes
- vv3.4.1.post19/16/2026
이번 업데이트는 HF, esm, mlx_lm 등 서로 다른 구현 방식 간의 호환성을 개선하여, 하나의 모델 가중치만으로도 다양한 환경에서 모델을 편리하게 활용할 수 있게 해줍니다. 특히 EsmFold2의 실험적 기능과 SAE(Sparse Autoencoder) 구현에서 발견된 버그를 해결하여, 단백질 구조 예측 및 특징 분석의 신뢰도를 높였습니다. 연구자분들은 이제 컴퓨팅 환경에 구애받지 않고 더욱 안정적이고 일관된 방식으로 생명공학 데이터 분석을 수행할 수 있습니다.
- vv3.4.19/9/2026
이번 ESMC v3.4.1 업데이트에서는 FoldCP 지원이 추가되어 단백질 구조 분석 작업의 유연성이 한층 높아졌습니다. 특히 Hugging Face 스타일의 ESMFold2 체크포인트를 직접 불러올 수 있게 되어, 기존에 공개된 고성능 모델들을 ESMC 환경에서 더욱 손쉽게 활용할 수 있습니다. 여러 버그 수정으로 안정성까지 개선되었으니, 모델 활용 범위를 넓히고 더욱 효율적인 단백질 언어 모델 연구를 진행하고 싶은 연구자분들께 이번 업데이트를 추천드려요.
- vv3.4.09/2/2026
이번 업데이트를 통해 최신 단백질 언어 모델인 ESMC와 ESMFold2를 완벽하게 지원하여, 더욱 정교한 단백질 구조 예측 및 분석이 가능해졌습니다. 특히 Hugging Face 포트와의 호환 레이어가 추가되어, 익숙한 Hugging Face 생태계의 모델들을 더욱 손쉽게 연구 워크플로우에 통합할 수 있습니다. 기존 레거시 모델들에 대한 하위 호환성도 유지되므로, 기존 연구 파이프라인을 변경할 걱정 없이 안정적으로 최신 기능을 도입해 보세요.
- vv3.2.2.post27/4/2026
이번 업데이트의 핵심은 단백질, 핵산(DNA/RNA), 리간드를 모두 포함하는 '모든 원자 수준(All-Atom)'의 분자 복합체 구조 예측 기능이 도입된 점이에요. 덕분에 단백질과 리간드 간의 결합이나 단백질-핵산 상호작용 같은 복잡한 생체 분자 시스템을 더욱 정밀하게 분석할 수 있습니다. 또한 MSA 처리 효율을 높인 새로운 도구들이 추가되어 대규모 서열 데이터 분석의 속도와 정확도가 한층 향상되었어요. 구조 예측의 안정성까지 개선된 만큼, 분자 상호작용 연구를 진행하는 연구원님들께 이번 업데이트를 적극 추천드려요.
🧪 Related Code of Life
No related Code of Life posts yet.