AI Identifies Binding Probabilities to Overcome Gene Editing Errors, and a Design Model Emerges to Prevent Off-Target Effects

Background
CRISPR gene editing and base editors, core technologies for gene therapies, are considered humanity's weapons to fundamentally overcome rare genetic diseases and intractable inherited diseases. However, the off-target effect, where these therapies randomly edit unintended genomic sequences when acting in real-time within a patient's body, has consistently hindered their safety. Previous research efforts focused on precisely analyzing the 3D conformation of proteins to prevent these errors. These efforts primarily involved designing modifications to regulate the charge at the interface where the protein, guide RNA (gRNA), or DNA interact, aiming to reduce non-specific binding. However, this approach often led to a vicious cycle where both the binding affinity for the original target sequence and the editing activity were simultaneously reduced. Furthermore, the subtle conformational changes in protein structure caused by even a single nucleotide mismatch could not be adequately explained by static structural maps alone. There was a critical need for a new computational biology technique to quantitatively analyze the principles of genomic target recognition at the molecular level through dynamic contact information.
Key Findings
A joint research team from Peking University and East China Normal University developed ContactSeek, an AI-driven design analysis framework that easily improves the precision of gene editing tools, based on the powerful computational capabilities of AlphaFold3, a protein structure prediction AI. The researchers focused directly on the contact probability (CP) data, a key parameter outputted by AlphaFold3 during the simulation of various nucleic acid binding forms within protein complexes. CP is a metric that quantifies the probability of two constituent residues, amino acids and nucleotides, physically contacting each other in molecular space. The analysis revealed that this metric remarkably sensitively distinguishes between the binding specificity of the on-target site and off-target sites with only slight sequence differences. To apply this structural calculation ability to real-world experiments, the research team combined the deep learning-based contact information prediction results with experimental dI-profiling data, which represents off-target detection genomic sequencing. Through this process, they identified the amino acid groups that ultimately determine the non-specific binding of the editing tool and named them Consensus Contact Regions (CCRs). CCRs are functionally similar to a variable resistor switch that controls the precision of gene editing. By manipulating this switch, fine-tuning the residues within the CCRs to match the binding affinity of a specific genetic locus, it is possible to create optimal protein variants that preserve activity at the original target site while almost completely blocking unwanted off-target binding. In fact, the research team successfully constructed several improved models of adenine base editor (ABE) and Cas12a cytosine base editor (CBE), the most widely used gene editing tools, with significantly reduced error rates by applying this design system.
Significance and Prospects
The emergence of ContactSeek represents an evolution of AI, shifting its role from simply focusing on static structural analysis to a precise design tool for novel protein engineering. Previous protein modification experiments required a laborious screening process involving thousands of random mutations, followed by individual experimental validation of activity. In contrast, ContactSeek accurately identifies the favorable interaction sites between proteins and genetic sequences through computational science, reducing the experimental validation time for candidate materials by dozens of times. However, there are additional challenges to overcome before this AI-driven design tool can be successfully translated into clinical drug development. It is necessary to demonstrate in various ways whether the contact probabilities derived in molecular virtual space are consistently expressed in the complex ionic states and protein crowding environments within actual cell cytoplasm. The potential for immune responses induced by structural changes and compatibility with in vivo delivery systems are also challenges to be addressed. The research team plans to expand the target gene editing tools and focus on in vivo pharmacological activity validation in the future.
Nature, Published online: 22 July 2026; doi:10.1038/s41586-026-10794-zContactSeek is an AlphaFold3-driven model that can improve the precision of genome-editing tools.
This research provides a practical roadmap to significantly accelerate the commercialization of gene therapies. In particular, it can dramatically improve safety in the development of personalized therapies for patients with inherited diseases. Specifically, in the case of diseases caused by a single nucleotide mutation, such as Huntington's disease or inherited retinal degeneration, it will be possible to design personalized base editors in real-time to minimize off-target effects. This involves acquiring off-target risk sequence information from patient-derived cells, inputting it into the ContactSeek analysis system, and inducing the creation of a custom Cas protein that is designed to respond specifically to that sequence. This is expected to significantly increase the likelihood of clinical approval for gene editing by eliminating the risk of the therapeutic agent targeting the wrong gene and causing cancer.