πŸ’»Code of Life

AI 'OptiPrime,' Trained on Biochemical Reactions, Reduces Uncertainty in Prime Editing Design

Nature BiotechnologyΒ·August 13, 2026AI Curation
AI 'OptiPrime,' Trained on Biochemical Reactions, Reduces Uncertainty in Prime Editing Design
✨AI Summary (Beta)Beta

Background

Prime editing (PE) allows for the recording of base substitutions and short insertions/deletions in the genome without DNA double-strand cleavage or separate donor templates. It involves Cas9 nickase, reverse transcriptase, and prime editing guide RNA (pegRNA) targeting DNA, followed by copying the sequence contained in the reverse transcriptase template (RTT) of the pegRNA. Clinical successes have been achieved in correcting NCF1 mutations in patients with chronic granulomatous disease, but practical application still involves extensive design exploration.

This requires adjusting not only the target sequence of the pegRNA but also the length of the RTT and the primer binding site (PBS), as well as silent mutations to avoid mismatch repair (MMR). For challenging targets, hundreds to thousands of combinations must be tested. Existing predictive models, such as PRIDICT and DeepPrime, learn efficiency from large sequence datasets but treat multiple biochemical steps as a single black box. While useful for filtering out low-efficiency candidates, they have limitations in selecting the optimal design from among already good candidates or generalizing to editing methods not included in the training data.

Key Findings

The researchers developed OptiPrime, which directly incorporates the molecular reactions that constitute PE into a mathematical structure. It separately predicts the 'apparent reaction rate' for each step, including DNA binding and cleavage, flap synthesis, heteroduplex resolution, and the binding/dissociation of MMR proteins, using neural networks and linear regression. These rates are then time-integrated using ordinary differential equations to calculate the final editing efficiency. This differs from existing models that input all sequence features at once.

The training data consists of 297,962 PE results obtained under 40 experimental conditions. Among these, the researchers directly measured 74,769 results from 1,290 targets in HEK293T and HeLa cells with different MMR statuses. HetFormer, which represents heteroduplex resolution, was pre-trained on 64 million simulated heteroduplexes. In a 5-fold cross-validation, OptiPrime's average Pearson correlation coefficient was 0.693, and the Spearman rank correlation coefficient was 0.745. The rank correlation coefficients of the comparison models were DeepPrime-FT 0.399 and PRIDICT2.0 series 0.538-0.590.

The biology of MMR was also reproduced within the model. In HeLa cells, PE4, which inhibits MMR, showed a 4.3-fold higher editing rate than PE2 based on the median, while in HEK293T cells with partial MMR deficiency, the increase was 1.5-fold. OptiPrime learned these differences and selected silent mutations that avoid MMR, and even predicted the results of PE3 and twinPE, which use two pegRNAs, without separate training data.

In patient-derived fibroblasts, the COL7A1 p.R185X mutation, which causes recessive dystrophic epidermolysis bullosa, was corrected by 37%. The six candidates selected by PRIDICT2.0 and DeepPrime all scored below 5%. In a Kif1a p.L181F mouse model, only 15 pegRNAs were tested to narrow down the initial candidates, and after four experiments and four weeks, the editing rate of embryonic fibroblasts was increased to a maximum of 64%. When this design was delivered to neonatal mouse brains using a dual AAV9 vector, an average of over 40% correction was observed in the entire cerebral cortex after four weeks, and over 70% in transduced cells.

Significance and Outlook

The strength of OptiPrime lies not only in its prediction accuracy. It uses known reaction pathways as inductive biases in the model, linking each predicted value to a specific step in PE biochemistry. This allows it to adapt to PE3 and twinPE, which are outside the training range, by rearranging the structure. It demonstrates the potential for prime editing design to shift from trial-and-error-based exploration to mechanism-based candidate selection.

However, the research results do not necessarily translate to clinical efficacy. The Kif1a experiment measured editing rates in mice treated shortly after birth, and neurological symptom recovery and long-term safety were not the primary evaluation targets of this study. In some brain regions, the correction rate was low, and the nicking guide RNA of PE3 still left room for trade-offs between efficiency and insertion/deletion byproducts.

The main limitation of the model is that its primary training data came from synthetic reporters randomly inserted into the genome. Cell-type-specific chromatin accessibility and the influence of non-coding regulatory sequences are not directly reflected. Predictive power may be reduced in targets where Cas9 cleavage scores do not match PE efficiency. Combining with chromatin state, splicing, and regulatory element prediction models, and performing prospective validation in various biological tissues, are the next challenges.

Nature Biotechnology, Published online: 12 August 2026; doi:10.1038/s41587-026-03261-7OptiPrime incorporates prime-editing biochemistry to predict prime-editing outcomes.

πŸ’¬Why it matters:

In the drug development field, it can be used as a design tool to reduce the initial exploration, which involves synthesizing and testing hundreds of pegRNAs to correct a specific mutation, to fewer than ten. For example, when a new mutation is identified in a rare disease patient, OptiPrime can prioritize MMR-avoiding silent mutations and RTT/PBS combinations, and then validate only the top candidates in patient-derived cells. In the Kif1a case, the researchers were able to obtain a candidate for in vivo delivery with 15 pegRNAs and 4 weeks of optimization.

In cell therapy development, it may also reduce the cost of screening candidates for primary cells that are difficult to edit, such as patient-derived fibroblasts or T cells. However, before it can be incorporated into the actual manufacturing process, it must undergo separate evaluation of off-target editing, insertion/deletion byproducts, vector toxicity, and long-term durability. The ranking on the web server should be used as a starting point for compressing the number of candidates to be validated, rather than as a value that replaces the experiment.

πŸ’¬ Comments

0 comments
Please log in to comment
Loading...