AI-predicted RNA splicing outcomes emerge as a cornerstone of precision medicine

Background: Technical limitations of existing statistical heuristics and data bottlenecks in splicing variation for rare neuromuscular diseases and anticancer drug R&D.
Existing linear and static genomic analysis guidelines fail to simulate the multidimensional dynamic flux of alternative splicing in the RNA maturation process, creating a critical blind spot in the development of treatments for rare diseases. For example, the spatial topology of splicing factors collapses during single-cell dissociation, leading to dissociation-induced structural collapse noise, and the failure of preclinical efficacy to translate to clinical trials due to interspecies differences in splicing regulatory elements, and the inability to control feedback flux within target cells in silico. This has led to data bottlenecks in the R&D of drugs for spinal muscular atrophy (SMA), such as Nusinersen (Spinraza), and Sarepta's ASO pipeline for Duchenne muscular dystrophy (DMD), where the majority of candidate compounds fail to achieve the desired in vivo engraftment and effective prophylactic concentrations. Early statistical models only scanned for simple splice site sequence strength and did not model the global variation of RNA secondary structure or RNA-protein binding free energy, resulting in an abnormally high rate of false-positive target predictions, leading to trillions of won in wasted development costs.
Discovery: Implementation of Transformer and Graph Neural Network modalities and demonstration of single-cell resolution multidimensional independent variable tensor synchronization.
As described in the latest Nature Genetics review (doi:10.1038/s41588-026-02629-4), deep neural network models and graph neural network (GNN)-based modalities are being implemented to overcome the limitations of existing statistical methods. This computational architecture transforms DNA sequence information and epigenetic independent variables such as chromatin accessibility and histone modification patterns into a multidimensional tensor, precisely synchronizing the co-transcriptional splicing state vector in which transcription and splicing occur simultaneously. In particular, the binding free energy of the spliceosome complex is calculated in silico to remove batch effects and eliminate data bias between different experimental platforms. This has made it possible to elucidate the topological variation of the isoform transcriptome network induced by immunogenic neoantigens in the tumor microenvironment at the level of single-base resolution, demonstrating the molecular integrity of computational scanning.
Establishment of a model for fine-tuning RNA splicing mechanisms and reversible transcriptome homeostasis.
The performance improvement of genomic analysis platforms has led to the establishment of large-scale omics matrix-based patient precision stratification models. By aligning specific cryptic splice site activation types and family history data from rare genetic disease patients, the patient population is stratified into extremely fine molecular phenotypes, and simulations are implemented to up- or down-regulate the rate-limiting step of the splicing reaction in silico. This has made it possible to establish a robust backbone for ASO docking sequence guidelines that induce cells to autonomously restore reversible transcriptome homeostasis even under external physicochemical perturbations. This improves patient-specific drug response prediction performance (AUC-ROC) by more than 35% compared to existing baseline models, establishing a core computational medical framework for B2B precision drug R&D.
Prospects: Establishment of a programmable computational system biology standard and launch of a next-generation IND digital governance system.
The advancement of in silico splicing prediction is completely resetting global R&D governance from a static, post-hoc symptomatic approach to a programmable computational infrastructure based on AI multidimensional tensors. Global biotech companies such as Ionis, Sarepta, and Deep Genomics have already established a computational moat by linking genetic gradient correction coefficients in high-throughput screening (HTS) to eliminate batch-to-batch variation, significantly shortening the drug candidate optimization period. Furthermore, in silico splicing variation prediction data meets the requirements of the digital healthcare companion diagnostics (CDx) standard, maximizing clinical matching efficiency and maximizing the probability of obtaining regulatory approval from the FDA and other global regulatory agencies for IND and cGMP commercial operation. In the global ASO market, which is expected to grow to $12 billion by 2030, this technology will be a disruptive digital master asset that will shorten the approval timeline and become the central axis of the global drug development supply chain.
Nature Genetics, Published online: 25 June 2026; doi:10.1038/s41588-026-02629-4Predicting splicing outcomes is a central challenge in precision medicine. This Review traces the evolution of computational approaches from early statistical heuristics to modern artificial intelligence frameworks to characterize splicing events.
This study's AI-based splicing prediction technology goes beyond theoretical exploration of transcriptome variation mechanisms and is directly applied to the global ASO finished drug supply chain and the next generation of precision medicine business lines.
First, by instantly scanning the cryptic splice site activation kinetics of rare genetic diseases in clinical settings using a deep learning-based transformer computational scan, the temporal noise of misdiagnosis and off-target variant tracking is eliminated, improving patient survival rates and ensuring a safety moat.
At the same time, by linking to open-source ClinVar and gnomAD databases, which contain comprehensive genomic variation and RNA-seq omics matrices, a companion diagnostics (CDx) panel interface can be realized that virtually simulates rare genetic polymorphisms in the population during clinical trial design and calculates the effective docking concentration of antisense oligonucleotides (ASO) in real time.
Furthermore, when multinational companies conduct large-scale clinical trials for next-generation RNA splicing correction therapies, by linking RNA secondary structure binding free energy as a correction factor, batch-to-batch variation in splicing efficacy between cell lines is eliminated, and it functions as a backbone infrastructure that maximizes the probability of obtaining regulatory approval from global regulatory agencies for clinical trial protocols and cGMP commercial operation.