Deep learning predicts extreme regulatory effects of noncoding variants across fetal and adult cell types

Background: Limitations of Single-Cell-Type Static GWAS Interpretation and Bottlenecks in Non-coding Regulatory Variant Data for Neurodevelopmental-Cardiovascular Disease R&D
Non-coding regions account for more than 98% of the human genome, and over 90% of disease-associated variants identified by GWAS are located in these non-coding regions. However, existing analytical paradigms have been constrained by a linear, coding-sequence-centric model of variant-to-phenotype connections. Despite the accumulation of cell-type-specific chromatin state catalogs by large-scale epigenomic consortia such as ENCODE and Roadmap Epigenomics, these datasets remain at the level of bulk tissue, providing only averaged accessibility profiles and failing to computationally control for noise arising from cell dissociation-induced structural disruption. There is a lack of high-resolution, quantitative, tensor-level predictive resources to determine how non-coding variants reconfigure the transcription factor binding free energy landscape in space-time-specific regulatory contexts, such as neurons and cardiomyocytes during fetal development. This deficiency creates a critical double bottleneck in R&D pipelines for the pathogenicity assessment of de novo mutations in autism spectrum disorder (ASD), common variants in schizophrenia, and non-coding variants in congenital heart disease, simultaneously generating false positive noise and false negative blind spots. Even the distribution of effect sizes for ultra-rare variants (based on gnomAD v4) across cell types lacks a baseline reference frame, rendering the governance of non-coding variant functional interpretation in multinational precision medicine pipelines virtually inoperable.
Discovery: ChromBPNet Deep Learning Sequence Model with 3 Billion Predicted Tensors in Operation and Empirical Demonstration of Synchronized Non-coding Variant Effects on Chromatin Accessibility in Fetal-to-Adult Multi-Cell Contexts
The Montgomery and Kundaje research groups operated the ChromBPNet deep learning sequence model on fetal brain and heart, and adult multi-organ single-cell ATAC-seq datasets, generating 3 billion predicted tensors of non-coding variant-chromatin accessibility effects, and constructed an unprecedented regulatory effect landscape across the human development-disease spectrum. By serially linking DeepLIFT contribution scoring and the TF-MoDISco transcription factor motif discovery pipeline, they elucidated the quantitative trajectory of how variants disrupt transcription factor occupancy binding free energy at the base-pair resolution in the cis-regulatory modules of key neurodevelopmental-neurodegenerative genes such as MEF2, TBR1, NeuroD1, CNTNAP2, FOXG1, RFX family, EBF3, ARID1B, and PICALM. A key algorithmic innovation is FLARE (Functional Lasso Analysis of Regulatory Evolution), which integrates ChromBPNet predicted regulatory effects, phylogenetic evolutionary constraints, and population genetic allele frequency spectra within a lasso regularization framework to destructively separate non-coding variants with extreme regulatory effects from benign background variants, establishing an in silico, priority-based pre-screening system. Empirical results reveal a stark dichotomy: common variants (MAF > 1%) exhibit highly cell-type-specific regulatory effects, while ultra-rare variants exhibit larger effect sizes and broader impact radii across multiple cell types. This finding completely surpasses existing single-cell-type, simple linear models, demonstrating that purifying selection pressure on regulatory variants is strongest in fetal neurons, as revealed by the downstream transcriptomic network topological variation curves.
Establishment of a Model for Precision Stratification of Non-coding Regulatory Network Topology and Reversible Developmental-Disease Continuum
By cross-integrating the 3 billion predicted omics matrices generated by the FLARE framework with GTEx multi-organ expression data, the ROSMAP brain expression outlier cohort, and the autism family trio de novo mutation dataset, they achieved precision stratification of disease-specific and cell-type-specific molecular phenotypes. They confirmed that de novo regulatory mutations in autism cases were significantly enriched in the top ranks of FLARE, and that the heritability of schizophrenia GWAS was significantly enriched in the high-ranking FLARE variant regions, cross-validating the molecular integrity of the framework across the common-rare variant continuum. The myocardial chromatin accessibility model extracted from fetal heart single-cell ATAC-seq was used to reverse-calculate the regulatory effects of non-coding mutations in congenital heart disease, directly linking them to the heart peaks in the ENCODE database. Furthermore, they demonstrated that the regulatory circuit rewiring induced by rate-limiting step up/down-clamping in intellectual disability-Helsmoortel-Van der Aa syndrome-associated loci such as NFIB and ARID1B operates reversibly over the developmental time axis. This provides a backbone for a complete upgrade of existing epigenomic interpretation systems, which have been confined to single-timepoint snapshots, to a spatio-temporal, multi-dimensional tensor-based model for precision stratification of the developmental-disease continuum.
Prospects: Establishment of a Standard for Programmable Non-coding Genomics and Launch of Next-Generation IND Digital Governance
This study establishes a declarative turning point, resetting the governance of non-coding variant functional interpretation R&D from a static, post-hoc, symptomatic approach to a fully programmable infrastructure based on AI-driven, multi-dimensional tensors. The ChromBPNet model, FLARE score, and peak calls are publicly available as open-source resources with DOIs on Synapse, GitHub (kundajelab/neuro-variants), and Zenodo, enabling immediate integration into the non-coding target pipelines of global multinational pharmaceutical and biotechnology companies. By integrating long-read sequencing-based structural variant data from Illumina and PacBio, and downstream linking with 10x Genomics single-cell multi-omics (simultaneous scRNA-seq + scATAC-seq profiling), a computational moat is established that eliminates batch-to-batch chromatin accessibility prediction bias by mapping cell-type-specific genetic gradient correction factors to the FLARE tensor in real-time during high-throughput screening. If the FLARE-based variant priority score is incorporated as a companion diagnostic (CDx) standard in the submission of Investigational New Drug (IND) applications for next-generation non-coding target modalities, such as Roche and Genentech's autism gene therapy pipeline, Biogen and Ionis' antisense oligonucleotide (ASO) schizophrenia program, and Novartis' congenital heart disease gene editing Phase I/II candidates, the regulatory grade of non-coding variant pathogenicity evidence in the FDA and EMA regulatory review frameworks will be disruptively upgraded, functioning as a master asset that maximizes the probability of obtaining cGMP commercial approval.
Nature Genetics, Published online: 15 June 2026; doi:10.1038/s41588-026-02619-6This study contributes a resource of predicted effects of noncoding variants on chromatin accessibility and a method to identify noncoding variants with extreme regulatory effects, with application to disease variant discovery.
The 3 billion chromatin accessibility variant effect prediction tensors in this study go beyond theoretical exploration of non-coding regulatory mechanisms and are directly applied to global precision medicine pipelines and next-generation precision personalized neuro-cardiovascular bio-business lines.
First, in the clinical setting, the FLARE algorithm can immediately scan the pathogenicity landscape of de novo non-coding regulatory mutations in autism and schizophrenia patients, eliminating the temporal gap noise in genetic counseling and early diagnosis, and safeguarding against precision stratification of false positive variant classification errors.
At the same time, by linking to open-source databases that aggregate large-scale population omics matrices such as GTEx, ROSMAP, and gnomAD, a companion diagnostic (CDx) panel interface is realized that can virtually simulate confounding factors of inter-ethnic allele frequencies in clinical trial design and real-time reverse-calculate the effective transcription factor binding free energy changes of non-coding cis-regulatory targets.
Furthermore, when the FLARE extreme regulatory effect score is linked as a correction factor in the large-scale regulatory approval clinical trials of multinational companies' next-generation non-coding target gene therapies, it eliminates batch-to-batch chromatin accessibility prediction bias and functions as a backbone infrastructure that maximizes the probability of obtaining cGMP commercial approval from global regulatory agencies.