💻Code of Life

Identifying Causes of Rare Diseases in the 98% Noncoding Genome… A New Breakthrough Unlocked by Functional Genomics and AI

Nature Genetics·September 8, 2026AI Curation
Identifying Causes of Rare Diseases in the 98% Noncoding Genome… A New Breakthrough Unlocked by Functional Genomics and AI
AI Summary (Beta)Beta

Background

With advancements in human genome decoding technology, the cost of Whole Genome Sequencing (WGS) has significantly decreased, yet more than half of rare disease patients still undergo diagnostic odysseys without identifying a genetic cause. Until now, genetic research and clinical diagnosis have focused on the exome region, which directly encodes proteins. The exome accounts for only about 1.5% to 2% of the entire genome. It has been presumed that numerous disease-causing variants exist in the vast noncoding region, which makes up the remaining 98%, but cases identified as actual causes have been extremely rare.

The clinical interpretation of noncoding region variants has been suboptimal due to the absence of clear decoding rules. In protein-coding sequences, functional losses such as those caused by amino acid substitutions or premature termination can be predicted relatively easily based on the codon system. In contrast, it is difficult to intuitively gauge the impact of mutations on phenotypes even when regulatory sequences are altered. The 3D chromatin structure, in which regulatory elements such as enhancers or promoters control genes located hundreds of thousands of base pairs away, also increases the difficulty of analysis. In essence, there was a lack of verification tools to distinguish lethal, disease-causing variants from harmless benign variants within the vast noncoding sequences.

Key Findings

This review published in Nature Genetics summarizes the molecular mechanisms by which noncoding variants cause rare diseases and highlights the fundamental reasons for the low diagnostic discovery rate. The authors analyzed that noncoding regulatory abnormalities primarily induce quantitative changes in gene expression levels, making them act more tissue-specifically and subtly than complete protein loss. They pointed out that existing methods, which relied solely on evolutionary conservation analysis, struggled to capture regulatory variants that are activated only in specific cellular environments.

As a key solution to accelerate the discovery of noncoding variants, the researchers presented a combination of Massively Parallel Reporter Assay (MPRA), high-resolution chromatin structural analysis (Micro-C), and machine learning-based prediction algorithms. By utilizing Saturation Genome Editing (SGE), all single-nucleotide variants in noncoding sequences can be functionally evaluated simultaneously in a laboratory setting. The trend is for functional prediction accuracy to improve dramatically, as deep learning sequence models that directly interpret sequence information to quantify chromatin accessibility and transcriptional activity by cell type are combined. It is assessed that the screening efficiency for narrowing down actual disease-causing variants from tens of thousands of candidate variants has increased by dozens of times compared to the past.

Significance and Outlook

The precise identification of noncoding regulatory networks is expected to be a turning point in increasing the diagnosis rate of rare genetic diseases. This is because it can provide a clear molecular etiology for patients who missed the appropriate treatment window due to unknown causes. Accumulated noncoding variant data provides a blueprint for developing precision therapeutics that directly target noncoding regulatory sequences, such as antisense oligonucleotides (ASOs) or epigenome editing.

Challenges also remain clear. To bridge the gap between the predictions of computer algorithms and the actual phenotypes in patients, precise comparison with single-cell transcriptome maps across different organs and developmental stages must follow. Establishing a high-speed functional verification system using disease-model organoids and standardizing clinical interpretation guidelines for noncoding variants are also identified as tasks to be solved. The trend is for genome data interpretation capabilities to expand significantly beyond the boundaries of coding regions into noncoding regions.

Nature Genetics, Published online: 08 September 2026; doi:10.1038/s41588-026-02743-3This Review discusses how rare-disease-causing variants in the noncoding genome impact gene regulation, why these examples are so few and how new approaches could accelerate discovery of noncoding variants affecting human health.

💬Why it matters:

The approach presented in this study can significantly improve the diagnosis rate in undiagnosed rare disease clinics where causes remain unidentified even after next-generation sequencing. In clinical settings, by linking a patient's WGS data with noncoding functional prediction AI models and single-cell multi-omics data, a system will be established to prioritize transcriptional regulatory site variants directly linked to the disease, shortening interpretation time from months to days.

From the perspective of the drug development industry, this provides an opportunity to identify a large number of new regulatory factor targets for intractable diseases where drug development was impossible due to the lack of protein targets. It is expected that the design of targeted RNA therapeutics, which inhibit disease-causing enhancers or selectively restore silenced gene expression, will become a reality.

💬 Comments

0 comments
Please log in to comment
Loading...

BioPlayground

Reading, linking, and lawful quotation stay open; high-speed bulk collection and unauthorized redistribution do not.

Unless stated otherwise, content rights belong to BioPlayground or the relevant rights holder.