πŸ’»Code of Life

AlphaGenome, which predicts the impact of 9 billion single-nucleotide variants in the human genome, paves the way for elucidating disease mechanisms in non-coding regulatory regions.

NatureΒ·September 8, 2026AI Curation
AlphaGenome, which predicts the impact of 9 billion single-nucleotide variants in the human genome, paves the way for elucidating disease mechanisms in non-coding regulatory regions.
✨AI Summary (Beta)Beta

Background

Despite rapid advancements in genome analysis technology, human genomics has faced a long-standing challenge. Although it has been over 20 years since the Human Genome Project decoded 3 billion base pairs, the regions actually encoding proteins account for only about 1.5% to 2% of the entire sequence. The vast non-coding regions, comprising the remaining 98%, act as switches to finely regulate gene expression, but their functional map remains uncharted territory.

Previous studies, including Genome-Wide Association Studies (GWAS), have succeeded in cataloging numerous genetic variants associated with complex diseases. However, since more than 90% of discovered variants are concentrated in non-coding regulatory regions, it has not been easy to prove the causal relationship between specific base changes and disease induction at the molecular level. High-throughput screening experiments that introduce mutations into cell lines also face limitations, such as enormous costs and the difficulty of fully reproducing the complex chromatin structure in vivo. While AlphaFold and AlphaMissense have advanced missense variant analysis in the field of protein structure prediction, a computational model capable of elucidating the vast regulatory network of non-coding regions has long been missing.

Key Discovery

To fill this gap, researchers at Google DeepMind introduced 'AlphaGenome,' an AI-based sequence model covering the entire human genome. The researchers succeeded in predicting the biological impact of a total of 9 billion single nucleotide variants (SNVs) by exhaustively analyzing the three possible substitution cases for each of the approximately 3 billion base pairs constituting the human genome. This means they have calculated, at single-nucleotide resolution, how fluctuations in a single base affect gene expression levels, chromatin accessibility, transcription factor binding affinity, and splicing patterns.

The model architecture is designed to simultaneously compute long genomic contexts spanning hundreds of thousands of base pairs to capture interactions between distant regulatory elements. The researchers quantified the ripple effects on distant regulatory elements, such as promoters and enhancers, by training on whole-genome data and large-scale functional genomics experimental results. Benchmark evaluation results showed that the accuracy of pathogenicity prediction for non-coding variants improved by more than two orders of magnitude compared to existing statistical genetics tools and machine learning models. While existing models often missed new regulatory variants by relying on conserved sequence ratios, AlphaGenome distinguishes itself by reading the three-dimensional binding changes of regulatory factors directly from the sequence information itself.

Significance and Outlook

This achievement marks a new turning point for research tracing the pathogenesis of rare intractable genetic diseases, cancer, and complex diseases such as diabetes caused by non-coding variants. By resolving the bottleneck in variant interpretation that was difficult to verify individually in the laboratory, a way has opened to pre-identify the pathogenicity of numerous non-coding region variants that remained as Variants of Uncertain Significance (VUS) in clinical genomic sequencing. For researchers handling large-scale genomic cohorts, it provides a powerful selection criterion to narrow down the pool of analysis candidates.

However, clear practical limitations of computational predictions also exist. Since the regulatory variant effects calculated by AI can manifest differently depending on the actual in vivo tissue environment, developmental stage, or epigenetic state, precise experimental verification must necessarily follow. Follow-up work is also required to augment genomic data from diverse racial groups to avoid bias toward specific population data. It is expected that, when combined with single-cell analysis technology, computational biology will establish itself as a key driver for accelerating drug target discovery.

Nature, Published online: 08 September 2026; doi:10.1038/d41586-026-02835-4AI model AlphaGenome forecasts the consequences of altering every single DNA letter in the human genome.

πŸ’¬Why it matters:

The functional map of 9 billion variants presented by AlphaGenome advances the timeline for drug development and precision medicine. In clinical settings, it will enable rapid molecular diagnosis for patients carrying non-coding variants that were previously classified as having unknown function after Next-Generation Sequencing (NGS) testing. A representative clinical scenario is establishing a treatment strategy on the day of testing by predicting that a non-coding promoter mutation in a patient suspected of having a rare genetic disease suppresses specific gene expression.

In the biopharmaceutical industry, it reduces trial and error in the stages of target discovery and gene therapy design. When designing antisense oligonucleotides (ASOs) or CRISPR-based corrective therapies targeting specific non-coding regions that cause disease, it helps prevent off-target mutations in advance. It is expected to provide practical benefits by significantly reducing the repetitive cell experiment processes in the candidate substance derivation stage, thereby saving drug development costs.

πŸ’¬ Comments

0 comments
Please log in to comment
Loading...

BioPlayground

Reading, linking, and lawful quotation stay open; high-speed bulk collection and unauthorized redistribution do not.

Unless stated otherwise, content rights belong to BioPlayground or the relevant rights holder.