Genomic AI GPN-Star, Directly Reflecting Evolutionary Phylogeny, Sets New Standard for Whole-Genome Variant Effect Prediction

Background
Out of the 3 billion base pairs in the human genome, only about 1.5% directly encode proteins. The remaining 98% non-coding regions act as critical switches regulating gene expression, yet it has been virtually impossible to experimentally verify the impact of sequence variations on diseases. While attempts have been made to measure evolutionary conservation using computational algorithms, they have revealed limitations in failing to adequately reflect the surrounding nucleotide sequence context. Recently, Genomic Language Models (gLMs), which learn large-scale sequences like natural language, emerged as a promising alternative. However, most existing gLM models either analyze single-species sequences or fail to consider evolutionary distance when aligning multi-species sequences. Base pair matches between closely related species, such as humans and chimpanzees, acted as redundant data causing bias, while variations in distantly related species were often underestimated. This created a growing demand to directly integrate biological evolutionary phylogenies into learning algorithms.
Key Findings
A research team led by Professor Yun S. Lee from the Department of Computer Science at UC Berkeley Professor Song’s research team developed GPN-Star, a genomic model that combines phylogenetic tree distances with Whole-Genome Alignment (WGA) data, and published it in the international journal Nature. The core driver is a phylogeny-aware architecture that directly maps evolutionary divergence times between species into cross-attention operations, moving beyond previous methods that relied on simple list-based multiple alignments. By moving away from the practice of focusing only on the reference human sequence and instead applying random masking across the aligned sequences of multiple species, the density of learning signals was significantly increased. Analysis of the evolutionary lineages of 241 mammal species and 100 vertebrate species led to the precise quantification of functional constraints acting on all alleles within the genome. Performance metrics show a clear gap. In ClinVar variant identification experiments, GPN-Star significantly outperformed the existing conservation metric PhyloP and the ensemble score CADD. In the ClinVar clinical genetic variant database variant identification experiment, GPN-Star significantly outperformed the existing conservation metric PhyloP and the ensemble score CADD. It is evaluated that the accuracy of finding causal variants in genome-wide association studies (GWAS) in non-coding regulatory regions also surpassed that of previous state-of-the-art models. This demonstrates the flexibility to transcend species boundaries, extending to the analysis of five major model genomes including mice, chickens, fruit flies, Caenorhabditis elegans, and Arabidopsis.
Significance and Outlook
The emergence of GPN-Star provides a powerful guide for drug discovery researchers tracking the disease associations of non-coding genetic variants. Even when performing Whole-Genome Sequencing (WGS) for patients with rare intractable diseases where causes were not found in exome sequencing, the task of sifting through millions of non-coding variants for the true cause has been likened to finding a needle in a haystack. Implementing GPN-Star, which precisely calculates evolutionary pressure, is expected to increase the efficiency of compressing potential candidates several times over. The researchers' decision to release the calculated scores as UCSC Genome Browser tracks and to fully open the model weights is also a factor injecting vitality into the ecosystem. This opens a way for small and medium-sized biotechs without high-performance computing equipment to immediately determine variant risk with just a few clicks. However, practical challenges remain. Prediction reliability tends to fluctuate in repetitive sequences or structural variation regions where inter-species base alignment is difficult. Furthermore, because it relies on base information preserved by evolutionary history, it has a clear limitation in not being able to capture real-time epigenetic expression dynamics occurring in specific cells or tissues. This is why integration with single-cell chromatin accessibility data has emerged as a key next-generation challenge.
Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-11005-5GPN-Star, a genomic language model with a phylogeny-aware architecture for whole-genome alignment data, is shown to be a scalable and flexible tool for genetic variant effect prediction across species.
In clinical genomic diagnostics, it is expected to be immediately implemented in interpretation pipelines to confirm non-coding pathogenic variants in patients with rare genetic diseases of unknown etiology. When performing whole-genome analysis on patients who failed diagnosis with standard exome testing, applying GPN-Star scores allows for the rapid filtering of pathogenic candidates hidden deep within promoters or introns from tens of thousands of Variants of Uncertain Significance (VUS). In the pharmaceutical and biotech industries, it will serve as a decisive clue for developing next-generation gene therapies targeting non-coding regions. If key enhancers or silencers regulating disease gene expression are accurately identified, a framework for the much more precise design of antisense oligonucleotide (ASO) or gene-editing-based transcriptional regulation therapies will be established. The agricultural and livestock biotech sectors will also benefit practically by shortening breeding cycles through the rapid selection of beneficial variants that determine crop disease resistance or livestock productivity.