πŸ’»Code of Life

Significant Improvement in Polygenic Prediction Accuracy for Complex Diseases in Underrepresented Groups through Integration of Multi-ancestry Biobank Data

Nature GeneticsΒ·September 14, 2026AI Curation
Significant Improvement in Polygenic Prediction Accuracy for Complex Diseases in Underrepresented Groups through Integration of Multi-ancestry Biobank Data
✨AI Summary (Beta)Beta

Background

The polygenic risk score (PRS), which calculates the future risk of disease onset by aggregating individual DNA variants, has been recognized as a key tool in precision medicine. It is a method that statistically sums the effects of hundreds of thousands to millions of single nucleotide polymorphisms (SNPs) to assess individual susceptibility. However, existing genome dynamics research was trapped by a structural limitation of severe sample imbalance. This is because more than 80% of participants used in Genome-Wide Association Studies (GWAS) were biased toward European ancestry.

This distortion in genomic data has resulted in a sharp decline in prediction model performance for non-European populations. When applying existing PRS algorithms to African or Hispanic groups, prediction accuracy can drop to less than half that of European groups. This is because linkage disequilibrium (LD) patterns and allele frequencies vary across populations when assessing disease susceptibility. Consequently, continuous criticism has been raised that advancements in genomic research are deepening healthcare inequalities by benefiting only specific racial groups.

Key Findings

Researchers adopted a large-scale multi-ancestry meta-analysis strategy by combining data from the National Institutes of Health (NIH) All of Us Research Program (AoU) and the UK Biobank (UKB). The AoU is a cohort that has secured a proportion of underrepresented minorities of over 50%, groups historically excluded from medical research. The researchers reconstructed models for complex traits and major chronic diseases by integrating Whole Genome Sequencing (WGS) and genotyping data from hundreds of thousands of individuals registered in both large biobanks.

By applying a multi-ancestry Bayesian regression technique that simultaneously learns the genetic architectures of diverse populations, the predictive performance of the models in non-European groups improved significantly. In major chronic diseases such as type 2 diabetes, coronary artery disease, and hypertension, the prediction accuracy (AUC) for African American and admixed populations significantly improved compared to existing single-ancestry-based models. Notably, in the African-ancestry cohort, the prediction performance for continuous traits such as systolic blood pressure and body mass index (BMI) improved by up to 40%. This was achieved through an algorithm that selects causal genetic variants acting commonly across multi-ancestry groups and finely corrects for the unique LD structures of each group. This proves that securing genetic diversity, rather than just the absolute expansion of sample size, is the decisive factor in determining the performance of prediction algorithms.

Significance and Outlook

This study demonstrates that expanding data scale and collecting multi-ancestry samples are practical solutions to bridging the ancestry gap in clinical genomics. A turning point has been reached in expanding the scope of precision preventive medicine, which was previously limited to specific racial groups. As multi-ancestry algorithms become widespread, they will further enhance the accuracy of high-risk group identification by reducing false-negative errors occurring in early screening systems. In the global pharmaceutical industry, a path has opened to pre-verify the genetic validity of target discovery across diverse ethnic groups.

However, a cautious approach is required for clinical application. Even though non-European samples have increased, data density is still insufficient to fully reflect the high genetic heterogeneity within African-ancestry groups. Additionally, the interaction of non-genetic factors, such as socioeconomic factors or living environments, with complex diseases remains a task to be addressed. In the future, international data linkage with local cohorts from Asia, South America, and the African continent must follow to achieve true global genomic equity.

Nature Genetics, Published online: 14 September 2026; doi:10.1038/s41588-026-02734-4Leveraging multiancestry and multi-biobank data from the All of Us Research Program and the UK Biobank improves polygenic prediction models for complex traits and diseases, especially for under-represented populations.

πŸ’¬Why it matters:

The multi-ancestry polygenic prediction models derived from this study can be immediately utilized in early screening programs for chronic diseases in primary healthcare institutions. This provides a pathway to early detection of patients with multi-ancestry and multicultural backgrounds who were previously missed by the conventional single-lineage model when applied to high-risk groups, allowing for the initiation of lifestyle improvements or preventive medication. There are also economic benefits, such as preventing waste in national healthcare budgets by accurately classifying high-risk groups for cardiovascular disease or metabolic syndrome.

In the field of new drug development, it enhances the stratification of clinical trial participants and the efficiency of biomarker discovery. In global Phase 3 clinical trials, it reduces uncertainty in predicting drug response rates by minimizing genetic risk bias when selecting patient cohorts of diverse ethnicities. Genomic analysis companies can build commercial pipelines that provide genetic testing reports with uniform reliability to consumers of diverse ancestries.

πŸ’¬ Comments

0 comments
Please log in to comment
Loading...

BioPlayground

Reading, linking, and lawful quotation stay open; high-speed bulk collection and unauthorized redistribution do not.

Unless stated otherwise, content rights belong to BioPlayground or the relevant rights holder.