A Paradigm Shift in Multi-ethnic Omics: Mega Exome-Wide Association Analysis of 1.1 Million Individuals Reveals Rare Coding Variants in 209 Genes and the Epidemiology of Next-Generation Lipid Metabolism Targets

-
Limitations of European‑biased genomic research and the blind spot for rare coding variants in minority populations. Blood lipid metabolism, a core driver of cardiovascular disease (CVD), is a complex phenotype shaped by numerous genetic weights and environmental factors. However, previous genome‑wide association studies (GWAS) have been overwhelmingly skewed toward European‑ancestry cohorts. Consequently, rare coding variants and loss‑of‑function (LoF) mutations that exert strong phenotypic penetrance in African, Hispanic, Asian, and other under‑represented groups have been systematically missed due to insufficient statistical power. Diagnostic markers lacking genetic diversity create a fatal blind spot that undermines the reliability of polygenic risk scores (PRS) across global populations.
-
Mega‑cohort exome binding of 1.1 million individuals: an integrated multi‑ethnic association framework (MVP‑UKB‑All of Us). In a study published on 25 May 2026 in Nature Genetics, we assembled an unprecedented multi‑ethnic cohort of 1,158,017 participants to address extreme racial bias in genomics. We integrated whole‑genome data from three leading biobanks—the U.S. Department of Veterans Affairs Million Veteran Program (MVP), the UK Biobank, and the NIH All of Us Research Program. By filtering non‑coding noise and applying a high‑resolution exome‑wide association study (EWAS) pipeline, we pinpointed protein‑altering regions with allele‑specific effects and mathematically combined effect sizes that differ across ancestries.
-
Identification of novel variants in 209 genes and empirical validation of ~30 next‑generation candidate drug targets. Multivariate regression of the mega‑exome dataset uncovered numerous rare coding‑variant hotspots across 209 core genes that programmatically regulate blood lipid traits. In addition to established lipid‑lowering targets such as APOB and PCSK9, we introduced more than 30 novel candidate drug targets that fundamentally remodel the kinetics of lipid metabolic pathways, presented for the first time in silico. These loci, validated across diverse ancestral backbones, provide robust biophysical evidence for the genetic causality of hyperlipidemia and atherosclerosis that disproportionately affect specific populations.
-
Establishment of a global multi‑ethnic precision‑medicine platform and acceleration of drug‑development timelines. This population‑genomics white paper delivers a disruptive impact on next‑generation biopharmaceutical R&D and personalized digital‑health business models. By converting cardiovascular‑prevention guidelines from a fixed, European‑centric polygenic score to a globally calibrated risk matrix that mathematically adjusts allele frequencies by ancestry, we enable a paradigm shift. The ~30 newly discovered drug targets constitute exclusive assets that can eliminate off‑target toxicity noise in nucleic‑acid therapeutics such as ASOs and siRNAs. Moreover, incorporating ancestry‑specific drug‑response weights into multinational trial designs will maximize translational success rates and dramatically shorten regulatory approval timelines.
Nature Genetics, Published online: 25 May 2026. DOI: 10.1038/s41588-026-02613-y
Summary: Bypassing the ancestry-biased bottlenecks of historical genome-wide studies, this milestone exome-wide association study (EWAS) synthesizes deep genomic profiles from 1,158,017 individuals across the Million Veteran Program, UK Biobank, and All of Us cohorts. By deploying a high-throughput multi-ancestry analytical matrix, the framework successfully uncovers highly penetrant, rare coding variants across 209 genes fundamentally linked to systemic blood lipid traits. Beyond validating classical benchmarks like APOB and PCSK9, the multi-omic metadata isolates over 30 novel candidate drug targets showing robust allelic effect sizes across non-European sub-populations. This diverse population genetics repository eliminates confounding missing heritability, providing a programmable, low-noise computational baseline for universal cardiovascular multi-genic risk stratification and targeted nucleotide therapeutic design.
This study addresses the two greatest challenges in whole‑genome genetics—missing heritability due to lack of ethnic diversity and the statistical‑noise barrier that obscures the expression thresholds of rare genotypes—by mathematically quantifying them through a 1.1‑million‑person multi‑cohort exome‑binding approach. The resulting dataset constitutes a top‑tier R&D asset (the “code of life”) that includes ancestry‑specific variant‑weight tensors and lipid‑phenotype variance curves. Consequently, it serves as a powerful proprietary reference for future AI‑driven, multi‑ethnic, multi‑gene polygenic risk‑score algorithms and enables global biopharma pipelines to achieve world‑leading molecular‑design resolution.