๐Ÿ’ปCode of Life

Enhanced Genome-Wide Association Study for Admixed Populations: A New Algorithm Improves Detection of Ancestry-Specific Variants in Related Individuals

Nature GeneticsยทJuly 21, 2026AI Curation
Enhanced Genome-Wide Association Study for Admixed Populations: A New Algorithm Improves Detection of Ancestry-Specific Variants in Related Individuals
โœจAI Summary (Beta)Beta

Background

Genome-Wide Association Studies (GWAS) have become a cornerstone for identifying genetic variants associated with diseases, particularly with the rise of large biobanks. However, most existing GWAS studies have been limited by their focus on European populations. Analyzing admixed populations, such as those of African or Hispanic descent, presents challenges due to population structure bias, which can lead to a high number of false positive signals.

To address these limitations, the statistical algorithm 'Tractor,' which incorporates Local Ancestry Inference, was introduced in 2021. Tractor successfully classifies ancestry at specific chromosomal regions, enabling a more precise estimation of the impact of ancestry-specific variants.

However, the original Tractor model assumes that there is no relatedness between the individuals being analyzed. Modern large biobanks, such as the Mexico City Prospective Study cohort and the UK Biobank, often contain samples with familial relationships. Excluding related individuals arbitrarily leads to significant loss of genomic data and reduced detection power, while including them without adjustment leads to statistical errors. This has been an ongoing dilemma.

Key Findings

A team of American researchers has developed a new algorithm, 'Tractor-Mix,' that performs precise analysis even in populations with complex familial relationships and mixed ancestry. The research findings were published in the latest issue of the international journal Nature Genetics.

Tractor-Mix is a framework that integrates a Linear Mixed Model (LMM) with local ancestry inference. It incorporates a Genetic Relationship Matrix (GRM), which quantifies the degree of genetic relatedness between samples, into the mathematical model. This allows for the accurate calculation and correction of statistical biases arising from family relationships. Simultaneously, the algorithm tracks the ancestry of each chromosomal segment in parallel.

The research team validated their hypothesis using data from the Mexico City cohort and the UK Biobank. The results showed that Tractor-Mix effectively controlled the false positive rate and significantly improved the detection power for variants enriched in specific ancestries, without excluding related samples. In particular, it accurately quantified the effect size of rare disease-related variants in non-European admixed populations without loss of genetic relative data.

Significance and Future Directions

This research establishes a foundation for expanding genomic research beyond its Eurocentric focus to include more diverse populations. It provides a statistical framework that allows for the full utilization of clinical genomic data from cohorts with complex familial relationships, such as those found in South America and Africa.

The researchers plan to conduct follow-up work to improve the computational speed of the algorithm. Due to the simultaneous processing of local ancestry inference and large-scale genetic relationship matrix calculations, the computational burden increases when processing datasets of hundreds of thousands of individuals. Improving computational efficiency through the implementation of a cloud-based distributed computing system is a key priority.

Nature Genetics, Published online: 20 July 2026; doi:10.1038/s41588-026-02689-6. Tractor-Mix is a GWAS method for admixed cohorts with high relatedness, improving ancestry-specific effect estimation, controlling false positives and boosting power to detect ancestry-enriched loci across diverse populations.

๐Ÿ’ฌWhy it matters:

The introduction of Tractor-Mix is expected to bring about direct changes in the development of personalized medicine and targeted therapies for diverse populations. Until now, pharmaceutical companies and research institutions have relied on data from European populations when identifying genetic variants associated with drug response and disease risk. This has often led to inaccuracies in target identification and the calculation of Polygenic Risk Scores (PRS) for diverse patient populations.

In the future, it will be possible to analyze global population databases with diverse genetic backgrounds and familial relationships without loss of information. This will accelerate the discovery of disease-causing variants specific to certain ethnicities or admixed populations, and provide practical benefits for the selection of participants in global clinical trials and the development of personalized treatment strategies.

๐Ÿ’ฌ Comments

0 comments
Please log in to comment
Loading...