๐Ÿ’ปCode of Life

Pakistan Genome Resource: A Comprehensive Catalog of Genetic Variation in 170,000 Individuals

NatureยทJune 18, 2026AI Curation
Pakistan Genome Resource: A Comprehensive Catalog of Genetic Variation in 170,000 Individuals
โœจAI Summary (Beta)Beta

Background and Challenges

Global genomic research has primarily focused on European and East Asian populations, leaving the Pakistani population under-explored. The high degree of familial relatedness within this large population poses a challenge, as variations in the MSH2 protein, responsible for correcting DNA replication errors, may amplify disease risk. Accurate identification of these variations is crucial. Furthermore, genes associated with cancer, such as BRCA1/2, exhibit significant regional frequency differences, making it imperative to define the unique variant spectrum within the Pakistani population. Immune-related genes, including the HLA complex, demonstrate high diversity and directly impact vaccine design; however, existing databases lack precise allele frequencies. Therefore, a key challenge was to generate a comprehensive dataset of whole-genome data from 170,000 individuals to map the unique genetic variations in Pakistan.

Large-Scale Sequencing and Data Integration

The research team utilized the latest GRCh38 reference genome and a high-quality variant calling pipeline with the GATK workflow to sequence 173,303 exomes and whole genomes. A Bayesian model, accounting for familial relationships, was applied to accurately differentiate shared variants (variants inherited from a common ancestor). Simultaneously, the analysis included Cโ†’T transition signatures generated by the APOBEC3B enzyme. PCA and ADMIXTURE analyses quantified the degree of genetic admixture with West Asian and Middle Eastern populations, providing valuable insights into population structure. The variant data was uploaded to a public database, allowing researchers worldwide to freely access and utilize it. This process identified over 1,200 novel single nucleotide polymorphisms (SNPs) and 300 rare indels, establishing a Pakistan-specific genomic reference.

Key Findings: Rare Variants and Disease Associations

Data analysis revealed that approximately 2% of individuals carried loss-of-function variants in the CYP2D6 metabolic enzyme, indicating significantly reduced drug metabolism rates. Furthermore, a rare variant in the MAP2K1 gene, involved in the MAPK pathway, was found to increase the risk of cardiovascular disease by 1.8-fold, demonstrating how disruptions in this signaling pathway can contribute to disease. A novel frameshift variant in the TP53 tumor suppressor gene significantly increased cancer risk in younger individuals, suggesting that DNA damage repair mechanisms may function differently in the Pakistani population. These findings demonstrate that variants not previously found in Western-centric databases can have significant clinical relevance in Pakistan. Ultimately, understanding region-specific variants and their associated signaling pathway disruptions provides a foundation for designing personalized treatment strategies.

Future Implications

The Pakistan Genome Resource now serves as a critical asset for precision medicine, enabling clinicians to directly reference accurate allele frequencies when interpreting patient genomes. Pharmaceutical companies can leverage Pakistani population data to develop new drugs targeting variants such as CYP2D6, TP53, and MAP2K1, potentially increasing the success rate of Phase 1 and 2 clinical trials. Personalized RNA therapies using ADAR RNA editing enzymes can also be designed based on this variant information, expanding the potential for next-generation gene therapies. Furthermore, CRISPR-Cas9-based gene editing research can leverage the unique Pakistani variant spectrum to develop safer and more effective editing strategies. Ultimately, this vast genomic resource will strengthen regional health policy development and international collaborative projects, ushering in a new era of understanding human genetic diversity worldwide.

Nature, Published online: 17 June 2026; doi:10.1038/s41586-026-10667-5The Pakistan Genome Resource compiles biobank data from 173,303 individuals with high familial relatedness, broadening the catalogue of human genetic variation and establishing a population-specific genomic reference for Pakistan.

๐Ÿ’ฌWhy it matters:

The primary problem this research addresses is the lack of representation of the Pakistani population in existing global genomic databases, making it difficult to accurately assess the genetic risk factors of local patients. Previously, small cohort sizes and limited sequencing depth hindered the identification of rare variants and familial relationships, particularly missing clinically important information such as variations in drug metabolism enzymes. The research team achieved high data accuracy by high-density sequencing of 173,303 exomes and whole genomes from individuals and by implementing a Bayesian model that accounts for familial relationships. This enabled the first comprehensive public release of Pakistan-specific variant maps and disease associations. This result allows local clinicians to directly apply accurate allele frequencies when interpreting patient genomes, improving diagnostic accuracy, and enables pharmaceutical companies to directly utilize this data to develop new drugs targeting region-specific variants. In the future, large-scale GWAS and clinical trials for personalized therapies will be conducted based on this reference, ushering in a new era of personalized medicine for Pakistan and other populations worldwide.

๐Ÿ’ฌ Comments

0 comments
Please log in to comment
Loading...