Pan-genome and Genetic Variation Map for Flowering and Dwarfism Traits Completed via Genomic Analysis of 2,320 Cultivated Peanuts

Background
Cultivated peanut (Arachis hypogaea) is a key legume crop supporting global edible oil and protein supplies. It holds high economic value in tropical and subtropical agriculture due to its ability to grow in dry and poor soil environments. However, it has a vulnerability of narrow genetic diversity due to its formation through natural hybridization and polyploidization among wild species. Relying solely on a single reference genome presented clear limitations in interpreting the entire complex allotetraploid structure.
Existing crop research has assembled genomic sequences based on a single standard variety. This approach results in missing data issues, failing to capture genes that are lost or newly emerged between varieties, as well as large-scale structural variations (SV). In particular, despite dramatic differences among peanut varieties in key agronomic traits such as flowering cycle, pod development, and plant architecture, there was a lack of comprehensive genetic information to explain these at the molecular level. This was the background in which a genomic blueprint was needed to precisely track where useful genes are located and which variations induce phenotypes.
Key Findings
An international research team constructed an integrated pan-genome by performing de novo assembly of the genomes of 10 representative cultivated peanut varieties. Furthermore, they conducted whole-genome resequencing (WGRS) on 2,320 germplasm accessions collected worldwide to comprehensively survey vast genetic polymorphism.
As a result of the analysis, numerous SVs, including large-scale insertions, deletions, and chromosomal inversions that were missing from the single reference genome, were revealed. Based on the pan-genome data, the researchers conducted a genome-wide association study (GWAS) to identify key trait loci directly linked to farm profitability. Notably, they identified flowering time-regulating gene loci that induce early flowering and uniform fruit set, as well as key dwarfism-related genes that inhibit stem elongation to enhance lodging resistance.
Compared to previous analyses that were limited to the single nucleotide polymorphism (SNP) level, this study is notable for proving the mechanism by which large-scale structural variations directly regulate changes in gene expression levels. In strains exhibiting the dwarf phenotype, structural variations within specific regulatory regions were observed to alter the transcriptional efficiency of growth hormone signaling genes. The spectrum of 2,320 germplasm accessions clearly illustrates the flow of gene clusters that were either conserved or lost during variety differentiation.
Significance and Outlook
This pan-genome map will serve as a molecular accelerator for customized crop improvement in response to climate change. In modern agricultural environments where drought and high-temperature stress are becoming permanent, technologies to freely regulate flowering periods and refine plant architecture to be more compact are essential conditions for securing yields. Converting the identified dwarf loci into molecular markers can shorten the selection period for dwarf varieties by several years.
Improvements in farming efficiency are also expected. Varieties with shorter, denser plant architecture are advantageous for mechanized harvesting and significantly reduce pest management costs. Additionally, by controlling flowering concentration to resolve the uneven pod development that previously occurred at different maturity stages, it becomes possible to harvest high-quality seeds with high marketability.
However, challenges remain before this vast genomic data can lead directly to the commercialization of varieties. Due to the genetic redundancy characteristic of tetraploid crops, the possibility of unintended metabolic interference during the editing of specific alleles cannot be ruled out. Subsequent field trials must be conducted to empirically verify the phenotypic expression of the identified gene loci across various cultivation environments.
Nature Genetics, Published online: 15 September 2026; doi:10.1038/s41588-026-02765-xPan-genome analyses of cultivated peanut using 10 newly assembled genomes, along with whole-genome resequencing of 2,320 germplasm accessions, highlight structural variations and loci associated with agronomic traits such as flowering and dwarfism.
This study provides a genetic resource navigation tool that can be directly applied to the field of molecular crop breeding. Breeders can use the 2,320-variety resequencing database to perform virtual screening of whether beneficial alleles related to flowering time and plant architecture are present when selecting breeding parents. While traditional field breeding required monitoring growth for several months after sowing to confirm flowering characteristics and dwarfism, it is now possible to identify superior individuals early through DNA chip testing at the seedling stage, increasing the rate of generational advancement by more than three times. Furthermore, it provides direct genetic coordinates for developing peanut lines optimized for smart farms—suitable for high-density cultivation and mechanized harvesting—and for breeding new varieties that maximize yield per unit area.