💻Code of Life

scE2G: A Single-Cell Data-Driven Model for Mapping Enhancer-Target Gene Regulatory Landscapes

Nature Genetics·August 4, 2026AI Curation
scE2G: A Single-Cell Data-Driven Model for Mapping Enhancer-Target Gene Regulatory Landscapes
AI Summary (Beta)Beta

Background

The human genome contains millions of enhancers that regulate gene expression in specific cell types. A significant proportion of disease-associated variants reside in enhancers, rather than protein-coding regions; however, identifying the target genes and cell types regulated by these enhancers is challenging. This is because enhancers can be located tens or hundreds of kilobases away from their target genes, and they may act by regulating genes other than the closest one.

Existing Activity-by-Contact (ABC) and ENCODE-rE2G models predict regulatory relationships by leveraging chromatin activity and three-dimensional contact information. However, bulk tissue analysis mixes signals from multiple cells, making it difficult to distinguish regulatory circuits in rare or transient cell states. Single-cell chromatin accessibility analysis (scATAC-seq) and single-cell RNA sequencing (scRNA-seq) have emerged as alternatives, but unsupervised learning approaches that rely on simple correlations between accessibility and expression levels lack objective standards for evaluating accuracy.

To address these limitations, the researchers developed scE2G, a family of single-cell enhancer-gene prediction models. The results were published on August 3, 2026, in Nature Genetics.

Key Findings

scE2G consists of scE2G-ATAC, which uses only scATAC-seq data, and scE2G-Multiome, which uses data from scRNA-seq and ATAC-seq performed on the same cells. Both models are supervised learning-based logistic regression classifiers. The researchers trained the models using 13,420 enhancer-gene candidate pairs identified through CRISPR screening in K562 erythroleukemia cells. Among these, 466 represented positive associations, where gene expression significantly decreased after enhancer inhibition, and 9,876 represented negative associations.

scE2G-ATAC calculates six features: ABC score, chromatin accessibility of enhancers and promoters, distance and gene density between the two regions, and promoter type. scE2G-Multiome adds the Kendall correlation coefficient of enhancer accessibility and gene expression within individual cells, and uses this in combination with the ABC score to create an ARC-E2G feature. The possibility of overfitting was also reduced through cross-validation, where one chromosome was excluded at a time.

The researchers compared the performance of the two scE2G models with ten existing single-cell models using three independent datasets: CRISPR perturbation, fine-mapped expression quantitative trait loci (eQTL), and genome-wide association study (GWAS) variant-gene associations. Both scE2G models achieved the highest precision and area under the precision-recall curve, not only in K562 cells but also in 4,175 additional CRISPR datasets from five other cell types. In eQTL evaluation, scE2G-Multiome had a recall of 13.9% and a 14.9-fold enrichment of variants. The next best performing non-ABC-based model, SCENIC+, had a recall of 1.4% and a 10.3-fold enrichment.

By constructing regulatory maps for 45 cell types from peripheral blood, bone marrow mononuclear cells, and islets, the researchers predicted an average of 48,758 connections per cell type in the 39 types that met the quality criteria. On average, 5.6 enhancers were associated with a single expressed gene, and the average distance between enhancers and their target promoters was 86.3 kilobases.

Significance and Implications

scE2G is significant because it can narrow down the cell types and target genes affected by disease variants, even in mixed-cell populations. The researchers linked 1,450 variants in 1,892 non-coding regions to 1,351 genes, and identified 458 regions that targeted genes further away from the nearest transcription start site.

For example, rs7696969, a variant associated with lymphocyte count, is located within INPP4B. scE2G suggests that this variant may regulate INPP4B and IL15 in natural killer cells and T cells. The distance from the variant to the promoters of the two genes is 441 kilobases and 769 kilobases, respectively. The posterior inclusion probability that the variant is an eQTL for INPP4B was also 66.6%, and the independent gene prioritization analysis, PoPS, also selected the two genes as top candidates. However, this is not a definitive confirmation of functional causality, but rather a hypothesis that requires further experimental validation.

For stable application of the model, it is recommended to have at least 100 cells per cell type, a minimum of 2 million ATAC fragments, and 1 million RNA unique molecular identifiers (UMIs). A limitation is that only 466 positive associations were used for training, and most of these were from K562 cells. Topological regulatory elements other than enhancers, such as CTCF binding sites, are not included as prediction targets. Accumulating large-scale CRISPR validation data from multiple primary cells and incorporating Hi-C or H3K27ac information will be necessary to expand the model into a clinically reliable regulatory map.

Nature Genetics, Published online: 03 August 2026; doi:10.1038/s41588-026-02695-8scE2G is a family of models that predict enhancer–gene regulatory interactions from single-cell datasets and enable mapping of these interactions across diverse cell types and tissues.

💬Why it matters:

Applying scE2G to single-cell multi-omics data from patient tissues can narrow down the cell types and target genes regulated by GWAS variants. For example, if a non-coding variant associated with an autoimmune disease overlaps with an enhancer in a specific T cell subtype, the connected target genes can be validated by CRISPR inhibition or by examining changes in expression in organoids or primary cells.

For pharmaceutical companies, scE2G has the potential to be a target discovery and patient stratification tool. By linking disease variants, acting cell types, and target genes, it can strengthen the genetic basis of drug targets and facilitate the design of appropriate biomarkers. However, the connections made by scE2G are probabilistic predictions, so they should be validated with CRISPR perturbation, protein expression, and three-dimensional chromatin contact data before being used for drug development decisions.

💬 Comments

0 comments
Please log in to comment
Loading...