๐Ÿ’ปCode of Life

AI Model for Predicting Gene Expression from Sequence: Generalizability and Interpretability as the Next Milestones

Nature GeneticsยทJuly 28, 2026AI Curation
AI Model for Predicting Gene Expression from Sequence: Generalizability and Interpretability as the Next Milestones
โœจAI Summary (Beta)Beta

Background

Following the decoding of the human genome, a major challenge in modern genetics has been to elucidate the specific gene regulatory activities of the 3 billion base pairs of DNA. Mainstream genetic research has focused on Genome-Wide Association Studies (GWAS) to identify correlations between specific variants and diseases, making it difficult to directly explain the molecular biological mechanisms of non-coding region variants. Recent advances in functional genomics data and artificial intelligence (AI) have led to the emergence of 'sequence-to-function models (seq2func models)' that predict gene expression levels from DNA sequence. This AI model predicts the regulatory mechanisms of target genes at the molecular level, opening a new paradigm in genomics research. However, existing models exhibit a vulnerability in that their predictive power drops sharply when applied to datasets outside their training range. Furthermore, the mechanisms by which gene regulatory principles operate within complex neural networks are opaque, and reliable validation remains a challenge.

Key Findings

This review article identifies model architecture, training data, prediction tasks, model interpretation, and evaluation strategies as key factors determining the performance and reliability of seq2func models. First, in terms of model architecture, Convolutional Neural Networks (CNNs) are good at identifying local regulatory patterns (motifs) within the DNA sequence. In contrast, Transformer-based models effectively capture long-range enhancer-promoter interactions of 100 kilobases (kb) or more by applying a self-attention mechanism. Recently, State Space Model (SSM)-based HyenaDNA, which significantly reduces computational complexity while handling long sequences of up to 1 megabase (Mb), has been studied as a new alternative. However, the limitations of training data and prediction tasks remain. Most seq2func models are trained on epigenome data that is biased towards specific cell lines or chromosomes, which can lead to the model learning statistical noise in the data rather than actual biological rules. The inability to distinguish between cell types in bulk data is also a factor that reduces generalization ability. Accordingly, the methods for verifying the AI's predictive basis are also diversifying. Integrated Gradients, which shows the important positions in the input sequence, and In Silico Mutagenesis (ISM) are representative examples. The researchers propose the introduction of Global Importance Analysis (GIA), which quantifies biological features such as the distance or order of regulatory patterns, beyond single base analysis. Furthermore, a rigorous evaluation strategy that utilizes untrained chromosomes or new cell states during model validation is necessary to accurately measure the model's generalization ability.

Significance and Prospects

This analysis shows that the field of genomic AI should move away from simple parameter expansion competitions and focus on interpretability and robustness that explain biological causality. In particular, the framework that comprehensively verifies the impact of model structure and data characteristics on generalization performance is expected to be a useful guide for researchers to design more sophisticated neural networks in the future. However, there are also practical challenges to be overcome. Most of the existing seq2func models do not fully reflect dynamic physiological phenomena, such as changes in the concentration of transcription factors (TFs) or three-dimensional structural changes in chromatin, during the learning process. Therefore, the introduction of context-based models that integrate cell state information into the input is essential. Furthermore, a standardized benchmark platform for evaluating whether the AI model has learned actual biological rules must be established in order to achieve tangible results in the field of precision drug development.

Nature Genetics, Published online: 27 July 2026; doi:10.1038/s41588-026-02670-3This Review surveys the current landscape of genomic artificial intelligence through the lens of sequence-to-function models, examining how architectural choices, training data, prediction tasks, model interpretation and evaluation strategies can shape their generalization.

๐Ÿ’ฌWhy it matters:

The genomic AI validation guidelines presented in this study can bring immediate changes to the personalized precision medicine and drug development industries. A typical scenario is the identification of novel non-coding genetic variants in patients that cause rare diseases. If AI can accurately predict pathogenic variants that directly impair gene expression, clinicians will be able to easily identify the actual cause of the disease from tens of thousands of genetic variants. Furthermore, the efficiency of designing synthetic promoters that optimize the activity of therapeutic genes in gene therapy development can be maximized. This involves designing a sequence with AI to be strongly expressed only in the desired cell type, and then verifying the stability of the design in real-time using GIA. This simulation will help screen effective candidate materials before animal experiments, resulting in reduced drug development time and costs.

๐Ÿ’ฌ Comments

0 comments
Please log in to comment
Loading...