๐Ÿ’ปCode of Life

Genomic AI overcomes limitations of extremely small data to elucidate the function of unknown enzymes

PNASยทAugust 5, 2026AI Curation
Genomic AI overcomes limitations of extremely small data to elucidate the function of unknown enzymes
โœจAI Summary (Beta)Beta

Background

Artificial intelligence (AI) is increasingly influential in protein structure prediction and genomics. Existing AI models require vast amounts of training data, ranging from thousands to tens of thousands, to ensure performance. However, in practical biological research, high-quality, experimentally validated data is extremely scarce, hindering the adoption of machine learning (ML) techniques. In particular, identifying unknown enzymes that modify phenazine, an organic compound affecting ecosystems and the human body, has been a long-standing challenge. Phenazine is a toxic substance secreted by bacteria, and tracing the biochemical reactions that render it harmless requires numerous control experiments. However, the genetic information related to this is very limited, making it nearly impossible to build a predictive system using conventional AI techniques, according to the researchers.

Key Findings

The research team, led by Professor Dianne Newman at the California Institute of Technology (Caltech), developed 'ML-CITO', an AI framework that operates with extremely small data by utilizing genomic contextual information. The model started by using only 14 known phenazine-modifying enzyme gene sequences as initial training data (seed data). Subsequently, the researchers introduced a genomic data augmentation technique to track homologous genes located near the phenazine biosynthetic gene cluster (BGC), increasing the training data to approximately 600. The augmented data was processed through a pre-trained protein language model (PLM), 'ESM Cambrian 600M', and transformed into a high-dimensional vector of 1,152 dimensions. This vector information was input into a three-layer multi-layer perceptron (MLP) neural network consisting of 512, 128, and 64 units, followed by the application of contrastive learning. The ML-CITO model precisely classifies candidate enzymes that react with phenazine within the protein space. The researchers successfully identified a protein from Pantoea agglomerans W2I1, a soil bacterium, that exhibits actual enzyme activity among the model's predicted candidates. The newly identified enzyme is a phenazine-thiol conjugating enzyme (PTC) that directly binds phenazine with glutathione (GSH), which regulates intracellular redox status. Previously, the conjugation reaction between phenazine and GSH was considered a non-enzymatic chemical reaction. The researchers demonstrated through biochemical experiments that the PTC they discovered catalyzes the reaction, directly mitigating the toxicity of phenazine. The researchers searched over 200,000 bacterial genomes and confirmed that approximately 3,415 PTC gene homologs are widely distributed across more than 30 phyla.

Significance and Prospects

This research presents a new breakthrough for research fields where protein data is extremely limited, making it difficult to apply ML models. By combining genomic contextual information with PLMs, the study demonstrates that high-performance prediction is possible even with small amounts of data. This framework is expected to be introduced in various bio-industrial fields requiring the elucidation of unknown enzyme functions, such as drug development, biomanufacturing, and environmental remediation. However, the fact that the ML-CITO model relies on data augmentation based on conserved physical locations in the genome (e.g., BGC) is a challenge to be overcome. Proteins with unclear genetic context or those scattered throughout the genome may not fully benefit from the data augmentation effect. The researchers plan to expand the model's applicability to a wider range of protein clusters and to verify the detailed biochemical characteristics of the 3,415 gene homologs identified in this study.

Proceedings of the National Academy of Sciences, Volume 123, Issue 31, August 2026. SignificanceMachine learning excels when large, well-labeled datasets are available, yet many biologically important problems lack sufficient experimental data to support such approaches to discovery. This limitation is particularly acute for identifying ...

๐Ÿ’ฌWhy it matters:

The practical value of this research is particularly evident in the development of new therapeutic strategies to control antibiotic-resistant strains. Phenazine toxins secreted by Pseudomonas aeruginosa, an opportunistic pathogen, induce chronic inflammation in patients with lung diseases and are key factors in increasing antibiotic resistance. By using this model to design inhibitors that interfere with the action of phenazine or regulate the activity of PTC, it may be possible to improve the efficacy of drugs for treating resistant bacteria. Furthermore, in the field of environmentally friendly bioprocesses, it is possible to quickly find unknown useful enzymes and improve the efficiency of synthesizing high-value-added chemicals. Redesigning the metabolic pathways of beneficial microorganisms that decompose organic waste or purify recalcitrant compounds is also a promising application.

๐Ÿ’ฌ Comments

0 comments
Please log in to comment
Loading...