AI Deciphers the Genome Language, Mapping Symbiosis of 100,000 Microbial Species on Earth

Background
Microbial symbiosis plays a wide range of roles, from the host's nutrient acquisition and stress resistance, disease defense, to carbon and nitrogen cycling in ecosystems. However, many bacteria and archaea cannot be cultured in the laboratory, making it difficult to confirm how they survive and with whom. Known symbiont genomes have been biased toward species isolated from humans or animals and those that are easy to culture, making it difficult to estimate the scale of symbiotic relationships in real ecosystems.
Symbionts tend to lose unnecessary metabolic genes and genomic regions as their dependency on the host increases. However, it is not possible to determine symbiosis solely based on small genomes. Free-living microbes can also reduce their genomes during environmental adaptation, and species loosely associated with hosts tend to retain relatively complete metabolic capabilities. The research team believed that a machine learning system that reads the overall functional composition of genomes, rather than individual genes, could distinguish between free-living, host-associated, and obligate intracellular types, thereby overcoming these limitations.
Key Findings
The research team developed a symbiont prediction framework called 'symclatron.' First, they inferred orthologous gene groups from 792 symbiont proteomes selectively sampled from major clades, and created 26,300 profile hidden Markov models in groups with five or more members. They then used 6,751 microbial genomes with confirmed lifestyles, categorized into 5,959 free-living, 409 host-associated, and 383 obligate intracellular, as training data.
The classifier 'symcla' and the regression model 'symreg,' which represents host dependency as a continuous value, each used 1,000 selected influential features. The final neural network integrated seven factors, including the predicted values of the two models, genome completeness, and distance from training clades. The research team also performed clade-level cross-validation by sequentially removing 11,025 clades from the training data, a design that not only tests whether it can correctly identify close relatives of known species but also whether it can be applied to unseen clades.
The F1 scores for free-living, host-associated, and obligate intracellular types in the validation data were 0.982, 0.699, and 0.908, respectively. It accurately classified 97.6% of obligate intracellular symbionts, with 5.2% of host-associated types misclassified as obligate intracellular and 2.3% of obligate intracellular types misclassified as host-associated. This also revealed the biological characteristic that the boundary between the two symbiotic categories is continuous.
The symclatron was applied to metagenome-assembled genomes (MAGs) and reference genomes recovered from 31,152 environmental sequencing projects. After quality control and deduplication based on an average sequence identity threshold of 95%, 107,067 bacterial and archaeal genomes were obtained, and 14,070 were classified as symbionts at a confidence threshold of 0.725, accounting for about 14% of the total, or roughly one in seven. The proportion of symbionts was 15β23% when considering only MAGs, and candidates were found in half of the known bacterial and archaeal phyla. These genomes were published as the 'Symbiont Genomes (SymGs)' catalog.
Significance and Prospects
This study opens the way to large-scale screening of symbiotic potential based solely on functional genomic traces, without the need for cultivation or host observation. The model classified the lineage of 'Candidatus Azoamicus ciliaticola,' which supplies energy to a flagellate host, and the newly reported archaeon 'Candidatus Sukunaarchaeum mirabile' as obligate intracellular symbionts, indicating that the model can explore dependencies not only among eukaryotes but also between bacteria and archaea.
Interpretation requires caution. The training data were biased toward bacteria associated with eukaryotic hosts, and archaeal cases were limited. Performance decreased in the most phylogenetically distant groups, and the F1 score for host-associated types was also lower than for the other two categories. It is also not possible to determine the interaction partner or whether the symbiosis is mutualistic or parasitic based on predictions alone. SymGs is closer to a candidate map for prioritizing follow-up cultivation, microscopic observation, single-cell analysis, and metabolic experiments, rather than a finalized list of verified symbionts.
Nature Biotechnology, Published online: 02 September 2026; doi:10.1038/s41587-026-03212-2A machine-learning framework called symclatron uses genome sequencing content to predict whether uncultivated bacteria and archaea live independently or in close association with hosts. Applied to a global genome collection, it suggests that host-dependent microbes are widespread across Earthβs biomes and microbial phyla.
By inputting microbial genomes obtained from environmental samples into symclatron, candidates with high potential for host dependency can be prioritized. For example, in plant root zones, candidates related to nutrient uptake or pathogen suppression can be selected for use in designing synthetic microbial communities, or in marine samples, symbiotic pairs mediating carbon and nitrogen cycling can be tracked with focus. In human and livestock microbiomes, the metabolic deficiencies of host-dependent microbes that are difficult to culture can be predicted, aiding in the design of customized media and co-cultivation conditions. However, actual host verification, functional validation, and biosafety evaluation are required for industrial strains or therapeutic targets.