The UCE single-cell AI foundation model breaks down species and tissue boundaries with 36 million cell data points

Background
Single-cell genomics, which analyzes individual cells, has revolutionized biological research. In particular, single-cell RNA sequencing (scRNA-seq) is generating vast amounts of gene expression data from numerous cells every year. However, it has been challenging to fully integrate and utilize this massive dataset. Batch effects, which are technical errors arising from differences in analysis instruments, research environments, or reagent types, distort biological signals.
Existing batch correction algorithms have limitations because they rely on pre-aligned pairs of datasets. Whenever new research data is added, the existing model must be retrained from scratch or undergo a large-scale fine-tuning process, which exponentially increases the cost of analysis.
Barriers between species are also difficult to overcome. Although humans and laboratory mice have high genetic similarity, the subtle differences in their gene sequences and expression patterns make it nearly impossible to simply compare cell data from different species. Existing approaches that compare homologous gene information one-to-one do not reflect genes that have been lost or newly created during evolution, resulting in lower accuracy. For these reasons, the biological community has long sought the development of a universal analysis tool that can integrate data from various species and laboratories without restriction.
Key Findings
The Universal Cell Embedding (UCE) model, developed by researchers at Stanford University, solves these challenges at once. UCE is defined as a transformer-based cell foundation model that has been pre-trained on 36 million single-cell atlas data points from eight species, including humans.
The model's core algorithm involves vectorizing the amino acid sequence of each gene using the protein language model ESM-2 to create a gene embedding, and then combining it with the gene expression information of a single cell. By using the gene sequence itself as a primary source of information, UCE acquires a universal biological representation that transcends species boundaries. Even if the organisms are evolutionarily distant, such as humans and zebrafish, if the cells have similar functions, the AI places them in the same high-dimensional latent space.
To verify the performance of UCE, the researchers conducted a zero-shot validation by analyzing a new dataset that was not included in the pre-training dataset. UCE successfully removed batch effects and accurately classified cell types in the new brain tissue cell data without any additional training or fine-tuning. In the process, an Integrated Mega-scale Atlas (IMA) encompassing 36 million cells was naturally created. This is the first case of hundreds of independent research datasets being combined into a single, massive cell map.
Significance and Prospects
The emergence of UCE is expected to mark a turning point in single-cell biology, shifting from individual experiment-centric analysis to a universal exploration system based on large language models. Scientists can now map their sequenced small datasets to the 36 million cell map defined by UCE and compare them. This is expected to significantly shorten the research cycle and maximize the efficiency of data sharing and integrated analysis.
It will also open the way for detailed elucidation of how the developmental processes of animals or the differentiation pathways of immune cells have been preserved in evolutionary history through interspecies comparative analysis. However, there are also opinions that UCE is not a perfect solution.
Current UCE is trained primarily on gene expression information, so its ability to organically integrate spatial transcriptomics or multi-omics data at the protein level still needs to be improved. The massive graphics processing unit (GPU) resources and infrastructure maintenance costs required to compute 36 million cells are also challenges that the research community needs to address together. The reliability of the cell connections suggested by the model will only be established when subsequent experimental validation supports them.
Nature Genetics, Published online: 07 August 2026; doi:10.1038/s41588-026-02718-4Universal cell embedding for single-cell biology
The UCE technology has the potential to revolutionize the paradigm of drug development in the pharmaceutical industry and clinical medicine. The most direct application scenario is the prediction of the efficacy of drug candidates across species. In the past, drug candidates that have been proven in animal models such as mice have often failed in human clinical trials, which is due to inconsistencies in cellular responses between species. By using UCE's common cell space, it is possible to mathematically and precisely predict how the state changes of mouse cancer cells in response to a specific drug will manifest in tumor cells of human patients. This is expected to contribute to the selection of promising drug candidates in the drug screening stage and to reduce the risk of failure in clinical trials.
It will also be a powerful tool in the study of rare diseases. In the case of rare diseases with a small number of patients worldwide, it is difficult to obtain sufficient single-cell samples. UCE's zero-shot function can be used to input a small amount of patient cell data into an existing large healthy cell atlas, allowing for ultra-precise tracking of deviations in expression levels between normal and diseased cells. This is a key to accelerating the elucidation of the pathological mechanisms at the level of dysfunctional cells and the discovery of personalized targeted therapies.