Genomics, proteomics, systems biology — decoding the blueprint of life.

Background The vaginal microbiome is closely linked to women's reproductive health, including bacterial vaginosis, Human Papillomavirus (HPV) infection, and preterm birth. However, the 16S rRNA analysis commonly used in existing studies struggles to distinguish bacteria at the species/strain level and fails to capture viruses and fungi. Furthermore, sample compositions centered on Western populations and pregnant women have limited the ability to explain microbial diversity in the global population. The sample contains over 90% human DNA, making microbial genome recovery difficult. Paradoxically, leveraging this human DNA allows for the simultaneous analysis of host genetic variants and microbial composition from the same sample. The research team integrated 2,967 public datasets and 1,433 bacterial isolates centered on the metagenome of 10,005 Chinese individuals to trace the connections between hosts and microbes. The study was published in Nature Genetics in June 2026, and a September correction notice only rectified the author names and the position of Figure 5b. The research findings and conclusions remained unchanged. Key Findings The Global Viral Metagenome Assembled Genomes (GVMG) constructed by the research team consists of a total of 65,055 genomes. This includes 36,059 genomes from 890 prokaryotic species, 43 genomes from 11 fungal species, and 28,953 genomes from 6,577 viral taxonomic units. The number of genomes is 1.9 times greater than that of the existing VMGC, and lineages comprising 13.0% of prokaryotic species and 79.0% of viral taxonomic units were not present in existing public repositories. Among the prokaryotic genomes, the proportion of high-quality or near-complete genomes was 68.1%, surpassing the 48.2% observed in VMGC. Differences between populations were distinct. The community type dominated by BVAB1, a bacterium related to bacterial vaginosis, was 0.46% in the Chinese cohort but 12.38% in the US cohort. Among viruses, 56.5% could not be found in other major vaginal, intestinal, or oral virome databases, and phage-host associations corresponding to the dominant lactobacilli, Lactobacillus iners and L. crispatus, were also identified. The researchers analyzed the relationship between 5.46 million human genetic variants and 54 species of gut microbiota using data from a discovery cohort of 3,137 individuals and a validation cohort of 3,227 plus 506 individuals. As a result, they identified seven host gene loci that met the study-wide significance level and were replicated in an independent cohort. The strongest signal was the association between an OPRK1-adjacent variant on chromosome 8 and the potential pathogen Ureaplasma urealyticum. The effect size β was 1.24, with a P-value of 1.50×10⁻⁵⁵. Associations involving ADAP1-Lactobacillus mulieris and an PRAMEF1-adjacent variant-Bifidobacterium piotii were also replicated in the same direction across both validation cohorts. Significance and Outlook This study expands the view of the vaginal microbiome from simple bacterial proportions to an ecosystem encompassing viruses, fungi, and intra-strain genetic variations. It revealed that the same microbial species can differ in genetic and functional composition depending on the population of origin, suggesting that diagnostic criteria or probiotics developed in specific regions may not be directly applicable to other populations. However, association does not imply causation. Whether OPRK1 variants directly regulate the colonization of U. Whether urealyticum* directly regulates colonization, or whether hormones, immunity, and environmental factors intervene, must be confirmed through cellular and animal experiments. A limitation remains that the large-scale new data is centered on Chinese populations. Longitudinal studies encompassing diverse ancestral groups, followed by functional validation, are necessary to enable genotype-based infection risk prediction or personalized microbial therapies.
💡 The GVMG can serve as a reference resource for identifying bacterial strains, phages, and fungi that existing databases miss during shotgun metagenomic analysis of vaginal samples. For example, it enables precision diagnostics, such as differentiating Gardnerella lineages when analyzing samples from patients with recurrent bacterial vaginosis or tracking which strains and phages remain after antibiotic administration. In the industry, it can be used to select probiotic and phage therapeutic candidates that reflect regional microbial differences. However, it is premature to use tests that judge infection or disease risk based solely on host genetic variants. Before clinical application, reproducibility across diverse populations, microbial community fluctuations over time, and confounding factors such as medication, sexual activity, and hormonal status must be validated.

Background Lung adenocarcinoma (LUAD) is the most common type of non-small cell lung cancer and exhibits extreme molecular heterogeneity. While the introduction of Epidermal Growth Factor Receptor (EGFR) or Anaplastic Lymphoma Kinase (ALK) inhibitors has improved treatment outcomes, many patients progress due to a lack of targetable mutations or the acquisition of drug resistance. This creates an urgent need for new target discovery for patient groups with exhausted treatment options. As large-scale CRISPR functional screening data, such as the Cancer Dependency Map (DepMap), has accumulated, the search for essential genes for cancer cell survival has become active. However, cell line-based screening in culture dishes fails to fully reflect the complexity of the actual patient tumor microenvironment. Relying solely on simple dependency scores exposes the flaw of including a large number of false-positive targets that appear effective in vitro but are not reproducible in actual patient tissues. This background necessitates a computational screening strategy to bridge the gap between laboratory-level functional vulnerabilities and clinical patient cohorts. Key Findings The researchers constructed DepPrior, a computational framework based on conjunctive ranking rules that combines multiple criteria. It is structured to strictly prioritize genes that simultaneously satisfy three complementary evaluation criteria. The first criterion is the separation of CRISPR gene dependency. Clear dependency variations among cell lines allowed for the primary selection of genes that exhibit selective vulnerability only in specific cancer cells. The second criterion for molecular predictability involved training linear and nonlinear machine learning models using gene expression levels and copy-number variation (CNV) data from cell lines as input features. The researchers predicted DepMap dependency scores to calculate the area under the receiver operating characteristic curve (AUROC) and the coefficient of determination (R²) at the gene level. The third is multi-cohort reproducibility. Validation was performed to confirm whether consistent expression patterns were maintained at the transcriptome and proteome levels of patient tumors when compared with large-scale independent clinical databases such as TCGA-LUAD, GEO, and CPTAC. The researchers designed DepScore, an index combining AUROC and R², to prioritize candidate genes. As a result of the analysis, FERMT2, CRKL, MYC, and CHMP4B were identified as the top lung adenocarcinoma target candidates. These genes show a pattern of robustly forming gene expression modules directly linked to cancer cell proliferation in TCGA-LUAD patient data. To prove the validity of the computational predictions, orthogonal experimental validation using the lung cancer cell line HCC827 was also performed. Knocking down the expression of FERMT2 and CRKL within the cells resulted in a decrease in target protein levels along with a distinct increase in apoptosis-related signaling proteins. In particular, a synergistic effect was observed where the change in apoptosis-inducing proteins was significantly more amplified under simultaneous inhibition of both genes compared to single-gene inhibition conditions. Significance and Outlook DepPrior serves as a bridge smoothly connecting the noise of cell line screening with the complexity of patient tissues. This is due to its design, which integrates functional genomics and patient multi-omics reproducibility into a single pipeline, moving away from the conventional reliance on simple transcriptomic correlation analysis or dependence on single cell lines. It is expected to function as a hypothesis-generating platform that reduces R&D costs and failure rates by narrowing down effective target candidates in the early stages of drug development. The amplification of apoptosis signals induced by the simultaneous inhibition of FERMT2 and CRKL presents a new possibility for combination therapies in lung adenocarcinoma, which had previously been limited to single-target approaches. Since the study focused on computer-based data prediction and functional inhibition at the cell line level, it also has the limitation that additional rescue experiments are required to demonstrate the structural binding of target proteins. Verification of in vivo drug efficacy and safety using patient-derived cancer organoids and animal xenograft models is also a task to be addressed in the future.
💡 DepPrior can be immediately deployed in the early target discovery stage of the drug development pipeline. For non-small cell lung cancer patients who do not respond to existing EGFR inhibitors or have developed resistance, developing dual-target inhibitors or small-molecule combination therapies targeting FERMT2 and CRKL—whose synergistic induction of apoptosis has been verified—represents a promising scenario. Pharmaceutical companies can preemptively secure high-purity target candidates with verified patient omics reproducibility before engaging in laborious and costly laboratory screening. It is evaluated that this will effectively serve as a computational filter, significantly reducing trial and error in the target validation stage and shortening the timeline for identifying active compounds.

Background Personality is a unique pattern of behavior through which individuals perceive and react to the world, and has long been a major area of exploration in psychology and behavioral science. In clinical settings, observations have steadily accumulated showing that personality traits such as neuroticism or extraversion are closely linked to the risk of developing depression, cardiovascular disease, and metabolic syndrome. However, observational studies alone made it difficult to clearly determine whether specific personality traits cause diseases, or whether underlying diseases or common environmental factors influenced personality formation. This hit a limit in clearly proving the sequence of causality. Existing genomic studies also failed to fully clarify the polygenic nature of personality due to constraints in sample size. Because human behavior and traits are shaped by the complex interplay of numerous minor genetic variants, it was difficult to achieve statistical significance with cohorts in the tens of thousands. This lack of samples has also been a long-standing obstacle in research revealing the shared biological mechanisms between mental and physical diseases. Furthermore, there was also a lack of large-scale analysis infrastructure to correct for the subjective bias inherent in self-reported survey data. Researchers have now embarked on a study to elucidate the genetic structure of personality traits and rigorously verify their potential causality with diseases by integrating large-scale genomic data from over one million people. Key Findings Researchers performed a Genome-Wide Association Study (GWAS) on a multinational cohort of over one million people to identify numerous genetic variants involved in personality traits. This achievement precisely identified genomic loci associated with major personality scales such as extraversion, neuroticism, and conscientiousness and verified their statistical associations. The study is evaluated to have dramatically increased statistical power compared to previous studies by combining vast genotypic data with personality phenotypes. Subsequently, they applied Mendelian Randomization (MR), using genetic variants as instrumental variables, to trace the causal pathways between personality traits and health outcomes. Analysis results indicated that genetic variants associated with neuroticism traits had potential causal effects not only on major depressive disorder and anxiety disorders but also on coronary artery disease and elevated chronic inflammation markers. This result supports the idea that high emotional instability can cause disturbances in the neuroendocrine and immune systems in the long term, increasing the risk of physical diseases. Conversely, genetic indicators associated with conscientiousness and extraversion were directly linked to protective effects, such as regular physical activity, reduced smoking rates, and longer lifespan. This point statistically demonstrates that, beyond simple correlations between personality and physical diseases, common genetic variants and biological pathways can directly contribute to disease pathogenesis. Significance and Outlook This study holds academic value in converting personality traits, which were treated as psychological concepts, into molecular genetic and statistical genetic data. This demonstrates that mental health and physical illnesses interact along a single biological axis, based on large-scale genomic big data. In the future, integrating personality genetic variants into Polygenic Risk Score (PRS) models will lay a solid foundation to significantly increase the predictive precision of personalized preventive medicine. However, challenges remain to be solved before application in clinical settings. A clear limitation is that the study cohort is predominantly of European ancestry, making it difficult to generalize the results to multi-ethnic populations. Critics also point out that population stratification in complex trait studies, as well as subtle residual confounding factors, cannot be completely excluded. Follow-up research is expected to significantly expand data from non-European populations and link it with brain imaging data and single-cell transcriptome analysis to elucidate the actual mechanisms by which genetic variants affect neural circuit development.
💡 The results of this study can be directly utilized in establishing early disease screening and customized intervention strategies in clinical practice and the digital healthcare industry. In frontline medical institutions, by integrating personality-based polygenic scores with a patient's physical examination indicators, a system can be established to precisely screen high-risk groups for the co-occurrence of depression and cardiovascular disease. In the field of Digital Therapeutics (DTx), scenarios could become reality where cognitive behavioral therapy algorithms are designed to align with an individual's genetic predispositions for neuroticism or conscientiousness, thereby maximizing patient medication and treatment adherence. This heralds the birth of a next-generation integrated preventive management platform that combines genomic information with individual psychological and behavioral tendencies.

Background During outbreaks of novel infectious diseases or variant viruses, genomic sequence data serves as a critical resource for identifying transmission pathways and mutation rates. Health authorities reconstruct the evolutionary history of the virus using phylogenetic trees to estimate when and where the virus was introduced and how it spread. The technique most widely recognized as the most accurate statistical framework in this process is Bayesian phylogenetics. This is because it can simultaneously estimate key epidemiological metrics, such as evolutionary rate and reproduction number, while systematically accounting for uncertainty. The problem is computational load. Existing Bayesian phylogenetic tools use Markov Chain Monte Carlo (MCMC) algorithms. As sample sizes exceed hundreds, computation time increases exponentially, often taking days or weeks to complete analysis. During a pandemic, when tens or hundreds of thousands of whole-genome sequencing datasets are rapidly generated, existing methods struggle to contribute in time to real-time epidemic prevention decision-making. Developing nations or regional health agencies lacking high-performance computing infrastructure are unable to even attempt to apply the latest phylogenetic analysis results directly to disease control efforts. Key Findings The research team has unveiled Delphy, a new computational framework capable of processing accumulating large-scale viral genomic data in near real-time. The core of Delphy is an online Bayesian update structure that organically integrates new sequences into the previously learned posterior probability distribution of the phylogenetic tree, rather than recalculating the entire tree from scratch every time new data arrives. Thanks to algorithmic optimization, Delphy drastically reduces computational complexity even as samples accumulate. While existing tools took dozens of hours on high-performance clusters to process thousands of sequences, Delphy completed Bayesian analysis within dozens of minutes on standard workstation-class computers. It maintained the highest level of statistical accuracy while increasing speed. Verification through simulations and actual infectious disease genomic datasets demonstrated that the posterior probability distributions of divergence time estimates and phylogenetic tree topologies achieved accuracy consistent with traditional full MCMC methods. This means high-level phylodynamic analysis can be performed with minimal computational cost even in resource-limited environments. Significance and Outlook The emergence of Delphy marks a turning point in shifting the paradigm of genomic surveillance from centralized post-hoc analysis to decentralized, real-time local surveillance. This is because local laboratories can perform phylogenetic analysis immediately as data is generated and feed the results back into quarantine operations. It provides a foundation for public health agencies worldwide to independently operate standardized Bayesian precision analysis while maintaining their own data sovereignty. Challenges remain. For pathogens with high recombination rates or complex insertion/deletion mutations, additional phylogenetic modeling beyond simple point mutation models is required. The risk of bias during long-term updates cannot be ruled out if quality variations or sequencing errors in new data accumulate. Follow-up research will require expanding pipelines to flexibly accommodate various mutation mechanisms and automate the data cleaning process.
💡 Delphi provides immediate utility for establishing strategies to block initial influxes and community spread during a pandemic. When a new variant enters through airports or seaports, its evolutionary origin and transmission speed can be confirmed within half a day of analyzing the collected genome. In particular, low- and middle-income countries or regional health centers that cannot rely on large-scale supercomputer centers will be able to independently operate precision infectious disease surveillance networks using standard desktop equipment. In hospital epidemiological settings, it is expected to demonstrate high practical utility by determining on the same day whether a hospital outbreak is due to intra-hospital transmission or multiple external introductions, allowing for precise decisions regarding isolation wards and disinfection scope.

Background Clostridioides difficile infection (CDI) is considered one of the most challenging healthcare-associated infections in clinical settings. When the gut microbiota ecosystem collapses in patients who have received long-term broad-spectrum antibiotics, the anaerobic spore-forming bacterium C. difficile overgrows and destroys intestinal epithelial tissue. This bacterium, which can cause everything from mild diarrhea to intestinal perforation and life-threatening pseudomembranous colitis, poses a fatal threat to the elderly and immunocompromised patients. Currently, specific antibiotics such as vancomycin or fidaxomicin are used as first-line treatments in clinical practice. However, antibiotic treatment leads to a vicious cycle where the destruction of normal intestinal flora causes recurrence in over 20% of patients. While attempts have been made for a long time to develop vaccines that neutralize toxins or fundamentally block bacterial proliferation, traditional attenuated live vaccines or recombinant protein toxoid approaches have failed to advance beyond the clinical stage due to insufficient responsiveness to variant strains and inadequate immunogenicity. As clinical limitations have revealed that single antigens are insufficient to simultaneously block toxin secretion and intestinal colonization, the need for a next-generation vaccine platform that precisely combines multiple immune epitopes has emerged. Recently, a breakthrough for solving this challenge is opening up through the combination of messenger RNA (mRNA) technology, proven in infectious disease response, and virtual design technology using high-performance computers. Key Findings Instead of synthesizing numerous antigens one by one on laboratory benches, the research team adopted a reverse vaccinology approach that precisely analyzes pathogen genomes and protein structures in a computer-based virtual environment. We designed a multi-epitope mRNA vaccine structure that targets the core protein sequences involved in Diplocystis infection and pathogenicity expression, thereby simultaneously inducing both humoral and cellular immunity. Using computational models, they precisely selected B cell epitopes that stimulate antibody production, along with Helper T lymphocyte (HTL) epitopes and Cytotoxic T lymphocyte (CTL) epitopes that directly eliminate infected cells. For the identified candidate sequences, safety was ensured by strictly screening their potential to induce in vivo immunity, as well as their allergenicity, cytotoxicity, and risk of attacking self-tissues (autoimmune risk), using computational algorithms. The selected epitopes were connected with optimal linkers and reconstructed into a single fusion construct. This was followed by evaluations of the vaccine protein's physicochemical properties and population coverage across diverse ethnic and genetic backgrounds. As a result of performing 3D molecular docking to confirm binding affinity with immune receptors and immune response simulations, it was observed that the designed vaccine could effectively induce strong antibody responses and interferon-gamma secretion. Finally, by applying molecular dynamics (MD) simulations on a scale of tens of nanoseconds, the molecular dynamics stability was confirmed, showing that the protein's three-dimensional structure remains stable under conditions similar to the biological environment. Significance and Prospects This study demonstrates that rational design of multi-antigen-target mRNA vaccines is feasible, significantly reducing the substantial costs and time associated with laboratory-stage development. It is noteworthy that computer calculations have provided a blueprint capable of precisely targeting the complex immune evasion strategy of Clostridioides difficile, which possesses a triple mechanism comprising spore formation, intestinal tract attachment, and exotoxin secretion. Clear challenges also remain. It is difficult to conclude that in-silico simulation results perfectly match in vivo immune responses. Wet-lab validation is essential to confirm whether the antibody titers and T-cell immune responses predicted by computational models are accurately reproduced in actual animal models. Research on optimizing mucosal administration formulations to enhance lipid nanoparticle (LNP) encapsulation efficiency and local intestinal immunity must be conducted in parallel. Because the immune environment on the surface of colonic epithelial cells differs from that in the systemic circulation, subsequent preclinical studies demonstrating the extent of local mucosal immune induction, including secretory IgA production, will be the critical key to commercialization.
💡 This study provides a concrete turning point for protection strategies for high-risk patients in general hospitals and nursing homes where urgent management of hospital-acquired infections is required. This is because the foundation has been laid for an mRNA vaccine pipeline that can be administered as a prophylactic pre-procedure dose to elderly hospitalized patients or those on long-term antibiotic therapy. It overcomes the limitations of existing antibiotic therapies that cause frequent recurrences by destroying beneficial intestinal bacteria, and it can establish a preventive immune barrier that pre-emptively controls bacterial colonization and toxin activity. Industrially, it has significant potential to evolve into a next-generation vaccine development platform that can identify multi-valent vaccine candidates against new variant strains in silico within days of securing genomic analysis data and proceed directly to the synthesis stage.

Background The polygenic risk score (PRS), which calculates the future risk of disease onset by aggregating individual DNA variants, has been recognized as a key tool in precision medicine. It is a method that statistically sums the effects of hundreds of thousands to millions of single nucleotide polymorphisms (SNPs) to assess individual susceptibility. However, existing genome dynamics research was trapped by a structural limitation of severe sample imbalance. This is because more than 80% of participants used in Genome-Wide Association Studies (GWAS) were biased toward European ancestry. This distortion in genomic data has resulted in a sharp decline in prediction model performance for non-European populations. When applying existing PRS algorithms to African or Hispanic groups, prediction accuracy can drop to less than half that of European groups. This is because linkage disequilibrium (LD) patterns and allele frequencies vary across populations when assessing disease susceptibility. Consequently, continuous criticism has been raised that advancements in genomic research are deepening healthcare inequalities by benefiting only specific racial groups. Key Findings Researchers adopted a large-scale multi-ancestry meta-analysis strategy by combining data from the National Institutes of Health (NIH) All of Us Research Program (AoU) and the UK Biobank (UKB). The AoU is a cohort that has secured a proportion of underrepresented minorities of over 50%, groups historically excluded from medical research. The researchers reconstructed models for complex traits and major chronic diseases by integrating Whole Genome Sequencing (WGS) and genotyping data from hundreds of thousands of individuals registered in both large biobanks. By applying a multi-ancestry Bayesian regression technique that simultaneously learns the genetic architectures of diverse populations, the predictive performance of the models in non-European groups improved significantly. In major chronic diseases such as type 2 diabetes, coronary artery disease, and hypertension, the prediction accuracy (AUC) for African American and admixed populations significantly improved compared to existing single-ancestry-based models. Notably, in the African-ancestry cohort, the prediction performance for continuous traits such as systolic blood pressure and body mass index (BMI) improved by up to 40%. This was achieved through an algorithm that selects causal genetic variants acting commonly across multi-ancestry groups and finely corrects for the unique LD structures of each group. This proves that securing genetic diversity, rather than just the absolute expansion of sample size, is the decisive factor in determining the performance of prediction algorithms. Significance and Outlook This study demonstrates that expanding data scale and collecting multi-ancestry samples are practical solutions to bridging the ancestry gap in clinical genomics. A turning point has been reached in expanding the scope of precision preventive medicine, which was previously limited to specific racial groups. As multi-ancestry algorithms become widespread, they will further enhance the accuracy of high-risk group identification by reducing false-negative errors occurring in early screening systems. In the global pharmaceutical industry, a path has opened to pre-verify the genetic validity of target discovery across diverse ethnic groups. However, a cautious approach is required for clinical application. Even though non-European samples have increased, data density is still insufficient to fully reflect the high genetic heterogeneity within African-ancestry groups. Additionally, the interaction of non-genetic factors, such as socioeconomic factors or living environments, with complex diseases remains a task to be addressed. In the future, international data linkage with local cohorts from Asia, South America, and the African continent must follow to achieve true global genomic equity.
💡 The multi-ancestry polygenic prediction models derived from this study can be immediately utilized in early screening programs for chronic diseases in primary healthcare institutions. This provides a pathway to early detection of patients with multi-ancestry and multicultural backgrounds who were previously missed by the conventional single-lineage model when applied to high-risk groups, allowing for the initiation of lifestyle improvements or preventive medication. There are also economic benefits, such as preventing waste in national healthcare budgets by accurately classifying high-risk groups for cardiovascular disease or metabolic syndrome. In the field of new drug development, it enhances the stratification of clinical trial participants and the efficiency of biomarker discovery. In global Phase 3 clinical trials, it reduces uncertainty in predicting drug response rates by minimizing genetic risk bias when selecting patient cohorts of diverse ethnicities. Genomic analysis companies can build commercial pipelines that provide genetic testing reports with uniform reliability to consumers of diverse ancestries.

Background The treatment paradigm for Non-Small Cell Lung Cancer (NSCLC) has changed dramatically with the advent of Immune Checkpoint Inhibitors (ICIs). Prescriptions have been guided by biomarkers such as programmed death-ligand 1 (PD-L1) expression levels or tumor mutational burden (TMB). However, in clinical settings, a significant number of patients still show outcomes different from predictions. Some patients show no response despite high PD-L1 expression, while others achieve long-term survival despite negative PD-L1 test results. This is because single molecular biomarkers struggle to fully capture the complex interactions between the Tumor Microenvironment (TME) and the host immune system. To address this discrepancy, efforts to integrate medical imaging, pathology slides, and genomic data through multi-omics approaches have continued. However, existing machine learning models have failed to cross the threshold for clinical adoption due to their 'Black Box' structure, where internal computational processes are unknown. It is difficult for medical staff to accept algorithms that cannot justify why a specific patient was classified as a responder, which in turn makes it hard to lead to actual changes in prescription. Achieving universality to overcome data disparities across various institutions and ethnic groups has also repeatedly proven to be a stumbling block. Key Findings A global collaborative research team has constructed a multimodal Explainable AI (XAI) model based on large-scale international Real-World Evidence (RWE) data. This study validated the algorithm's effectiveness on NSCLC patient cohorts from multiple countries with different healthcare systems. The method involves the fusion and analysis of digital pathology images (H&E stained slides), Computed Tomography (CT) scans, Whole Exome Sequencing (WES), transcriptome profiling, and clinical Electronic Health Records (EHR) into a single network. The model proposed by the research team significantly outperformed existing single indicators with a clear margin in the area under the receiver operating characteristic curve (AUC), an index of predictive accuracy for immune-oncology treatment response. It demonstrated discriminative ability exceeding 0.80, greatly surpassing the PD-L1 immunohistochemistry-based prediction value in the mid-0.60s and the TMB-based prediction value in the early-0.60s. Progression-Free Survival (PFS) and Overall Survival (OS) stratification also clearly separated the high-risk group from the low-risk group statistically. The core competitiveness lies in interpretability. Rather than merely outputting a prediction score, the model visualized the spatial density of Tumor-Infiltrating Lymphocytes (TIL) within pathology images, the infiltration characteristics of the tumor boundary in CT images, and specific chemokine expression pathways in the form of Attention Maps. Through this, the researchers confirmed that the algorithm places higher weight on the immune activation patterns at the tumor stroma boundary rather than the tumor parenchyma. In multi-center decision-support experiments involving several medical oncologists, the concordance and accuracy of physicians' treatment decisions significantly improved after reviewing the evidence provided by the model. Significance and Outlook Precision oncology, which previously relied on fragmented genetic tests or a few types of immunostaining, is evolving into a data-fusion-based diagnostic system. Organically weaving multi-dimensional data obtained from the patient's body revealed hidden clinical value. The fact that it possesses a proprietary interpretation module, thereby opening a pathway to alleviate medical distrust in AI, is also noteworthy. By providing the rationale for predictions, it has laid the foundation for enhancing trust in the treatment selection process between doctors and patients. The challenges to be overcome for commercialization are clear. A pipeline to standardize technical variations arising from differing H&E staining conditions across institutions, resolution discrepancies among CT scanner manufacturers, and sequencing platforms is essential. Follow-up work is also required to determine how closely the biological mechanisms proposed by the explainability module align with functional validation at the actual laboratory level. Furthermore, prospective randomized controlled trials meeting the regulatory standards for Software as a Medical Device (SaMD) must be completed.
💡 The results of this study may bring about practical changes for patients with advanced non-small cell lung cancer facing a choice in first-line treatment selection. Currently, the criteria for deciding between immunotherapy monotherapy and combination with cytotoxic chemotherapy in standard care are incomplete. If multimodal AI is integrated into hospital pathology and imaging reading systems, it can identify responders who do not require complex combination therapies, thereby reducing the risk of toxic side effects and saving treatment costs. For patients with a high probability of non-response, it becomes possible to establish customized treatment strategies, such as recommending early entry into other targeted therapies or clinical trials. For diagnostic kit developers and software medical device companies, it provides a clear basis for developing companion diagnostic (CDx) products based on multi-biomarkers.

Background Immune Checkpoint Inhibitors (ICIs), a cornerstone of cancer immunotherapy, have achieved near-curative long-term survival in some advanced cancer patients. Following cases of dramatic tumor shrinkage after drug administration, ICIs have gained attention as an alternative to overcome the limitations of conventional cytotoxic chemotherapy. However, in actual clinical practice, the proportion of patients receiving substantial therapeutic benefits remains around 20-30%. Many patients experience disease progression or face severe immune-related toxicities similar to autoimmune diseases despite receiving expensive drugs. Currently, clinical decisions are made using single biomarkers such as Programmed Death-Ligand 1 (PD-L1) protein expression levels or Tumor Mutational Burden (TMB). However, it has been pointed out that these single indicators cannot fully explain the complex tumor microenvironment and the patient's unique systemic immune ecosystem. In fact, frequent reports show starkly different treatment outcomes even within patient groups with identical TMB levels. This has led to the emergence of patient-level multimodal data integration analysis models that encompass genetic mutations in cancer tissue, transcriptome expression patterns, tumor-infiltrating lymphocyte distribution, and systemic inflammatory markers as an alternative. Key Findings Researchers constructed an ICI treatment response prediction model by integrating multimodal biological data at the patient level collected from several clinical cohorts. The configuration integrates Whole Exome Sequencing (WES)-based genomic variants, gene expression levels measured by RNA sequencing, immune cell infiltration quantified from pathological tissue slides, and patient peripheral blood test values using a single machine learning algorithm. Compared to the individual application of existing single biomarkers, the Area Under the ROC Curve (AUC) improved significantly by more than 0.15 when multimodal data were integrated and analyzed. This achievement allowed for the accurate identification of treatment-refractory patient groups that were difficult to distinguish using single markers alone. The explanation is that biological indicators forming a complex signaling network acted complementarily to offset the classification error of the prediction model. Despite this performance enhancement, a clear technical limitation—a decrease in predictive power—was exposed in external validation cohorts. The AUC, which exceeded 0.80 in the internal training dataset used to train the model, plummeted to the 0.60–0.65 range when applied to independent validation cohorts from other medical institutions. Technical discrepancies in specimen preprocessing methods and sequencing platforms at each hospital were cited as the primary causes hindering the model's generalization performance. Racial differences in patient populations and imbalances in detailed clinical stage distribution were also factors that increased prediction error. This demonstrates the persistent risk of overfitting, where the model becomes excessively tailored to the local characteristics of the training data during the forced integration of multimodal data. Significance and Outlook This study provides a clue for enhancing the predictive power of cancer immunotherapy response through the precise combination of multidimensional biological signals, while simultaneously revealing the technical bottlenecks hindering actual clinical application. It has been praised for highlighting the harsh reality in clinical practice, where algorithms developed in a single research environment are difficult to directly apply to datasets from other medical institutions. A prediction model excessively fitted to a specific dataset is difficult to function fully in a complex, real-world clinical environment. In future research, establishing normalization technologies that precisely correct for data discrepancies and batch effects between medical institutions is identified as the top priority. There is an opinion that a standard data collection specification compatible with medical sites worldwide must be established through multi-center clinical validation. The research paradigm is expected to shift toward developing robust Artificial Intelligence (AI) pipelines that maintain consistent predictive power even in heterogeneous environments, beyond merely increasing the complexity of algorithms.
💡 The multimodal biomarker analysis system holds the potential to evolve into a precision diagnostic panel for selecting candidates for high-cost ICI administration. A representative clinical application scenario involves screening out non-responsive patients prior to administration to preemptively prevent unnecessary immune toxicity side effects, thereby simultaneously reducing the economic burden on patients and healthcare expenditure. It can also contribute to establishing customized treatment strategies, such as suggesting combination therapies or alternative targeted therapies in a timely manner for patients with specific immune deficiency factors. However, to implement this in actual hospital clinics or as Software as a Medical Device (SaMD), standardization of specimen analysis protocols and verification of reproducibility across institutions must be supported.

Background During the COVID-19 pandemic, the large-scale emergency deployment of messenger RNA (mRNA) vaccines revealed structural limitations in existing pharmacovigilance systems. Traditional passive surveillance, which relies on voluntary reports from healthcare providers and vaccine recipients, suffers from structural flaws such as frequent reporting delays and high underreporting rates. The difficulty in precisely calculating the actual incidence of vaccine-related adverse events within the general population has also challenged health authorities. Whenever new vaccine platforms are introduced, surveillance channels are flooded with unstructured descriptive data and duplicate complaints. To early detect abnormal signs and clearly prove safety, a new approach was needed to rapidly integrate data from multiple sources. Machine Learning (ML) has emerged as a significant alternative for processing Real-World Data (RWD), encompassing vast Electronic Health Records (EHR), medical billing data, and adverse event reporting systems. Key Findings The researchers comprehensively collected English-language literature indexed in PubMed, Embase, and Web of Science from the inception of the database up to June 2026. Two reviewers independently reviewed the literature, and risk of bias assessment was based on a modified QUADAS-2 tool for assessing the quality of diagnostic accuracy studies. Due to the high level of heterogeneity across machine learning tasks (signal detection, text extraction, risk prediction, prognosis stratification), algorithms, and evaluation metrics, a narrative synthesis approach was inevitable. A total of 43 studies passed the strict criteria to be included in the final analysis. In the field of adverse event prediction, tree-based models in the decision tree family recorded stable values with an Area Under the Curve (AUC) between 0.85 and 0.87. However, it is difficult to determine direct superiority because the clinical settings and datasets used vary across studies. When Natural Language Processing (NLP) was applied to descriptive adverse event records in the vaccine reporting system, duplicate signals decreased by 17%. By filtering out unnecessary duplicate reports, the workload of surveillance personnel was reduced, while the speed at which meaningful safety signals are identified was significantly enhanced. In the case of myocarditis, cited as a major adverse reaction to mRNA vaccines, the prediction performance of ML models trained on cardiovascular cohorts reached an AUC of up to 0.899. However, studies directly validating these models in actual mRNA vaccine-vaccinated patient groups remain rare, and most are limited to retrospective analyses. For next-generation platforms such as self-amplifying mRNA vaccines or vaccines for cancer treatment, ML applications are identified as being in the proof-of-concept stage due to a lack of post-marketing real-world data. Implications and Outlook This systematic review clearly demonstrates that vaccine safety surveillance systems, which previously relied on post-event reporting, are evolving into active and intelligent real-time monitoring systems. Algorithms that learn from vast medical big data in real-time have the potential to detect subtle signs of rare side effects, which are difficult to identify during the early stages of vaccination. However, significant hurdles remain before integration into actual clinical practice. Discrepancies in data formats and quality across medical institutions, and the 'black box' problem where the derivation process of algorithms is difficult to explain clearly, remain obstacles to gaining trust in clinical settings. The current lack of specific licensing guidelines and certified validation standards from regulatory agencies is also cited as a factor delaying commercialization. Only by ensuring data transparency and establishing explainable artificial intelligence models will intelligent pharmacovigilance systems be positioned as a reliable shield for public health.
💡 The results of this study suggest specific directions for improvement in vaccine administration sites and health authority safety management systems. In medical institutions, high-risk groups for myocarditis can be precisely identified by pre-analyzing the underlying diseases, past drug reaction history, and cardiovascular risk factors of vaccine recipients using EHR-based ML models. This involves linking a Clinical Decision Support System (CDSS) that automatically provides intensive observation schedules and customized post-vaccination management guidelines to selected high-risk vaccine recipients. Health regulatory agencies can significantly reduce the administrative burden on dedicated personnel by introducing NLP algorithms into adverse event reporting systems to rapidly eliminate 17% of duplicate signals. This makes it a reality to establish a nationwide active surveillance network that captures statistical anomalies without missing them during the initial phase of new vaccine distribution.

Background The spread of multidrug-resistant bacteria (MDR), which do not respond to existing antibiotics, is a major health crisis facing modern medicine. Health authorities warn that annual deaths due to antibiotic resistance could reach 10 million by 2050. Conversely, the development of new antibiotics is stagnant due to massive costs and the rapid acquisition of resistance. In this deadlock, bacteriophages—viruses that selectively infect and kill specific bacteria—have emerged as an alternative. However, the academic consensus is that using natural bacteriophages as therapeutics has clear limitations. Natural phages have an excessively narrow host range and are easily neutralized by bacterial CRISPR-Cas defense systems. Some phages also cause side effects by remaining latent in the host genome and spreading toxins or resistance genes. Existing synthetic biology was limited to modifying parts of proteins. It was impossible for humans to individually reconstruct the entire regulatory network of a genome spanning tens of millions of base pairs (bp). This created a demand for technology capable of designing the entire viral life cycle from scratch. Key Findings The research team constructed a large-scale genomic language model to design bacteriophage genomes that target and destroy specific bacteria from scratch on a computer. This was achieved by the AI model deeply learning viral sequence data from hundreds of thousands of species, thereby autonomously mastering gene arrangement rules. Establishment of a genome generation system that comprehensively covers gene overlap structures and transcriptional regulatory factors at once. The resulting artificial phage genome has a size of approximately 42,000 base pairs, with sequence homology to natural phages being less than 60%. Despite being a completely new sequence, the endolysin gene, which decomposes capsid proteins and bacterial cell walls, is arranged according to precise rules. The researchers assembled the synthesized long-chain DNA in an in vitro expression system, injected it into bacterial hosts, and successfully resurrected actual virus particles with infectivity. The result of computer code manifesting as a physical virus. In vitro experiments showed that it completely lysed 28 out of 30 clinical strains of multidrug-resistant Pseudomonas aeruginosa. The target range was 2.6 times wider than that of natural phages. By optimizing the binding site, the frequency of bacterial resistance mutations was suppressed to one-tenth of the original level. Significance and Prospects This study is recognized as the first case in which a virus's entire genome was designed from scratch using artificial intelligence and expressed as a functional entity. It lays the technical groundwork for rapidly fabricating custom-designed viruses via computer commands. There is hope that the supply period for patient-specific phage therapy can be reduced from several months to within a few weeks. However, there are still significant challenges to overcome for its clinical adoption. The process cost of synthesizing and assembling long-chain DNA with tens of thousands of base pairs without error remains burdensome. It is difficult to rule out the risk that when artificial antigens are administered into the body, the patient's immune system recognizes them as foreign invaders and generates neutralizing antibodies. Proof of in vivo safety that does not harm beneficial microbiota must be established first. It is also urgent to establish a biosecurity verification system. The need for a software safety mechanism to block the creation of harmful viruses has been pointed out. Subsequent research is expected to focus on enhancing immune evasion capabilities and increasing DNA synthesis rates.
💡 A clinical scenario providing rapid, customized treatment options for patients with severe hospital-acquired infections, for whom existing antibiotics have become ineffective, could become a reality. Representative applications include chronic Pseudomonas aeruginosa pneumonia in cystic fibrosis patients or Staphylococcus aureus infections at artificial joint surgery sites. The workflow involves medical professionals analyzing the bacterial genome from a patient sample and inputting it, after which the AI model immediately designs an artificial phage genome capable of overcoming the receptor structure and defense mechanisms of that specific strain. When combined with automated synthesis platforms, an ultra-fast precision treatment pathway could be opened, completing and administering a target phage cocktail to a patient within 1 to 2 weeks. Beyond the medical field, analysis suggests high application value in the eco-friendly biological control market, such as suppressing antibiotic overuse in livestock and fisheries and removing pathogenic foodborne bacteria in food processing facilities.

Background Out of the 3 billion base pairs in the human genome, only about 1.5% directly encode proteins. The remaining 98% non-coding regions act as critical switches regulating gene expression, yet it has been virtually impossible to experimentally verify the impact of sequence variations on diseases. While attempts have been made to measure evolutionary conservation using computational algorithms, they have revealed limitations in failing to adequately reflect the surrounding nucleotide sequence context. Recently, Genomic Language Models (gLMs), which learn large-scale sequences like natural language, emerged as a promising alternative. However, most existing gLM models either analyze single-species sequences or fail to consider evolutionary distance when aligning multi-species sequences. Base pair matches between closely related species, such as humans and chimpanzees, acted as redundant data causing bias, while variations in distantly related species were often underestimated. This created a growing demand to directly integrate biological evolutionary phylogenies into learning algorithms. Key Findings A research team led by Professor Yun S. Lee from the Department of Computer Science at UC Berkeley Professor Song’s research team developed GPN-Star, a genomic model that combines phylogenetic tree distances with Whole-Genome Alignment (WGA) data, and published it in the international journal Nature. The core driver is a phylogeny-aware architecture that directly maps evolutionary divergence times between species into cross-attention operations, moving beyond previous methods that relied on simple list-based multiple alignments. By moving away from the practice of focusing only on the reference human sequence and instead applying random masking across the aligned sequences of multiple species, the density of learning signals was significantly increased. Analysis of the evolutionary lineages of 241 mammal species and 100 vertebrate species led to the precise quantification of functional constraints acting on all alleles within the genome. Performance metrics show a clear gap. In ClinVar variant identification experiments, GPN-Star significantly outperformed the existing conservation metric PhyloP and the ensemble score CADD. In the ClinVar clinical genetic variant database variant identification experiment, GPN-Star significantly outperformed the existing conservation metric PhyloP and the ensemble score CADD. It is evaluated that the accuracy of finding causal variants in genome-wide association studies (GWAS) in non-coding regulatory regions also surpassed that of previous state-of-the-art models. This demonstrates the flexibility to transcend species boundaries, extending to the analysis of five major model genomes including mice, chickens, fruit flies, Caenorhabditis elegans, and Arabidopsis. Significance and Outlook The emergence of GPN-Star provides a powerful guide for drug discovery researchers tracking the disease associations of non-coding genetic variants. Even when performing Whole-Genome Sequencing (WGS) for patients with rare intractable diseases where causes were not found in exome sequencing, the task of sifting through millions of non-coding variants for the true cause has been likened to finding a needle in a haystack. Implementing GPN-Star, which precisely calculates evolutionary pressure, is expected to increase the efficiency of compressing potential candidates several times over. The researchers' decision to release the calculated scores as UCSC Genome Browser tracks and to fully open the model weights is also a factor injecting vitality into the ecosystem. This opens a way for small and medium-sized biotechs without high-performance computing equipment to immediately determine variant risk with just a few clicks. However, practical challenges remain. Prediction reliability tends to fluctuate in repetitive sequences or structural variation regions where inter-species base alignment is difficult. Furthermore, because it relies on base information preserved by evolutionary history, it has a clear limitation in not being able to capture real-time epigenetic expression dynamics occurring in specific cells or tissues. This is why integration with single-cell chromatin accessibility data has emerged as a key next-generation challenge.
💡 In clinical genomic diagnostics, it is expected to be immediately implemented in interpretation pipelines to confirm non-coding pathogenic variants in patients with rare genetic diseases of unknown etiology. When performing whole-genome analysis on patients who failed diagnosis with standard exome testing, applying GPN-Star scores allows for the rapid filtering of pathogenic candidates hidden deep within promoters or introns from tens of thousands of Variants of Uncertain Significance (VUS). In the pharmaceutical and biotech industries, it will serve as a decisive clue for developing next-generation gene therapies targeting non-coding regions. If key enhancers or silencers regulating disease gene expression are accurately identified, a framework for the much more precise design of antisense oligonucleotide (ASO) or gene-editing-based transcriptional regulation therapies will be established. The agricultural and livestock biotech sectors will also benefit practically by shortening breeding cycles through the rapid selection of beneficial variants that determine crop disease resistance or livestock productivity.

Background With advancements in human genome decoding technology, the cost of Whole Genome Sequencing (WGS) has significantly decreased, yet more than half of rare disease patients still undergo diagnostic odysseys without identifying a genetic cause. Until now, genetic research and clinical diagnosis have focused on the exome region, which directly encodes proteins. The exome accounts for only about 1.5% to 2% of the entire genome. It has been presumed that numerous disease-causing variants exist in the vast noncoding region, which makes up the remaining 98%, but cases identified as actual causes have been extremely rare. The clinical interpretation of noncoding region variants has been suboptimal due to the absence of clear decoding rules. In protein-coding sequences, functional losses such as those caused by amino acid substitutions or premature termination can be predicted relatively easily based on the codon system. In contrast, it is difficult to intuitively gauge the impact of mutations on phenotypes even when regulatory sequences are altered. The 3D chromatin structure, in which regulatory elements such as enhancers or promoters control genes located hundreds of thousands of base pairs away, also increases the difficulty of analysis. In essence, there was a lack of verification tools to distinguish lethal, disease-causing variants from harmless benign variants within the vast noncoding sequences. Key Findings This review published in Nature Genetics summarizes the molecular mechanisms by which noncoding variants cause rare diseases and highlights the fundamental reasons for the low diagnostic discovery rate. The authors analyzed that noncoding regulatory abnormalities primarily induce quantitative changes in gene expression levels, making them act more tissue-specifically and subtly than complete protein loss. They pointed out that existing methods, which relied solely on evolutionary conservation analysis, struggled to capture regulatory variants that are activated only in specific cellular environments. As a key solution to accelerate the discovery of noncoding variants, the researchers presented a combination of Massively Parallel Reporter Assay (MPRA), high-resolution chromatin structural analysis (Micro-C), and machine learning-based prediction algorithms. By utilizing Saturation Genome Editing (SGE), all single-nucleotide variants in noncoding sequences can be functionally evaluated simultaneously in a laboratory setting. The trend is for functional prediction accuracy to improve dramatically, as deep learning sequence models that directly interpret sequence information to quantify chromatin accessibility and transcriptional activity by cell type are combined. It is assessed that the screening efficiency for narrowing down actual disease-causing variants from tens of thousands of candidate variants has increased by dozens of times compared to the past. Significance and Outlook The precise identification of noncoding regulatory networks is expected to be a turning point in increasing the diagnosis rate of rare genetic diseases. This is because it can provide a clear molecular etiology for patients who missed the appropriate treatment window due to unknown causes. Accumulated noncoding variant data provides a blueprint for developing precision therapeutics that directly target noncoding regulatory sequences, such as antisense oligonucleotides (ASOs) or epigenome editing. Challenges also remain clear. To bridge the gap between the predictions of computer algorithms and the actual phenotypes in patients, precise comparison with single-cell transcriptome maps across different organs and developmental stages must follow. Establishing a high-speed functional verification system using disease-model organoids and standardizing clinical interpretation guidelines for noncoding variants are also identified as tasks to be solved. The trend is for genome data interpretation capabilities to expand significantly beyond the boundaries of coding regions into noncoding regions.
💡 The approach presented in this study can significantly improve the diagnosis rate in undiagnosed rare disease clinics where causes remain unidentified even after next-generation sequencing. In clinical settings, by linking a patient's WGS data with noncoding functional prediction AI models and single-cell multi-omics data, a system will be established to prioritize transcriptional regulatory site variants directly linked to the disease, shortening interpretation time from months to days. From the perspective of the drug development industry, this provides an opportunity to identify a large number of new regulatory factor targets for intractable diseases where drug development was impossible due to the lack of protein targets. It is expected that the design of targeted RNA therapeutics, which inhibit disease-causing enhancers or selectively restore silenced gene expression, will become a reality.

Background Despite rapid advancements in genome analysis technology, human genomics has faced a long-standing challenge. Although it has been over 20 years since the Human Genome Project decoded 3 billion base pairs, the regions actually encoding proteins account for only about 1.5% to 2% of the entire sequence. The vast non-coding regions, comprising the remaining 98%, act as switches to finely regulate gene expression, but their functional map remains uncharted territory. Previous studies, including Genome-Wide Association Studies (GWAS), have succeeded in cataloging numerous genetic variants associated with complex diseases. However, since more than 90% of discovered variants are concentrated in non-coding regulatory regions, it has not been easy to prove the causal relationship between specific base changes and disease induction at the molecular level. High-throughput screening experiments that introduce mutations into cell lines also face limitations, such as enormous costs and the difficulty of fully reproducing the complex chromatin structure in vivo. While AlphaFold and AlphaMissense have advanced missense variant analysis in the field of protein structure prediction, a computational model capable of elucidating the vast regulatory network of non-coding regions has long been missing. Key Discovery To fill this gap, researchers at Google DeepMind introduced 'AlphaGenome,' an AI-based sequence model covering the entire human genome. The researchers succeeded in predicting the biological impact of a total of 9 billion single nucleotide variants (SNVs) by exhaustively analyzing the three possible substitution cases for each of the approximately 3 billion base pairs constituting the human genome. This means they have calculated, at single-nucleotide resolution, how fluctuations in a single base affect gene expression levels, chromatin accessibility, transcription factor binding affinity, and splicing patterns. The model architecture is designed to simultaneously compute long genomic contexts spanning hundreds of thousands of base pairs to capture interactions between distant regulatory elements. The researchers quantified the ripple effects on distant regulatory elements, such as promoters and enhancers, by training on whole-genome data and large-scale functional genomics experimental results. Benchmark evaluation results showed that the accuracy of pathogenicity prediction for non-coding variants improved by more than two orders of magnitude compared to existing statistical genetics tools and machine learning models. While existing models often missed new regulatory variants by relying on conserved sequence ratios, AlphaGenome distinguishes itself by reading the three-dimensional binding changes of regulatory factors directly from the sequence information itself. Significance and Outlook This achievement marks a new turning point for research tracing the pathogenesis of rare intractable genetic diseases, cancer, and complex diseases such as diabetes caused by non-coding variants. By resolving the bottleneck in variant interpretation that was difficult to verify individually in the laboratory, a way has opened to pre-identify the pathogenicity of numerous non-coding region variants that remained as Variants of Uncertain Significance (VUS) in clinical genomic sequencing. For researchers handling large-scale genomic cohorts, it provides a powerful selection criterion to narrow down the pool of analysis candidates. However, clear practical limitations of computational predictions also exist. Since the regulatory variant effects calculated by AI can manifest differently depending on the actual in vivo tissue environment, developmental stage, or epigenetic state, precise experimental verification must necessarily follow. Follow-up work is also required to augment genomic data from diverse racial groups to avoid bias toward specific population data. It is expected that, when combined with single-cell analysis technology, computational biology will establish itself as a key driver for accelerating drug target discovery.
💡 The functional map of 9 billion variants presented by AlphaGenome advances the timeline for drug development and precision medicine. In clinical settings, it will enable rapid molecular diagnosis for patients carrying non-coding variants that were previously classified as having unknown function after Next-Generation Sequencing (NGS) testing. A representative clinical scenario is establishing a treatment strategy on the day of testing by predicting that a non-coding promoter mutation in a patient suspected of having a rare genetic disease suppresses specific gene expression. In the biopharmaceutical industry, it reduces trial and error in the stages of target discovery and gene therapy design. When designing antisense oligonucleotides (ASOs) or CRISPR-based corrective therapies targeting specific non-coding regions that cause disease, it helps prevent off-target mutations in advance. It is expected to provide practical benefits by significantly reducing the repetitive cell experiment processes in the candidate substance derivation stage, thereby saving drug development costs.

Background Human African Trypanosomiasis (HAT), transmitted by the tsetse fly, is a fatal endemic disease caused by the protozoan Trypanosoma brucei, leading to central nervous system paralysis and sleep disorders. While initial infection presents as simple fever and headache, the parasite's penetration of the blood-brain barrier results in severe neurological damage. Current commercial treatments suffer from low administration convenience and risks of neurotoxicity, and no FDA-approved preventive vaccine exists. The primary obstacle to vaccine development lies in the pathogen's unique immune evasion strategy. The Variant Surface Glycoprotein (VSG) covering the parasite surface continuously switches its antigenicity, often neutralizing the host immune system's ability to form memory against a single antigen. Traditional attenuated live vaccines or protein subunit methods have struggled to overcome this extreme antigenic variation. Consequently, alternative vaccine designs combining Reverse Vaccinology—utilizing genome and protein structural data analysis—with the messenger RNA (mRNA) platform, which allows for rapid antigen combination, have begun to gain attention. Key Findings The research team targeted three proteins essential for persistent infection and survival in T. brucei: VSG, the central axis of antigenic variation; Heat Shock Protein 70 (HSP70), which assists in stress response and protein folding; and the Vacuolar Transporter Chaperone (VTC) complex, responsible for membrane transport and vesicle movement. We derived optimal peptide epitopes with high immunogenicity while excluding allergenic potential and cytotoxicity, after passing through a bioinformatic screening pipeline. The selected epitope combinations cover major histocompatibility complex (MHC) alleles across global population groups, achieving 100% global population coverage. Physicochemical evaluations showed an Aliphatic index of 71.23, ensuring structural thermal stability, and a Grand Average of Hydropathy (GRAVY) score of -0.719, indicating excellent solubility. The predicted tertiary structure of the vaccine protein recorded a TM-score of 0.65 ± 0.13 and a C-score of -0.50; following structural refinement, it demonstrated thermodynamic stability with a Ramachandran plot favored region ratio of 86.8% and a Z-score of -5.26. The ability to activate the immune system was also demonstrated at the molecular level. Molecular docking with Toll-like receptor 2 (TLR-2) and Toll-like receptor 4 (TLR-4), key mediators of innate immune recognition, formed very low binding energies of -1013.5 kJ/mol and -1002.8 kJ/mol, respectively. Binding stability was further confirmed via Molecular Dynamics (MD) simulation, Principal Component Analysis (PCA), Dynamic Cross-Correlation Matrix (DCCM), and Molecular Mechanics Generalized Born Surface Area (MM-GBSA) calculations. Furthermore, to maximize the translation efficiency of the mRNA molecule, codon optimization was performed for the Escherichia coli expression system. A Codon Adaptation Index (CAI) of 0.9688 and a GC content of 44.70% were secured, and the structural integrity of the transcript itself was confirmed through Minimum Free Energy (MFE) analysis. Immune simulations observed the proliferation of activated B lymphocytes and T lymphocytes, along with high Immunoglobulin M (IgM) and Immunoglobulin G (IgG) antibody titers. Significance and Outlook This study is significant in that it completed a blueprint for a multi-epitope vaccine designed to strike intracellular and surface membrane proteins from multiple angles, effectively turning the pathogen's complex antigenic variation against itself. In particular, by going beyond the limitations of targeting VSG alone and including the essential intracellular survival factors HSP70 and VTC, the study proposes a design that fundamentally blocks vaccine evasion pathways caused by antigenic variation. As the structural stability and immunogenic potential have been clearly identified through computer modeling, this could serve as a catalyst to shorten the research period for neglected tropical disease vaccine development. However, as these are in-silico results based on computer calculations, the critical hurdle of actual in-vivo efficacy verification remains. Prior evaluation of delivery efficiency upon loading the designed mRNA into lipid nanoparticles (LNPs), intracellular translation expression levels, and infection protection efficacy in animal models must be conducted. Monitoring long-term immune responses in response to the potential emergence of pathogen variants is also identified as a task to be addressed in the future.
💡 This vaccine candidate presents a practical alternative to improve health security in sub-Saharan Africa, the habitat of the tsetse fly. Existing chemotherapeutic agents require long-term hospitalization and intravenous administration, making them difficult to apply in remote villages with poor medical infrastructure. If this multi-epitope mRNA vaccine is commercialized in combination with freeze-drying technology or thermostable formulations, it could become a public health weapon to block disease spread through mass vaccination. Furthermore, since it was designed with 100% population coverage in mind, its potential for wide application regardless of racial background increases the possibility of collaboration with global vaccine production companies.

Background David Botstein, who passed away on February 27, 2026, at the age of 83, was a geneticist who did not merely explain specific biological phenomena but designed tools to interpret complex biological systems. A commemorative article published in PNAS in September 2026 highlights his academic journey, from studying the Salmonella phage P22 to mapping the human genome, developing DNA microarrays, and advancing quantitative biology education. During the 1960s and 1970s, when Botstein began his research, there were no reliable methods to locate disease-related genes within the human genome. DNA sequencing technology was in its infancy, and there were few markers to link phenotypes to causal genes. He used conditional lethal and suppressor mutations to group genes involved in the same function and estimate their order of action within pathways. This approach expanded genetics from the study of individual genes to the analysis of interactions and systems. Key Discoveries His most widely recognized achievement is the 1980 proposal, along with Ronald Davis and others, of a human gene linkage map based on restriction fragment length polymorphisms (RFLPs). The concept was to use DNA fragments that varied between individuals as chromosomal markers and calculate their co-inheritance frequency with disease phenotypes within families to narrow down the location of causal genes. This reversed the traditional approach, which required knowing the gene itself before determining its location. The RFLP map accelerated the identification of genes responsible for hereditary diseases such as Huntington's disease and laid the conceptual foundation for BRCA1 discovery and the Human Genome Project. It effectively transformed human chromosomes into searchable maps using recombination frequencies as coordinates, even before genome sequencing was available. For this contribution, Botstein was among the 11 inaugural recipients of the Breakthrough Prize in Life Sciences in 2013. Princeton University described him as a key leader of the Human Genome Project and a pioneer in gene mapping methods. His research extended from gene location to expression patterns. Collaborating with Patrick Brown's team, he advanced DNA microarray technology, clustering expression levels of thousands of genes to distinguish functionally similar genes and tumor subtypes. He conducted genome-wide analyses of yeast cell cycles and metabolic responses and contributed to classifying diffuse large B-cell lymphoma and breast cancer based on molecular features. His work was notable for including not only experimental equipment but also algorithms and visualization systems for interpreting large-scale data. Significance and Outlook Botstein's legacy lies less in the RFLP technology itself and more in the research approach of 'creating measurable markers and reconstructing biological structures through computation.' The common principles underlying today's genome-wide association studies, single-cell transcriptome analysis, and cancer molecular diagnostics trace back to this lineage. Transferring concepts established in bacteria and yeast to human diseases also helped bridge the gap between model organism research and clinical genomics. His career, which spanned academia and industry, further broadened his influence. He conducted research at MIT and Stanford University and held leadership roles at Genentech and Calico as Chief Scientific Officer. From 2003 to 2013, he directed Princeton University's Lewis-Sigler Institute for Integrative Genomics, integrating mathematics, physics, and computer science into biology education. However, RFLP and early microarrays were limited in resolution and measurement range and have largely been replaced by high-throughput sequencing. While the technology has evolved, the principle of developing tools aligned with hypotheses and interpreting data at the systems level remains embedded in the design of precision medicine and systems biology research.
💡 The RFLP linkage map is the starting point for genetic cancer panels and rare disease family testing performed in hospitals today. The diagnostic workflow of narrowing down candidate regions by analyzing the co-inheritance of mutations and diseases among family members and then confirming causal mutations through sequencing is a direct extension of this approach. The tumor expression classification methods established through microarrays have evolved into multi-gene tests that assist treatment decisions, such as predicting breast cancer recurrence risk or distinguishing lymphoma subtypes. Pharmaceutical companies can apply the same principles to patient cohort selection and biomarker discovery. However, drug responses cannot be reliably determined by expression signals alone, and clinical validity must be separately validated across diverse population groups.

Background Spatial omics measures gene expression, proteins, chromatin accessibility, histone modifications, metabolites, and imaging data at specific coordinates within tissues. Overlaying multiple molecular layers from the same region allows for a three-dimensional understanding of cell states and tissue architecture, but differences in signal distribution and noise levels across modalities often lead to loss of unique signals during integration. Existing single-cell integration methods such as Seurat, MOFA+, and MultiVI do not fully account for spatial neighbor relationships. The spatial multimodal tool SpatialGlue is limited to a maximum of three molecular layers and incurs high computational costs when handling high-resolution imaging or datasets with millions of coordinates. Forcing dissimilar data into alignment can flatten biological differences, a problem known as 'over-alignment.' Additionally, these methods lack the ability to identify and flag unstable or unreliable locations in the analysis. Key Discovery SCIGMA, developed by researchers at Brown University, is an unsupervised deep learning model that combines a multi-view graph neural network with uncertainty-aware contrastive learning. The model merges spatial graphs connecting physically close locations and feature graphs linking locations with similar molecular profiles. A graph attention network (GAT) generates latent representations for each omics layer, and cross-attention integrates these representations into a shared space. A decoder reconstructs the original features to minimize information loss. A key innovation is the uncertainty-aware contrastive loss, which learns location-specific temperature parameters. This ensures that shared representations of the same coordinate are close to modality-specific representations without uniformly mixing signals unique to each omics layer. Coordinates with poor alignment are assigned high uncertainty, allowing researchers to identify complex regions such as tumor heterogeneity or immune cell compartments, as well as areas suspected of technical noise. The research team evaluated the model across 19 datasets from 10 tissue types and 9 platforms, including eight data types such as gene expression, protein, chromatin accessibility, histone modification, and metabolites. In comparable mouse brain spatial ATAC–RNA and spleen SPOTS data, SCIGMA outperformed Seurat v5, MOFA+ 1.13.0, MultiVI v1, and SpatialGlue v1 in spatial domain detection, modality-specific signal preservation, feature reconstruction, and reproducibility across repeated runs. Other methods failed to produce results in some large datasets even with over 400GB of CPU memory and two days of processing time, according to the researchers. SCIGMA processed over a million spatial locations, including 2-micrometer resolution data from 10x Visium HD, using graph sampling. In a 5-month-old mouse brain Spatial-Mux-seq dataset, SCIGMA integrated RNA, protein, ATAC, H3K27ac, and H3K27me3 data simultaneously. Previously indistinct anatomical regions such as the cerebral cortex layers, hippocampal dentate gyrus, caudate putamen, and striatum were clearly separated in the shared representation. The Nature Genetics paper notes that this framework is extensible to future platforms beyond five modalities. Implications and Outlook SCIGMA addresses three major challenges in spatial multimodal analysis—preservation of modality-specific signals, computational scalability for large datasets, and uncertainty quantification—within a single framework. The ability to assign uncertainty to individual coordinates allows researchers to selectively validate clusters that may require pathological review or further experimentation. The software and reproducible analysis code are publicly available. However, uncertainty scores reflect a mixture of biological heterogeneity and technical variability. High uncertainty alone cannot confirm new cell states or disease boundaries; tissue staining and independent marker validation are still required. Performance advantages are also limited to datasets with annotated ground truth or where competing methods are computationally feasible. Future validation is needed to assess how well SCIGMA handles differences in clinical sample layouts, missing modalities, and varying resolutions.
💡 In hospitals, SCIGMA can be used to integrate RNA, protein, and pathology imaging data from tumor sections to more precisely delineate cancer cell compartments, immune cell infiltration areas, and stromal boundaries. For example, in ovarian cancer tissue, regions with high uncertainty can be prioritized for re-examination, and the expression of immune checkpoint proteins and genes at those coordinates can be compared to narrow down potential therapeutic targets. Pharmaceutical companies can incorporate SCIGMA into analysis pipelines to compare spatial signal changes before and after drug treatment, identifying resistant microenvironments or toxicity hotspots. However, SCIGMA is currently a research computational tool. For clinical use in patient treatment decisions, standardized sample processing, external cohort validation, and clinical calibration of uncertainty thresholds must be established first.

Background The human cerebral cortex rapidly expands during fetal development as neural stem cells proliferate and differentiate into multiple lineages. At the core of this process are ventricular radial glia (vRG) and outer radial glia (oRG). Particularly in primates, the significantly increased oRG is considered a key cell type responsible for the uniquely broad and gyrencephalic human cortex, although the cell-specific gene regulatory mechanisms remain poorly understood. Previous single-cell studies have revealed which genes are expressed and which chromatin regions are open, but have struggled to precisely trace long-range regulatory contacts in 3D space. This limits the interpretation of disease-associated variants, which often regulate genes tens of kilobases away rather than the nearest gene. Researchers isolated vRG, oRG, oligodendrocyte precursor cells (OPC), and microglia (MG) from the human cerebral cortex at 15–24 weeks of gestation and performed an integrated analysis of gene expression, chromatin accessibility, DNA methylation, and 3D chromatin interactions. The different developmental stages were chosen because the peak appearances of RG and OPC·MG differ. Nature paper Key Findings The team identified 69,141 candidate cis-regulatory elements (cCREs) in vRG, 72,450 in oRG, 65,295 in OPC, and 69,508 in MG with high accessibility and low methylation. Over 60% of these were located outside promoters. Using promoter-linked chromatin interaction analysis by sequencing (PLAC-seq) with H3K4me3, they captured 135,000–144,000 significant interactions per cell type at 2-kilobase resolution, detecting approximately four times more contacts than the previous 5-kilobase analysis. The average interaction distance ranged from 188,000 to 233,000 base pairs. In mouse embryo reporter experiments, 14 out of 18 RG and intermediate progenitor regulatory sequences showed enhancer activity in the embryonic brain. Comparing vRG and oRG revealed 7,941 differentially accessible regions but only 756 differentially methylated regions, indicating that differences between the two cells were more pronounced in chromatin accessibility and 3D contacts. Transcription factor LHX2 was linked to the oRG regulatory network, while ASCL1 was associated with vRG. When LHX2 was inhibited using short hairpin RNA, oRG self-renewal decreased and oligodendrocyte-like differentiation increased. Machine learning analysis also narrowed the scope of disease variants. Among 11,360 schizophrenia-associated variants identified by DeepGWAS, 929 were located in open chromatin, and 112 were classified as candidates for altering accessibility. The risk allele T of rs4449074 was predicted to reduce the activity of the vRG enhancer hs3134, and in mouse forebrain, it showed weaker signals than the non-risk allele C. Alzheimer’s disease risk was enriched only in MG regulatory elements, while autism spectrum disorder risk was prominent in vRG and oRG. Implications and Outlook Evolutionary analysis showed that 72 of the differentially accessible regions in oRG overlapped with human accelerated regions (HAR), compared to only 13 in vRG, indicating a 2.43-fold higher concentration in oRG. The team evaluated 565 open HARs and 3,447 human-chimpanzee sequence differences in oRG, identifying 76 candidates where human sequences increased accessibility and 73 where chimpanzee sequences remained more open. HARsv21313 was connected to the promoter of ROCK2 through 3D contact. Inhibiting this region using CRISPR interference (CRISPRi) in induced pluripotent stem cell-derived neural progenitors reduced ROCK2 expression and increased Ki67-positive cells. Additionally, the human-specific sequence of HARsv21602 lost a binding site for the repressive transcription factor ZBT18, potentially enhancing EPHA4-related regulatory activity. This suggests that human cortical expansion and neuropsychiatric risk may share some non-coding regulatory circuits. However, the analysis values represent averages of cell populations isolated by fluorescence-activated cell sorting. Rare progenitor subtypes and spatial heterogeneity were not captured, and the accessibility differences among rs4449074 carriers were not statistically significant. Results from mouse embryos and cultured neural progenitors should not be directly extrapolated to human fetal development or patient pathology. Spatial and single-cell epigenomic analyses, along with precise editing experiments in human brain organoids, are needed to confirm causal relationships.
💡 This map can serve as a selective tool to connect non-coding variants identified in genome-wide association studies to the actual operating cells and target genes. For example, if a schizophrenia variant reduces vRG enhancer activity and interacts with specific developmental genes, pharmaceutical companies could evaluate candidate compounds using human cortical organoids that replicate the regulatory axis. Alzheimer’s disease variants could be prioritized for MG, while autism and schizophrenia variants could be assigned to RG and OPC, reducing the cost of cell model selection and functional validation. However, the data alone are not sufficient to determine diagnostic risk scores or therapeutic targets, and reproducibility must be confirmed in more donors with diverse genetic backgrounds.

Background Genetic prediction models that estimate the impact of genetic variants on RNA expression levels, protein concentrations, and metabolite levels have become essential tools for elucidating disease mechanisms and discovering new drug targets. Transcriptome-wide association studies (TWAS) and proteome-wide association studies (PWAS) link diseases to molecular traits by indirectly estimating them from genome-wide association study (GWAS) data. These approaches allow researchers to explore the pathway from genetic variants to molecular traits and diseases without directly sampling patient tissues. The challenge has been that these models are scattered across paper appendices, lab servers, and individual repositories like PredictDB. Variant coordinates, effect alleles, and weight formats vary, and information on training population ancestry, measured tissues, or platforms is often missing. This inconsistency made it difficult to apply the same model to independent cohorts or compare across omics layers. In particular, models developed in European-ancestry populations may yield different association power and effect estimates when applied to other ancestry groups. A research team from the University of Cambridge addressed this fragmentation by building OmicsPred, an open platform for registering, searching, and distributing genetic prediction models for multi-omics traits. The study was published in Nature Genetics on September 1, 2026. Key Findings OmicsPred assigns each predictive score an OPGS identifier and provides standardized formats for chromosomal positions, rsIDs, effect alleles, and variant-specific weights. It also links to the original genome build, molecular traits, measured tissues and platforms, model development methods, training and validation sample sizes and ancestry compositions, performance metrics, and usage conditions. Gene, transcriptome, proteome, and metabolome data are integrated with Ensembl, UniProt, and ChEBI, while tissues and phenotypes are normalized using established ontologies. The file format is designed based on the multi-polygenic score catalog (PGS Catalog) standard. Users can download models in pgsc_calc-compatible score files or PredictDB formats executable in MetaXcan. A REST application programming interface (API) is also provided for querying metadata within the platform. The implementation code for the platform and database is publicly available under the Apache 2.0 license. The research team validated the platform's utility using disease GWAS data from the U.S. Million Veteran Program (MVP). They analyzed 9 datasets of blood- and plasma-based transcriptome and proteome predictive scores, totaling 38,450 scores, using S-PrediXcan. In European-ancestry data, 1,233 PheCode disease traits were evaluated, 986 in African-ancestry data, and 719 in mixed-ancestry data. Only GWAS with ancestry and training cohorts that were similar were matched. Scores with less than 75% overlap with GWAS variants were excluded, and multiple testing was corrected per dataset. The analysis identified 194,138 false discovery rate-corrected significant associations. The association between CFH expression and age-related macular degeneration reached a P-value of 1.29×10⁻⁶¹, and that between PCSK9 expression and coronary artery disease reached 7.61×10⁻²⁵. Thirty-one gene-disease combinations were consistently replicated across all six European-ancestry datasets, and five combinations were commonly detected across ancestry groups in the SomaLogic proteome models. Implications and Outlook The value of OmicsPred lies in transforming existing models into a reusable research infrastructure, rather than introducing new prediction algorithms. Comparing predictions for the same gene or protein across tissues, measurement platforms, omics layers, and ancestry groups can help prioritize candidates repeatedly linked to specific diseases. The strong replication of known associations such as CFH and PCSK9 demonstrates the platform's ability to recover disease-related signals. The open API and standardized metadata are also useful for automating large-scale target discovery pipelines. Pharmaceutical companies can evaluate associations between genetically predicted molecular traits and hundreds of diseases in bulk, assessing not only potential indications but also possible adverse signals. If a candidate is consistently identified across multiple omics layers, it can strengthen the rationale for prioritizing follow-up functional experiments. However, associations between genetic predictions and diseases alone do not prove causality or drug efficacy. Linkage disequilibrium, horizontal pleiotropy, tissue specificity, and measurement platform differences can distort results. The continued dominance of European-ancestry models and data remains a limitation. To identify candidates with high potential for therapeutic targets, more models developed and validated in diverse ancestry groups and cell types are needed, combined with colocalization analysis, Mendelian randomization, and experimental validation.
💡 In drug development, researchers can compare disease GWAS with OmicsPred scores to prioritize RNA or protein candidates regulated by genetics. For example, in a cardiovascular disease program, candidates consistently identified across multiple transcriptome and proteome datasets can be passed to colocalization analysis and cell experiments, and their associations with other diseases can be checked to review indications and safety risks early. Hospital and biobank researchers can apply standardized score files to their own genomic data to estimate molecular traits in large patient populations where direct tissue measurements are difficult. However, to use these as clinical decision-making tools, separate validation in the target population and quantification of prediction accuracy and ancestry-specific biases are necessary.

Background Microbial symbiosis plays a wide range of roles, from the host's nutrient acquisition and stress resistance, disease defense, to carbon and nitrogen cycling in ecosystems. However, many bacteria and archaea cannot be cultured in the laboratory, making it difficult to confirm how they survive and with whom. Known symbiont genomes have been biased toward species isolated from humans or animals and those that are easy to culture, making it difficult to estimate the scale of symbiotic relationships in real ecosystems. Symbionts tend to lose unnecessary metabolic genes and genomic regions as their dependency on the host increases. However, it is not possible to determine symbiosis solely based on small genomes. Free-living microbes can also reduce their genomes during environmental adaptation, and species loosely associated with hosts tend to retain relatively complete metabolic capabilities. The research team believed that a machine learning system that reads the overall functional composition of genomes, rather than individual genes, could distinguish between free-living, host-associated, and obligate intracellular types, thereby overcoming these limitations. Key Findings The research team developed a symbiont prediction framework called 'symclatron.' First, they inferred orthologous gene groups from 792 symbiont proteomes selectively sampled from major clades, and created 26,300 profile hidden Markov models in groups with five or more members. They then used 6,751 microbial genomes with confirmed lifestyles, categorized into 5,959 free-living, 409 host-associated, and 383 obligate intracellular, as training data. The classifier 'symcla' and the regression model 'symreg,' which represents host dependency as a continuous value, each used 1,000 selected influential features. The final neural network integrated seven factors, including the predicted values of the two models, genome completeness, and distance from training clades. The research team also performed clade-level cross-validation by sequentially removing 11,025 clades from the training data, a design that not only tests whether it can correctly identify close relatives of known species but also whether it can be applied to unseen clades. The F1 scores for free-living, host-associated, and obligate intracellular types in the validation data were 0.982, 0.699, and 0.908, respectively. It accurately classified 97.6% of obligate intracellular symbionts, with 5.2% of host-associated types misclassified as obligate intracellular and 2.3% of obligate intracellular types misclassified as host-associated. This also revealed the biological characteristic that the boundary between the two symbiotic categories is continuous. The symclatron was applied to metagenome-assembled genomes (MAGs) and reference genomes recovered from 31,152 environmental sequencing projects. After quality control and deduplication based on an average sequence identity threshold of 95%, 107,067 bacterial and archaeal genomes were obtained, and 14,070 were classified as symbionts at a confidence threshold of 0.725, accounting for about 14% of the total, or roughly one in seven. The proportion of symbionts was 15–23% when considering only MAGs, and candidates were found in half of the known bacterial and archaeal phyla. These genomes were published as the 'Symbiont Genomes (SymGs)' catalog. Significance and Prospects This study opens the way to large-scale screening of symbiotic potential based solely on functional genomic traces, without the need for cultivation or host observation. The model classified the lineage of 'Candidatus Azoamicus ciliaticola,' which supplies energy to a flagellate host, and the newly reported archaeon 'Candidatus Sukunaarchaeum mirabile' as obligate intracellular symbionts, indicating that the model can explore dependencies not only among eukaryotes but also between bacteria and archaea. Interpretation requires caution. The training data were biased toward bacteria associated with eukaryotic hosts, and archaeal cases were limited. Performance decreased in the most phylogenetically distant groups, and the F1 score for host-associated types was also lower than for the other two categories. It is also not possible to determine the interaction partner or whether the symbiosis is mutualistic or parasitic based on predictions alone. SymGs is closer to a candidate map for prioritizing follow-up cultivation, microscopic observation, single-cell analysis, and metabolic experiments, rather than a finalized list of verified symbionts.
💡 By inputting microbial genomes obtained from environmental samples into symclatron, candidates with high potential for host dependency can be prioritized. For example, in plant root zones, candidates related to nutrient uptake or pathogen suppression can be selected for use in designing synthetic microbial communities, or in marine samples, symbiotic pairs mediating carbon and nitrogen cycling can be tracked with focus. In human and livestock microbiomes, the metabolic deficiencies of host-dependent microbes that are difficult to culture can be predicted, aiding in the design of customized media and co-cultivation conditions. However, actual host verification, functional validation, and biosafety evaluation are required for industrial strains or therapeutic targets.

Background Bacteria of the genus Leptospira are pathogens that cause leptospirosis in humans and livestock. They colonize the kidneys of various animals and are excreted in urine, contaminating the environment, but the genes responsible for pathogenicity and host adaptation remain poorly characterized. This is largely due to the difficulty in generating mutant strains to directly confirm gene function. In conventional bacterial genome editing, Cas9 cuts both strands of the target DNA, and the cell's repair process induces gene disruption. However, most Leptospira cannot properly repair double-strand breaks (DSBs). Even when Cas9 accurately cuts the target, edited cells often die, making it difficult to recover knockout strains. CRISPR interference using catalytically inactive Cas9 can suppress gene expression, but it does not permanently alter the DNA sequence itself. To overcome this limitation, the research team developed a dual CRISPR system combining CRISPR/Cas9-NHEJ and CRISPR prime editing (PE), which operate on different principles. The former provides DSB repair capability externally, while the latter avoids lethal DSBs altogether. Key Findings In the CRISPR/Cas9-NHEJ system, Cas9 and guide RNA are coexpressed with DNA repair proteins LigD and Ku from Mycobacterium smegmatis. Ku captures the cut DNA ends, and LigD connects them, thereby supplementing the deficient NHEJ function in Leptospira. This repair process is error-prone, resulting in insertions or deletions (indels) at the target site, which can disrupt the gene reading frame and generate knockout mutants. PE is a more precise approach that reduces the risks of cutting and repairing DNA. The researchers fused a Cas9 nickase with a reverse transcriptase. A prime editing guide RNA directs the system to the target, and the reverse transcriptase records the desired sequence into new DNA. Since only one DNA strand is cut, this method preserves cell viability in DSB-sensitive Leptospira. The two systems serve distinct roles. Cas9-NHEJ is well-suited for rapidly generating diverse indels through error-prone repair. In contrast, PE is ideal for experiments requiring precise, pre-designed outcomes, such as single-base substitutions. Unlike previous studies that relied on random mutagenesis or transient gene suppression, this new toolkit allows researchers to selectively choose between gene disruption and precise correction based on their experimental goals. Significance and Outlook The new editing systems provide a robust foundation for causally validating pathogenic factors, metabolic pathways, environmental survival, and host colonization genes in Leptospira. Comparing strains with complete gene deletions and those with specific base changes allows the effects of protein loss and individual amino acid alterations to be analyzed separately. The systems also offer potential for studying functional differences among various species and serovars. In vaccine development, the systems could be used to create attenuated live vaccines by removing pathogenic genes or to engineer strains with modulated antigen production. However, the current achievements focus on expanding the editing principles and tools. Factors such as editing efficiency per target, off-target mutations, and strain-specific plasmid delivery and expression require further validation. Indels generated by NHEJ are variable, necessitating individual sequence confirmation for candidate strains, and PE performance can also vary depending on the target sequence and guide design. To assess the practical applicability of these systems, the genetic stability and pathogenicity changes of edited strains must be confirmed in animal infection models.
💡 In the laboratory, researchers can first use Cas9-NHEJ to delete suspected pathogenic genes and screen for infection phenotypes, then use PE to modify specific bases or amino acids in active regions to narrow down functional mechanisms. For example, by administering edited strains to hamster infection models and comparing their ability to induce acute disease and colonize the kidneys, the roles of genes at different stages of infection can be distinguished. In industrial applications, designing vaccine strains by removing toxin-related genes without markers is a promising application. However, factors such as the stability of production strains during serial cultivation, off-target mutations, and environmental release risks must be evaluated. Since these pathogens are being precisely edited, biosafety management and regulatory standards must also be established in parallel.