Genomics, proteomics, systems biology โ decoding the blueprint of life.

Background Messenger RNA (mRNA)-based therapeutics have expanded beyond their success in COVID-19 vaccines to include cancer immunotherapy, gene editing, and protein replacement therapy for rare diseases. To ensure that injected mRNA remains intact and reaches target cells, a delivery system is essential for safe and effective transport. Lipid nanoparticles (LNPs), widely used in clinical settings, have demonstrated their utility but still require improvement. Key limitations include lower-than-expected intracellular delivery efficiency and potential immunogenicity due to accumulation in the body after repeated administration. Conventional LNPs tend to be trapped in endosomes after crossing the cell membrane, leading to mRNA degradation if they cannot escape. To efficiently release therapeutic genetic material into the cytoplasm, the development of intelligent materials that change their structure in response to specific microenvironments is crucial. This has led researchers to focus on developing smart nanocarriers that can precisely release drugs by recognizing physiological characteristics of disease sites or cells. Key Findings The researchers comprehensively analyzed various nanocarrier technologies that utilize redox (oxidation-reduction) responsiveness, considering the chemical environment both inside and outside cells. Compared to normal cells, the intracellular and tumor microenvironments exhibit significantly higher glutathione (GSH) concentrations (hundreds of times higher) and increased production of reactive oxygen species (ROS). Nanomaterials designed to release drugs in response to these concentration differences are the core of redox-responsive nanocarriers. The main technology platforms are classified into reduction-responsive, oxidation-responsive, and dual-responsive systems. Reduction-responsive carriers incorporate disulfide bonds into their molecular structure, allowing them to break down and rapidly release mRNA when exposed to high GSH concentrations within cells. Oxidation-responsive platforms utilize the characteristic of certain hydrophobic materials changing to hydrophilic in the excessive ROS environment of tumors or inflammatory sites, inducing the breakdown of the carrier. Recently, hybrid nanocarriers combining the advantages of polymers and lipids have been synthesized, successfully enhancing both stability and cell permeability. These carriers are precisely controlled to maintain their integrity during circulation in the body and disassemble upon entering cells. Significance and Outlook Redox-responsive smart nanocarriers have the potential to significantly improve the targeted delivery efficiency of mRNA therapeutics. By greatly improving the efficiency of endosomal escape, which is the process by which therapeutic molecules reach the cytoplasm, they can enhance therapeutic effects and reduce the required dosage. This, in turn, can mitigate systemic side effects associated with drug overdoses. However, there are clear challenges to be addressed before commercialization. The long-term safety of nanomaterials in terms of their degradation and excretion in the body must be definitively demonstrated. The technology also needs further research to control potential immunogenicity that may arise during repeated administration. Furthermore, the standardization of manufacturing processes for the large-scale production of uniform-quality nanoparticles is a critical industrial hurdle. In the future, researchers plan to introduce artificial intelligence (AI) to predict optimal responsive chemical structures and expand the scope of application to personalized theranostic systems that combine diagnosis and treatment.
๐ก This technology is expected to accelerate the commercialization of personalized immunotherapies and gene editing technologies for patients with difficult-to-treat cancers. By designing drugs to be released only in the unique redox environment of tumor tissues, it minimizes toxicity to normal cells. For example, nanoparticles loaded with patient-specific cancer vaccines can be injected intravenously and designed to induce the production of target proteins only within immune cells in the liver and other lymphoid organs. In genetic metabolic diseases requiring repeated administration, it can also improve safety and open up the possibility of long-term treatment. If an AI-based platform is established to rapidly screen for optimal lipid structures during the manufacturing process, the drug development period can also be significantly shortened.

Background Interdisciplinary collaboration among archaeology, genetics, and linguistics is essential when tracing human migration and the spread of civilizations. The expansion of the Bantu language family, encompassing the region south of the Sahara in Africa, is considered one of the largest language dispersal events in human history. Numerous scholars have endeavored to elucidate the early diversification process of this language family. Traditional phylogenetic approaches have primarily relied on tree-like analyses, borrowed from models of biological diversification. This methodology operates under the assumption that languages, during their divergence, experienced little to no interaction. However, unlike biological species, languages frequently undergo horizontal transmission, where vocabulary mixes due to contact with neighboring groups. The Bantu language family likely experienced extensive interaction and exchange of vocabulary among different groups throughout its history. Simple phylogenetic models that exclude this language contact can distort the true evolutionary trajectory of languages. Previous analyses treated language contact as mere noise, failing to provide a clear conclusion on how the Bantu language family actually diversified and spread thousands of years ago. It is crucial to fully incorporate the actual interaction, i.e., the mixing of vocabulary, into the computational model. Key Findings A multinational research team led by Dr. Patrรญcia Santos applied coalescent theory, originally developed in population genetics, to linguistics. The team proposes a novel mathematical computational model that treats language contact as a natural evolutionary process. This model incorporates the concepts of gene flow and mutation from genetics, representing them as lexical borrowing between languages and the rate of independent vocabulary change, respectively. To validate this model, the researchers employed an Approximate Bayesian Computation (ABC) approach, comparing it with actual Bantu vocabulary data. The analysis yielded surprising results. It revealed that the currently available vocabulary data of the Bantu language family has largely lost its early historical information. The rate of vocabulary change and the intensity of language contact within the Bantu language family were significantly higher than expected. This rapid change and frequent contact completely obscured the original signals from the early stages of language diversification. Consequently, it is now impossible to reconstruct the early diversification history of the Bantu language family using only the currently available vocabulary dataset. Significance and Prospects This research serves as a serious warning for interdisciplinary studies aimed at reconstructing human migration routes. Previous archaeological and genetic studies have used linguistic phylogenetic analysis as a triangulation tool to validate their hypotheses. However, if the inherent lexical mixing in language data makes early phylogenetic reconstruction fundamentally impossible, then the hypotheses based on it are also questionable. The research team explicitly states that the linguistic diversification structure of the Bantu language family should not be hastily cited as definitive evidence for reconstructing prehistoric human history. Future research should focus on refining the model by incorporating linguistic markers that are less susceptible to contact, such as grammatical structures or phonological features, in addition to vocabulary. Just as phylogenomic analysis in bioinformatics overcame the limitations of single-gene analysis, a multi-marker model should be developed in linguistics. The validated mathematical framework can be usefully applied not only in Africa but also in analyzing Native American languages and large language families in Eurasia. By acknowledging the limitations of linguistics, we can pave the way for more precise reconstructions of prehistory.
๐ก This modeling technique can be applied not only to the study of human history but also to improve the performance of natural language processing (NLP) artificial intelligence (AI) that handles multilingual translation. Modern large language models (LLMs) often make errors when evaluating the semantic similarity of words in multilingual environments with frequent historical contact. A promising alternative is to incorporate a horizontal contact calculation algorithm based on coalescent theory into the training process of LLMs. For example, in complex multilingual environments such as Swahili, which is a fusion of several Bantu languages and Arabic, the model can trace the origin and time of introduction of each vocabulary item, thereby minimizing translation errors. Furthermore, it can assist in predicting the vocabulary loss patterns of endangered minority languages and designing effective language restoration scenarios.

Background Artificial intelligence (AI) is increasingly influential in protein structure prediction and genomics. Existing AI models require vast amounts of training data, ranging from thousands to tens of thousands, to ensure performance. However, in practical biological research, high-quality, experimentally validated data is extremely scarce, hindering the adoption of machine learning (ML) techniques. In particular, identifying unknown enzymes that modify phenazine, an organic compound affecting ecosystems and the human body, has been a long-standing challenge. Phenazine is a toxic substance secreted by bacteria, and tracing the biochemical reactions that render it harmless requires numerous control experiments. However, the genetic information related to this is very limited, making it nearly impossible to build a predictive system using conventional AI techniques, according to the researchers. Key Findings The research team, led by Professor Dianne Newman at the California Institute of Technology (Caltech), developed 'ML-CITO', an AI framework that operates with extremely small data by utilizing genomic contextual information. The model started by using only 14 known phenazine-modifying enzyme gene sequences as initial training data (seed data). Subsequently, the researchers introduced a genomic data augmentation technique to track homologous genes located near the phenazine biosynthetic gene cluster (BGC), increasing the training data to approximately 600. The augmented data was processed through a pre-trained protein language model (PLM), 'ESM Cambrian 600M', and transformed into a high-dimensional vector of 1,152 dimensions. This vector information was input into a three-layer multi-layer perceptron (MLP) neural network consisting of 512, 128, and 64 units, followed by the application of contrastive learning. The ML-CITO model precisely classifies candidate enzymes that react with phenazine within the protein space. The researchers successfully identified a protein from Pantoea agglomerans W2I1, a soil bacterium, that exhibits actual enzyme activity among the model's predicted candidates. The newly identified enzyme is a phenazine-thiol conjugating enzyme (PTC) that directly binds phenazine with glutathione (GSH), which regulates intracellular redox status. Previously, the conjugation reaction between phenazine and GSH was considered a non-enzymatic chemical reaction. The researchers demonstrated through biochemical experiments that the PTC they discovered catalyzes the reaction, directly mitigating the toxicity of phenazine. The researchers searched over 200,000 bacterial genomes and confirmed that approximately 3,415 PTC gene homologs are widely distributed across more than 30 phyla. Significance and Prospects This research presents a new breakthrough for research fields where protein data is extremely limited, making it difficult to apply ML models. By combining genomic contextual information with PLMs, the study demonstrates that high-performance prediction is possible even with small amounts of data. This framework is expected to be introduced in various bio-industrial fields requiring the elucidation of unknown enzyme functions, such as drug development, biomanufacturing, and environmental remediation. However, the fact that the ML-CITO model relies on data augmentation based on conserved physical locations in the genome (e.g., BGC) is a challenge to be overcome. Proteins with unclear genetic context or those scattered throughout the genome may not fully benefit from the data augmentation effect. The researchers plan to expand the model's applicability to a wider range of protein clusters and to verify the detailed biochemical characteristics of the 3,415 gene homologs identified in this study.
๐ก The practical value of this research is particularly evident in the development of new therapeutic strategies to control antibiotic-resistant strains. Phenazine toxins secreted by Pseudomonas aeruginosa, an opportunistic pathogen, induce chronic inflammation in patients with lung diseases and are key factors in increasing antibiotic resistance. By using this model to design inhibitors that interfere with the action of phenazine or regulate the activity of PTC, it may be possible to improve the efficacy of drugs for treating resistant bacteria. Furthermore, in the field of environmentally friendly bioprocesses, it is possible to quickly find unknown useful enzymes and improve the efficiency of synthesizing high-value-added chemicals. Redesigning the metabolic pathways of beneficial microorganisms that decompose organic waste or purify recalcitrant compounds is also a promising application.

Background The human genome contains millions of enhancers that regulate gene expression in specific cell types. A significant proportion of disease-associated variants reside in enhancers, rather than protein-coding regions; however, identifying the target genes and cell types regulated by these enhancers is challenging. This is because enhancers can be located tens or hundreds of kilobases away from their target genes, and they may act by regulating genes other than the closest one. Existing Activity-by-Contact (ABC) and ENCODE-rE2G models predict regulatory relationships by leveraging chromatin activity and three-dimensional contact information. However, bulk tissue analysis mixes signals from multiple cells, making it difficult to distinguish regulatory circuits in rare or transient cell states. Single-cell chromatin accessibility analysis (scATAC-seq) and single-cell RNA sequencing (scRNA-seq) have emerged as alternatives, but unsupervised learning approaches that rely on simple correlations between accessibility and expression levels lack objective standards for evaluating accuracy. To address these limitations, the researchers developed scE2G, a family of single-cell enhancer-gene prediction models. The results were published on August 3, 2026, in Nature Genetics. Key Findings scE2G consists of scE2G-ATAC, which uses only scATAC-seq data, and scE2G-Multiome, which uses data from scRNA-seq and ATAC-seq performed on the same cells. Both models are supervised learning-based logistic regression classifiers. The researchers trained the models using 13,420 enhancer-gene candidate pairs identified through CRISPR screening in K562 erythroleukemia cells. Among these, 466 represented positive associations, where gene expression significantly decreased after enhancer inhibition, and 9,876 represented negative associations. scE2G-ATAC calculates six features: ABC score, chromatin accessibility of enhancers and promoters, distance and gene density between the two regions, and promoter type. scE2G-Multiome adds the Kendall correlation coefficient of enhancer accessibility and gene expression within individual cells, and uses this in combination with the ABC score to create an ARC-E2G feature. The possibility of overfitting was also reduced through cross-validation, where one chromosome was excluded at a time. The researchers compared the performance of the two scE2G models with ten existing single-cell models using three independent datasets: CRISPR perturbation, fine-mapped expression quantitative trait loci (eQTL), and genome-wide association study (GWAS) variant-gene associations. Both scE2G models achieved the highest precision and area under the precision-recall curve, not only in K562 cells but also in 4,175 additional CRISPR datasets from five other cell types. In eQTL evaluation, scE2G-Multiome had a recall of 13.9% and a 14.9-fold enrichment of variants. The next best performing non-ABC-based model, SCENIC+, had a recall of 1.4% and a 10.3-fold enrichment. By constructing regulatory maps for 45 cell types from peripheral blood, bone marrow mononuclear cells, and islets, the researchers predicted an average of 48,758 connections per cell type in the 39 types that met the quality criteria. On average, 5.6 enhancers were associated with a single expressed gene, and the average distance between enhancers and their target promoters was 86.3 kilobases. Significance and Implications scE2G is significant because it can narrow down the cell types and target genes affected by disease variants, even in mixed-cell populations. The researchers linked 1,450 variants in 1,892 non-coding regions to 1,351 genes, and identified 458 regions that targeted genes further away from the nearest transcription start site. For example, rs7696969, a variant associated with lymphocyte count, is located within INPP4B. scE2G suggests that this variant may regulate INPP4B and IL15 in natural killer cells and T cells. The distance from the variant to the promoters of the two genes is 441 kilobases and 769 kilobases, respectively. The posterior inclusion probability that the variant is an eQTL for INPP4B was also 66.6%, and the independent gene prioritization analysis, PoPS, also selected the two genes as top candidates. However, this is not a definitive confirmation of functional causality, but rather a hypothesis that requires further experimental validation. For stable application of the model, it is recommended to have at least 100 cells per cell type, a minimum of 2 million ATAC fragments, and 1 million RNA unique molecular identifiers (UMIs). A limitation is that only 466 positive associations were used for training, and most of these were from K562 cells. Topological regulatory elements other than enhancers, such as CTCF binding sites, are not included as prediction targets. Accumulating large-scale CRISPR validation data from multiple primary cells and incorporating Hi-C or H3K27ac information will be necessary to expand the model into a clinically reliable regulatory map.
๐ก Applying scE2G to single-cell multi-omics data from patient tissues can narrow down the cell types and target genes regulated by GWAS variants. For example, if a non-coding variant associated with an autoimmune disease overlaps with an enhancer in a specific T cell subtype, the connected target genes can be validated by CRISPR inhibition or by examining changes in expression in organoids or primary cells. For pharmaceutical companies, scE2G has the potential to be a target discovery and patient stratification tool. By linking disease variants, acting cell types, and target genes, it can strengthen the genetic basis of drug targets and facilitate the design of appropriate biomarkers. However, the connections made by scE2G are probabilistic predictions, so they should be validated with CRISPR perturbation, protein expression, and three-dimensional chromatin contact data before being used for drug development decisions.

Background In cancer genomic analysis, the focus is typically on identifying the presence of specific somatic mutations. However, the gene mutation dosage (GMD), which represents the number of gene copies with the mutation, may also be relevant to tumor evolution. Calculating GMD requires the joint interpretation of somatic mutations and copy number alterations. The researchers developed INCOMMON, a Bayesian tool that estimates mutation copy number and multiplicity from targeted sequencing data without requiring normal control samples, and applied it to large-scale clinical datasets of various cancer types. This method probabilistically estimates tumor purity and read counts per chromosome copy to calculate the gene state created by the combined effects of mutations and copy number alterations. Key Findings The researchers analyzed over 60,000 clinical cancer samples and over 500,000 mutations obtained from 39 major solid tumor types. In cross-validation, the proportion of predictions with a total copy number and mutation multiplicity error of less than one copy was 78.6% and 96.4%, respectively. Dividing over 20,000 patients into groups based on the GMD of multiple genes revealed 46 cancer-specific biomarkers that predicted overall survival. Among these, 13 were not identified by conventional methods that only distinguish the presence or absence of mutations. Additionally, 26 biomarkers were associated with metastatic spread, and 20 predicted the tendency to metastasize to specific organs. These results demonstrate statistical associations and predictive performance, but do not prove a causal relationship where GMD directly causes metastasis. Significance and Prospects INCOMMON is designed to estimate mutation copy number and multiplicity from read counts in clinical targeted sequencing without requiring raw FASTQ or BAM files or normal control samples. Adding GMD information to the presence or absence of mutations can broaden the scope of identifying prognostic and metastasis-related biomarkers. The ability to re-analyze previously accumulated clinical panel data also enhances the research's applicability. However, to use the analysis results in actual treatment decisions, reproducibility should be confirmed in an independent patient cohort, and clinical utility should be prospectively validated for each cancer type and treatment method. The current findings do not represent the immediate completion of precision medicine, but rather present a computational method that refines biomarker discovery.
๐ก Even tumors with the same driver mutation can have different clinical outcomes depending on the copy number of the mutation and the surrounding copy number alterations. The researchers showed that GMD can be independently associated with survival, metastatic potential, and organ-specific metastasis in over 60,000 samples. Importantly, they identified additional biomarkers that were missed by simple mutation presence/absence analysis. However, these are predictive results obtained from retrospective large-scale data, so external validation and prospective evaluation are needed before applying them to patient treatment.

Background Breast cancer is a heterogeneous disease characterized by significant differences in molecular profiles and treatment responses, even within the same organ. Estrogen receptor (ER), progesterone receptor (PR), and human epidermal growth factor receptor 2 (HER2) are established biomarkers that guide decisions regarding endocrine therapy and HER2-targeted treatment. BRCA1 and BRCA2 mutations are relevant for assessing hereditary risk and for the use of poly(ADP-ribose) polymerase (PARP) inhibitors. However, conventional classification alone is insufficient to fully explain tumor evolution and drug resistance. Even within the same HR-positive or HER2-positive breast cancer subtype, treatment outcomes and recurrence patterns can vary. Furthermore, biopsies, which involve sampling only a portion of the tumor, may not fully capture intratumoral heterogeneity. Image interpretation can also be influenced by factors such as breast density, imaging equipment, and the experience of the radiologist. To address these limitations, researchers have integrated recent advances in molecular biology, artificial intelligence (AI), and precision medicine. This review provides a comprehensive overview of breast cancer biomarkers, diagnostics, and treatment strategies, rather than presenting the results of a clinical study involving a new patient cohort. Key Findings The review describes the breast cancer decision-making process as a combination of 'established and novel biomarkers.' While ER, PR, HER2, and BRCA mutations currently guide treatment decisions, alterations in tumor suppressor genes such as TP53, PTEN, and STK11 are presented as potential markers for more detailed interpretation of tumor behavior and resistance mechanisms. However, these three genes are not yet established as independent criteria for treatment decisions across all breast cancers. Their prognostic value and predictive ability for treatment response should be prospectively validated in specific cancer subtypes. In the diagnostic arena, machine learning and deep learning are being applied to medical imaging, including mammography, to identify subtle lesions. These technologies are also being integrated into multimodal systems that combine clinical information, imaging data, and genomics. This approach differs from traditional methods that rely on single images or biomarkers by calculating the correlations between different data layers to support early detection, diagnostic assistance, and risk stratification. The abstract does not provide specific performance metrics such as accuracy or sensitivity, the size of the training dataset, or the names of specific AI models. Therefore, the improvements in diagnostic performance should be interpreted as a general trend, rather than as evidence of the clinical superiority of a specific algorithm. The therapeutic landscape has also expanded. In addition to surgery and cytotoxic chemotherapy, treatment strategies now include endocrine therapy, immunotherapy, targeted therapy, antibody-drug conjugates (ADCs), and gene-based approaches. In advanced breast cancer, cyclin-dependent kinase 4/6 (CDK4/6) inhibitors, PARP inhibitors, phosphoinositide 3-kinase (PI3K) inhibitors, and selective estrogen receptor degraders (SERDs) have emerged as important targeted therapies. Each of these agents is selected based on the patient's hormone receptor status, genetic mutations, and prior treatment history and response. Significance and Future Directions The key concept presented in this review is that AI should serve as a decision-support tool for physicians, rather than replacing them with automated diagnostic systems. By highlighting suspicious lesions in images and integrating this information with pathology and genomic data, AI can improve the consistency of patient selection and treatment sequencing. Nanotechnology-based drug delivery systems offer a strategy to increase drug exposure in the tumor and reduce toxicity to normal tissues. CRISPR/Cas9 is expected to be a valuable tool for identifying resistance genes and validating therapeutic targets. Several challenges remain before these technologies can be widely implemented in clinical practice. AI systems must be validated in external datasets to ensure that their performance is maintained across different hospitals, patient populations, and imaging equipment. The potential for bias in training data and the lack of transparency in decision-making algorithms must also be addressed. CRISPR/Cas9 therapy faces challenges related to off-target editing, tumor cell delivery efficiency, and long-term safety. Nanocarriers must meet requirements for in vivo distribution, manufacturing reproducibility, and large-scale production. Given the lack of direct comparative trials or survival data in the review, it is important to distinguish between the technologies presented as established clinical standards and those that are still in the development stage. The actual adoption of these technologies will depend on prospective studies that evaluate not only accuracy but also survival, reduction in unnecessary tests, toxicity, and cost-effectiveness.
๐ก In clinical practice, AI-assisted mammography can prioritize images for review, and multidisciplinary meetings can integrate findings from pathology and genomic testing, including ER, PR, HER2, BRCA, and PIK3CA mutations, to guide treatment decisions. This approach can help clinicians determine the appropriateness of CDK4/6 inhibitors or SERDs for patients with HR-positive advanced breast cancer and PARP inhibitors for patients with BRCA mutations. For the industry, this review highlights opportunities for developing standardized multimodal platforms that integrate imaging, pathology, and genomic data, as well as companion diagnostics that link biomarkers to treatment decisions. However, patient selection based on TP53, PTEN, and STK11, as well as CRISPR-based therapies and nanocarriers, must first demonstrate clinical efficacy and manufacturing quality. The actual adoption of these technologies will depend on prospective studies that evaluate not only accuracy but also survival, reduction in unnecessary tests, toxicity, and cost-effectiveness.

Background Humans are born with approximately one million nephrons, the functional units of the kidney, which will be used throughout their lifetime. These functional units are formed only during the fetal developmental stage, and their formation ceases completely after birth. The total number of nephrons in the two kidneys of an adult varies from 200,000 to 2,000,000, up to a 10-fold difference between individuals. If the number of nephrons falls below the standard range during the fetal period due to environmental or genetic factors, leading to oligonephronia, the risk of developing chronic kidney disease or hypertension increases in adulthood. This is the prevailing view in the medical community. Research to elucidate the detailed mechanisms of fetal kidney development has primarily relied on animal models such as mice. However, due to differences in kidney structure between humans and animals, there are limitations in directly applying the results of animal experiments to patients. Existing single-cell analysis techniques can distinguish between cell types but do not show the spatial information of individual cells or their interactions with surrounding cells. To clearly understand the developmental process of a highly organized organ like the kidney, a precise map is needed that simultaneously analyzes the spatial structure of cells and the signaling network between cells. Key Findings A joint research team from the University of Pennsylvania School of Medicine and the Children's Hospital of Philadelphia (CHOP) obtained more than 700,000 cells from fetal kidney tissue between 12.5 and 20.5 weeks of gestation. They combined single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics (ST) to create the first map of cell gene expression and spatial location. Analysis revealed that kidney development is guided not only by cell-specific genetic programs but also by soluble signals from neighboring cells. In particular, insulin-like growth factor 2 (IGF2) was found to be a key signal in maintaining and proliferating immature kidney progenitor cells (NPCs). The mechanism involves IGF2 secreted by surrounding cells binding to NPC receptors, thereby preserving the stem cell properties. This was demonstrated in an experiment in which IGF2 supply was blocked in the culture medium, causing NPCs to lose their stem cell ability and differentiate rapidly. Furthermore, analysis of a large population genetics database confirmed that IGF2 gene variants are associated with adult kidney size and nephron number. This is noteworthy in that it elucidates the genetic factors that regulate the proliferation limit of stem cells. Significance and Prospects This research lays the foundation for overcoming the limitations of existing developmental research that relies on animal models. In particular, by revealing the cell-to-cell signaling network during the fetal period, it provides a breakthrough for organoid technology, which aims to create mini-kidneys in the laboratory. Existing organoids have the disadvantages of having a small number of functional units and a simple internal structure. By reproducing the cell arrangement and IGF2 concentration changes revealed by ST, a more sophisticated artificial kidney can be constructed. However, there are still many points to be addressed before this can be linked to the development of organs for transplantation or actual treatment. Fetal kidney formation involves various factors, including the immune relationship with the mother, blood vessel formation, and blood flow. It is difficult to reproduce all of these complex biological responses in a planar culture environment by simply adding individual factors. In the future, the combination with an engineering platform that can increase the simulation accuracy, such as a microfluidic chip, is considered a key task. If research continues to overcome this barrier, it is expected that it will be possible to provide customized cell therapies for patients with intractable kidney diseases.
๐ก The fetal kidney development information obtained in this study is expected to be used in various scenarios in the pharmaceutical industry and clinical practice. First, it can greatly improve the efficiency of nephrotoxicity assessment in the drug development process. By creating organoids that are patterned similarly to the actual human kidney by reflecting the ST map information, the nephrotoxicity of drug candidates can be screened more accurately before clinical trials, reducing the failure rate of drug development. Second, early prediction and personalized management of pediatric and adult kidney diseases can be realized. By examining IGF2-related gene variants in the prenatal stage, high-risk individuals for oligonephronia can be identified early, and preventive kidney protection strategies, such as dietary improvements and restrictions on drug use, can be established from infancy. In the long term, it is expected to play a central role as a blueprint for the development of functional artificial kidneys to solve the problem of organ shortages.

Background Escherichia coli is mostly harmless in the gut, but some strains cause severe food poisoning with deadly toxins. Shiga toxin-producing E. coli (STEC) causes hemorrhagic colitis and hemolytic uremic syndrome (HUS). Recently, the surge in multidrug-resistant and hypervirulent strains has raised serious concerns for public health. E. coli has a high rate of gene exchange, making it difficult to treat with existing antibiotics alone, and antibiotic administration can promote toxin release, worsening symptoms. Therefore, the development of a next-generation vaccine to prevent infection is urgently needed. Conventional methods require a great deal of time and resources to culture and attenuate pathogens, but reverse vaccinology, which uses computer simulations, has emerged as an alternative. This technology designs optimal protein fragments that induce an immune response based on genomic data. Key Findings The researchers used computer simulations to design a multiepitope mRNA vaccine candidate that targets Shiga toxin 1 (Stx1), a key factor in STEC. They used an immunogenicity prediction algorithm to precisely identify sites that stimulate immune cells. The vaccine design focuses on simultaneously activating immune cells in the body. The researchers selected two cytotoxic T-lymphocyte (CTL) epitopes, three helper T-lymphocyte (HTL) epitopes, and two linear B-lymphocyte (LBL) epitopes. Strict filtering criteria were applied to exclude potential allergens and toxins and maximize antigenicity. The selected epitopes are linked with EAAK, AAY, GPGPG, and KK linkers. The adjuvant PefE protein was fused to amplify the immune response, creating a final multiepitope structure consisting of a total of 184 amino acid residues. The predicted molecular weight of the vaccine candidate is 19969.45 Da, and it recorded a high score of 0.9278 in the VaxiJen v.2.0 antigenicity analysis tool. The isoelectric point is 10.31, indicating stability in a basic environment. The three-dimensional structure analysis revealed an alpha-helix content of 21.74%. The binding force with human Toll-like receptor 2 (TLR2) and Toll-like receptor 4 (TLR4) was measured, showing weighted energy scores of -1096.2 kcal/mol and -1110.5 kcal/mol, respectively. This indicates that numerous hydrogen bonds and salt-bridge networks strongly interact with the receptor. Immune simulation results showed that after vaccine administration, there was a strong clonal expansion of B cells, T cell activation, and increased secretion of cytokines (IFN-ฮณ, IL-10). The codon adaptation index (CAI), which determines translation efficiency, is 0.87. The minimum free energy (MFE) of the mRNA secondary structure is -204.10 kcal/mol, which supports stable synthesis in cells. Significance and Implications This study demonstrates that vaccine candidates can be derived from genomic data and algorithms without directly culturing strains in the laboratory. Simulation-based design provides a basis for rapidly developing countermeasures in pandemic situations or when new variants emerge. It is also expected to be widely applicable to the development of vaccines for various multidrug-resistant bacteria, including E. coli. The results of this study are based on a predictive model using computer simulations, which has limitations. It is uncertain whether the numerical values and binding forces in the algorithm will be the same in the actual in vivo environment. Due to the nature of mRNA vaccines, there is also a risk that stability may change or unexpected immune hypersensitivity reactions may occur during the lipid nanoparticle (LNP) formulation process. In vitro expression tests and in vivo efficacy and toxicity tests in animals are essential for commercialization. Further research is needed to prove actual protective efficacy before clinical trials can begin.
๐ก This study presents a concrete scenario that can change the treatment paradigm for multidrug-resistant foodborne E. coli infections. In the event of a hypervirulent E. coli infection in a mass catering facility or livestock farm, it would be useful to proactively secure immunity with this vaccine in infected patients, rather than simply alleviating symptoms or prescribing antibiotics that may promote toxin release, thereby preventing progression to severe HUS. Industrially, in silico design can drastically reduce the development time from years to months, from actual vaccine synthesis to animal testing. Multiepitope design is expected to greatly contribute to simplifying vaccine production processes, increasing purification efficiency, and reducing manufacturing costs. This will establish it as a useful public health solution for foodborne disease prevention programs in developing countries with poor healthcare infrastructure.

Background Following the decoding of the human genome, a major challenge in modern genetics has been to elucidate the specific gene regulatory activities of the 3 billion base pairs of DNA. Mainstream genetic research has focused on Genome-Wide Association Studies (GWAS) to identify correlations between specific variants and diseases, making it difficult to directly explain the molecular biological mechanisms of non-coding region variants. Recent advances in functional genomics data and artificial intelligence (AI) have led to the emergence of 'sequence-to-function models (seq2func models)' that predict gene expression levels from DNA sequence. This AI model predicts the regulatory mechanisms of target genes at the molecular level, opening a new paradigm in genomics research. However, existing models exhibit a vulnerability in that their predictive power drops sharply when applied to datasets outside their training range. Furthermore, the mechanisms by which gene regulatory principles operate within complex neural networks are opaque, and reliable validation remains a challenge. Key Findings This review article identifies model architecture, training data, prediction tasks, model interpretation, and evaluation strategies as key factors determining the performance and reliability of seq2func models. First, in terms of model architecture, Convolutional Neural Networks (CNNs) are good at identifying local regulatory patterns (motifs) within the DNA sequence. In contrast, Transformer-based models effectively capture long-range enhancer-promoter interactions of 100 kilobases (kb) or more by applying a self-attention mechanism. Recently, State Space Model (SSM)-based HyenaDNA, which significantly reduces computational complexity while handling long sequences of up to 1 megabase (Mb), has been studied as a new alternative. However, the limitations of training data and prediction tasks remain. Most seq2func models are trained on epigenome data that is biased towards specific cell lines or chromosomes, which can lead to the model learning statistical noise in the data rather than actual biological rules. The inability to distinguish between cell types in bulk data is also a factor that reduces generalization ability. Accordingly, the methods for verifying the AI's predictive basis are also diversifying. Integrated Gradients, which shows the important positions in the input sequence, and In Silico Mutagenesis (ISM) are representative examples. The researchers propose the introduction of Global Importance Analysis (GIA), which quantifies biological features such as the distance or order of regulatory patterns, beyond single base analysis. Furthermore, a rigorous evaluation strategy that utilizes untrained chromosomes or new cell states during model validation is necessary to accurately measure the model's generalization ability. Significance and Prospects This analysis shows that the field of genomic AI should move away from simple parameter expansion competitions and focus on interpretability and robustness that explain biological causality. In particular, the framework that comprehensively verifies the impact of model structure and data characteristics on generalization performance is expected to be a useful guide for researchers to design more sophisticated neural networks in the future. However, there are also practical challenges to be overcome. Most of the existing seq2func models do not fully reflect dynamic physiological phenomena, such as changes in the concentration of transcription factors (TFs) or three-dimensional structural changes in chromatin, during the learning process. Therefore, the introduction of context-based models that integrate cell state information into the input is essential. Furthermore, a standardized benchmark platform for evaluating whether the AI model has learned actual biological rules must be established in order to achieve tangible results in the field of precision drug development.
๐ก The genomic AI validation guidelines presented in this study can bring immediate changes to the personalized precision medicine and drug development industries. A typical scenario is the identification of novel non-coding genetic variants in patients that cause rare diseases. If AI can accurately predict pathogenic variants that directly impair gene expression, clinicians will be able to easily identify the actual cause of the disease from tens of thousands of genetic variants. Furthermore, the efficiency of designing synthetic promoters that optimize the activity of therapeutic genes in gene therapy development can be maximized. This involves designing a sequence with AI to be strongly expressed only in the desired cell type, and then verifying the stability of the design in real-time using GIA. This simulation will help screen effective candidate materials before animal experiments, resulting in reduced drug development time and costs.

Background Pregnancy is a process characterized by rapid and profound physiological changes throughout the maternal system. The cardiovascular, metabolic, and immune systems must dynamically adapt to support fetal growth, ultimately leading to a healthy outcome. However, the genetic mechanisms regulating these physiological changes during pregnancy remain largely unknown. Previous genome-wide association studies (GWAS) have primarily focused on non-pregnant adults and have been limited by their reliance on single-time-point data during pregnancy. Conducting longitudinal studies with large cohorts of pregnant women and data collection throughout the entire gestational period poses significant logistical and financial challenges. To address these challenges, researchers propose an alternative approach: repurposing high-resolution genome sequencing data generated during non-invasive prenatal testing (NIPT). NIPT is a routine screening test that analyzes cell-free DNA in maternal blood to detect fetal chromosomal abnormalities. By linking the genetic data obtained from this test with clinical records, a large cohort can be effectively established. The researchers used this approach to elucidate the dynamic hormonal changes and metabolic processes that occur during pregnancy. Key Findings The researchers conducted a comprehensive analysis, linking genomic information from up to 121,579 Chinese pregnant women with data on 111 clinical phenotypes. These phenotypes include lipid and glucose metabolism markers, blood cell parameters, and pregnancy complications. The analysis identified thousands of independent genetic loci associated with physiological traits during pregnancy. A significant proportion of these genetic regions represent novel findings not previously reported in pregnancy-related traits. A particularly noteworthy discovery is the identification of dynamic genetic effects that vary across gestational weeks. Analysis of blood parameters measured repeatedly during pregnancy revealed that approximately 18% of the overall genetic signal exhibits interactions with gestational week, resulting in changes in effect size. Specific genetic variants demonstrate strong effects on maternal blood cell counts early in pregnancy, but these effects diminish or reverse later in pregnancy. These time-specific genes are strongly enriched in biological pathways involved in hormone response, immune regulation, and fetal growth. Furthermore, Mendelian randomization (MR) analysis was used to investigate the causal relationship between physiological abnormalities during pregnancy and the risk of chronic diseases in mothers after delivery. The results indicate that disruptions in metabolic markers during pregnancy are directly associated with an increased risk of cardiovascular disease and type 2 diabetes in mothers in the long term. Significance and Implications This research redefines pregnancy not as a static physiological state, but as a dynamic process of continuous genetic regulation. It supports the need to move beyond simply applying existing adult research findings to pregnant women and instead establish a dedicated genomic reference for pregnant women. Moreover, the study demonstrates a new paradigm for healthcare data utilization by repurposing NIPT data, which is typically used for a single-time-point test, into a valuable resource for research. The study demonstrates the feasibility of leveraging existing healthcare infrastructure to create a large, high-quality genomic cohort of 120,000 individuals. However, some limitations should be acknowledged. This study focused solely on Chinese pregnant women, and further validation is needed to determine whether the same effects are observed in other ethnic groups. Additionally, the analysis was based on data collected at a limited number of time points during NIPT, and future studies should incorporate long-term follow-up data after delivery. Why It Matters This research has the potential to be used as a precision medicine tool for personalized pregnancy management in the clinical setting. For example, during routine NIPT screening in early pregnancy, the fetal screening can be combined with an analysis of the mother's genomic information. This analysis can be used to calculate the risk of developing high-risk complications such as gestational diabetes and preeclampsia later in pregnancy, based on the individual mother's genetic profile and gestational week-specific prediction curves. Based on these predictions, healthcare providers can prescribe proactive dietary adjustments or design personalized clinical management plans to prevent complications. From an industrial perspective, this research is expected to change the paradigm of the existing NIPT screening market. Diagnostic companies are likely to expand their business areas from simple fetal abnormality detection to comprehensive healthcare solutions that cover the mother's entire lifespan. Given that the causal relationship between gestational metabolic diseases and chronic diseases after delivery has been genetically elucidated, research to identify new therapeutic targets in the development of female-targeted drugs is also expected to accelerate.
๐ก This research has the potential to be used as a precision medicine tool for personalized pregnancy management in the clinical setting. For example, during routine NIPT screening in early pregnancy, the fetal screening can be combined with an analysis of the mother's genomic information. This analysis can be used to calculate the risk of developing high-risk complications such as gestational diabetes and preeclampsia later in pregnancy, based on the individual mother's genetic profile and gestational week-specific prediction curves. Based on these predictions, healthcare providers can prescribe proactive dietary adjustments or design personalized clinical management plans to prevent complications. From an industrial perspective, this research is expected to change the paradigm of the existing NIPT screening market. Diagnostic companies are likely to expand their business areas from simple fetal abnormality detection to comprehensive healthcare solutions that cover the mother's entire lifespan. Given that the causal relationship between gestational metabolic diseases and chronic diseases after delivery has been genetically elucidated, research to identify new therapeutic targets in the development of female-targeted drugs is also expected to accelerate.

Background Mammarenaviruses, a group of rodent-borne viruses, are significant human pathogens causing severe viral hemorrhagic fevers. These include viruses responsible for Argentine hemorrhagic fever and Lassa fever, which pose a global public health threat. These infections are endemic in parts of Africa and South America, causing periodic outbreaks and significant mortality. However, currently available vaccines are limited, and therapeutic options are scarce, highlighting the urgent need for effective countermeasures. Traditional vaccine development approaches have focused on using whole viral proteins as antigens to elicit an immune response. This approach may not provide broad protection against viral variants or related viruses. Furthermore, it can induce unwanted immune responses and potential adverse effects. Therefore, there is a need for a next-generation, broadly protective vaccine that can target multiple mammarenavirus species safely and effectively. Key Findings The researchers focused on designing a novel approach using computer-aided reverse vaccinology and immunoinformatics. They began by collecting and analyzing the sequences of major structural proteins from ten mammarenavirus species to identify conserved regions. They then used the Immune Epitope Database (IEDB) to screen for highly conserved and immunogenic epitopes. The final selected combination exhibits excellent immune-inducing capabilities by simultaneously activating both B and T cells. The resulting broad-spectrum multi-epitope fusion messenger RNA (mRNA) vaccine candidate incorporates additional features to enhance immune response efficiency. A tissue plasminogen activator (tPA) signal peptide was added to promote antigen export and facilitate rapid delivery to cells such as macrophages. Additionally, a PADRE (Pan-HLA DR-reactive epitope) sequence, which acts as an adjuvant, and linker technology were combined to flexibly connect the epitopes. The entire designed mRNA sequence underwent algorithm optimization to ensure stable expression within host cells. The physical stability and activity of the vaccine were assessed through various computational simulation techniques. The three-dimensional structure of the protein was predicted, followed by molecular docking analysis to assess the interaction with immune receptors. The results confirmed that the designed vaccine strongly binds to receptors that induce immune responses. Molecular dynamics (MD) simulations showed that the complex maintains a stable conformation over time, with minimal structural distortion. Furthermore, computer-based immune simulations revealed a strong and sustained antibody response, along with the activation of diverse immune cells. Significance and Future Directions This design represents a significant step towards developing a universal vaccine that targets the entire mammarenavirus family. By simultaneously analyzing ten mammarenavirus species and identifying common features, the researchers have developed a strategy to overcome the immune evasion mechanisms of rapidly mutating RNA viruses. In the future, if new mammarenavirus species emerge, it may be possible to rapidly develop a response based on the existing vaccine design. However, it is important to note that this study is a preliminary design stage based on computer simulations and immunoinformatics analysis. Further validation is needed to determine whether the results of the simulations will translate to in vivo conditions and whether the immune cells will respond as expected. Subsequent studies, including preclinical trials in animals and clinical trials in humans, are essential to demonstrate the actual efficacy and safety of the vaccine.
๐ก This research can be applied as a platform technology to shorten vaccine development time by several months during emerging infectious disease outbreaks. Previously, significant time was required to isolate and purify antigens after a virus emerged. However, by using the reverse vaccinology platform used in this study, it is possible to design and optimize immunogenic epitopes and generate optimized mRNA candidates within a few days after the gene sequence is obtained. In particular, the multi-epitope design approach, which defends against multiple mammarenavirus subtypes simultaneously, greatly simplifies the vaccine manufacturing process for emerging variants. This leads to reduced production costs and flexible use of manufacturing facilities, which can be a useful tool for addressing the global imbalance in vaccine supply, including in low-income countries.

Background Limitations of precision medicine and emerging needs Traditional diagnostics primarily relied on microscopic observation of cellular or tissue morphology. This approach is prone to subjective interpretation by physicians and makes it difficult to assess intratumoral heterogeneity. In diseases such as cancer and autoimmune disorders, where the pathogenesis varies among patients, it is often challenging to provide appropriate treatment based solely on morphological analysis. Biotechnology, which integrates molecular biology, nanotechnology, and computer science, has provided a foundation for overcoming these limitations. Molecular diagnostics, which analyzes gene mutations and biomarker expression patterns, and multi-omics technologies are representative examples. These technologies elucidate disease mechanisms at the genetic level and support the design of personalized treatments. To achieve high precision and minimize adverse effects, a mechanism-based diagnostic system is essential for targeted therapies. Key Findings High-dimensional data-driven diagnostics and targeted therapies The advancement of biotechnology is transforming diagnostic and therapeutic strategies. Molecular diagnostics and multi-omics contribute to elucidating the pathogenesis of various diseases, including cancer, genetic disorders, infectious diseases, cardiovascular diseases, and autoimmune diseases. Researchers have established diagnostic protocols that integrate gene mutation information and biomarker expression patterns to evaluate cellular heterogeneity and immune regulatory mechanisms. In the therapeutic field, gene editing, nanomedicine, immunotherapy, and targeted drug delivery are driving progress. Nanoparticle-based targeted delivery systems have significantly reduced the toxicity of drugs to normal tissues. Gene editing plays a role in correcting the underlying genetic cause of the disease, paving the way for mechanism-based precision therapy. Introduction of artificial intelligence-driven digital pathology The development of digital pathology analysis, which combines artificial intelligence (AI) and bioinformatics, is also noteworthy. Machine learning and deep learning are key technologies for capturing subtle pathological information through whole-slide image (WSI) analysis. Digital pathology systems integrate multi-dimensional data and clinical information to discover new biomarkers and classify patient risk. The combination of morphological and molecular data in pathology techniques is expected to become a new standard for distinguishing disease subtypes. Significance and Prospects Overcoming challenges for commercialization Convergent biotechnology is a key driver in accelerating the implementation of precision medicine. The combination of genetic analysis and targeted therapy is expected to realize personalized medicine. Academia and industry anticipate that this approach can reduce diagnostic errors and improve treatment efficacy, particularly in high-cost cancer treatments and rare diseases. However, there are several challenges to be addressed before clinical implementation. High costs of analytical equipment and technical complexity increase the financial burden on healthcare institutions. The lack of standardization of multi-omics and AI learning data hinders data sharing between institutions. Furthermore, consensus is needed on issues related to the exposure of personal genomic information and bioethical legal issues. Technology democratization for reducing healthcare disparities Therefore, future research should focus on technology democratization. The development of point-of-care (POC) molecular diagnostics and biosensor platforms that can be used in resource-limited healthcare settings is a major challenge. The commercialization of low-cost platforms is expected to enable patients in underserved areas to benefit from precision diagnostics, thereby reducing healthcare disparities.
๐ก This study demonstrates a roadmap for transforming precision diagnostics, which has traditionally relied on expensive equipment and complex laboratory tests, into a healthcare-integrated approach. A specific application scenario is the use of microfluidics-based biosensor chips in healthcare facilities, including large hospitals, health centers, and local clinics, to identify cancer gene mutations in a patient's single drop of blood within 30 minutes. In addition, portable digital scanners equipped with AI models can perform primary slide readings in remote areas or developing countries with limited access to pathologists, reducing the rate of misdiagnosis. From a pharmaceutical industry perspective, this approach can be immediately incorporated into commercialization strategies by pre-selecting clinical trial participants using multi-omics analysis, thereby maximizing the efficacy of new drug candidates and increasing the success rate of new drug development.

Background Identifying therapeutic targets that selectively kill cancer cells is a central challenge in cancer research. CRISPR-Cas9-based gene knockout screening has been useful in determining the extent to which individual genes are essential for cell survival. The Cancer Dependency Map (DepMap) project, in particular, has made significant contributions to drug development by performing large-scale assessments of gene essentiality in hundreds of cancer cell lines. However, relying solely on essentiality data at the individual gene level has limitations. Genes essential for cancer cell survival may also be active in normal cells, leading to toxicity, and experimental noise can result in irrelevant candidates being ranked highly. Furthermore, single-gene information alone cannot fully capture the complex intracellular network interactions. Therefore, systems biology approaches that consider the functional context of genes have emerged as an alternative. Key Findings To address these challenges, Jung J and Yoo S's research team introduced the 'Neighbor-Correlation Essentiality Score (NCES)' framework, which combines cell-specific protein interaction networks. NCES is described as a framework that integrates the CERES score, which is the gene essentiality score from DepMap, with protein-protein interaction (PPI) networks. The researchers derived interaction weights between neighboring genes based on cell-line-specific gene expression patterns and drug response data. Subsequently, they performed accuracy assessments using the Therapeutic Target Database (TTD) and DrugBank gold standard data on seven cancer cell lines. The analysis revealed that NCES model variants consistently outperformed existing methods that analyze only individual gene essentiality. In particular, the 'CRISPR-weighted variant model,' which assigns weights based on gene knockout results, showed remarkable predictive power. This model achieved AUROCs of 0.794 and 0.779 in TTD and DrugBank gold standard predictions, respectively. The weighted NCES model demonstrated a significant improvement in predictive accuracy compared to the basic network model. Furthermore, literature analysis confirmed that several high-ranking genes not included in existing databases, such as CCNB1, CDC7, and WEE1, are either biologically relevant targets or can be inhibited by existing drugs. This demonstrates the potential of NCES to identify novel cancer therapeutic targets. Significance and Prospects This study demonstrates how the combination of big data and network biology can overcome bottlenecks in drug development. Since genes within cells are organically connected and interact, the cell-specific weighted network analysis proposed by the NCES framework is expected to significantly reduce the false-positive rate in target identification. This will reduce the time and cost spent in the target validation phase, thereby increasing the likelihood of successful drug development. Moreover, the strength of NCES in reflecting cell-line-specific characteristics can be a valuable tool for accelerating the realization of personalized precision medicine. By tailoring the selection of the most effective drug targets to the specific cancer type or the genomic environment of individual patients' cancer cells, a pathway can be opened to personalized treatment. The ability to propose novel target genes not listed in existing databases is also a positive factor in diversifying drug pipelines. However, there are some challenges to overcome before NCES can be implemented in actual drug development. The inherent incompleteness of protein-protein interaction networks may introduce errors in the predictions. Furthermore, it is essential to validate in vivo whether the predictions based on cell-line data are consistently reproduced in the complex tumor microenvironment of actual patients. Future research will likely involve integrating clinical data to address these issues.
๐ก This research can be implemented in the early stages of drug development, specifically in the 'Target Identification' phase, to maximize research efficiency. For example, a pharmaceutical research team developing a new drug candidate for a specific rare solid tumor faces the challenge of deciding which of thousands of genes to target. Previously, they relied on random, large-scale screening experiments, which cost hundreds of thousands of dollars and took several months. However, by implementing the NCES framework, they can analyze cell-line-specific protein interaction data to narrow down the most critical 10-20 core target genes for cancer cell death within just a few days. This allows researchers to minimize costly experiments and focus their efforts on the most promising targets. Furthermore, the pharmaceutical industry can link this with companion diagnostics biomarker development to design clinical trials that pre-select patient populations with activated specific gene networks, thereby significantly increasing the success rate of clinical trials.

Background Proteogenomics, which connects genomic information with actual protein levels in the body, has become a key tool for new drug development. However, previous studies have been biased towards populations of European ancestry. This has limited the prevention and treatment of diseases in non-European populations. Different populations have different genetic characteristics, and there was a risk of missing pathological mechanisms that are only activated in specific groups. Plasma proteomics, which precisely analyzes proteins in the blood, has also faced technical barriers. Most previous studies relied on affinity-based analysis methods. This method is fast, but there is a risk of measurement errors if the protein structure changes due to genetic mutations. In addition, the concentration difference between the most and least abundant plasma proteins is 10 billion-fold, making analysis very difficult. As a result, it has often failed to identify key disease target proteins with extremely low concentrations. Key Findings 1,400 Plasma Proteins Captured by Nanoparticles A joint research team, including the Wellcome Sanger Institute, began analyzing blood samples from approximately 1,400 British South Asians (Bangladeshi and Pakistani ancestry) residing in the UK to address the limitations of existing technologies. The researchers used nanoparticles to filter out high-concentration proteins in the plasma and concentrate low-concentration proteins, introducing a novel nanoparticle-enriched mass spectrometry (NP-MS) method. In this process, they successfully detected more than 5,700 protein groups using the Seer Proteograph XT platform. 895 Genetic Variants Revealed a Proteomic Map The researchers compared the obtained proteomic data with genomic information to analyze protein quantitative trait loci (pQTLs), which directly affect protein levels. As a result, they identified more than 1,200 significant genetic variant-protein associations, of which 895 were confirmed as cis-protein quantitative trait loci (cis-pQTLs). Notably, about half of the identified variants are new genomic information not observed in European ancestry studies. This demonstrates the importance of diverse studies that include non-European populations for understanding disease. Filling the Gaps Missed by Existing Platforms The researchers demonstrated the technical complementarity by comparing the same samples using existing affinity-based analysis methods, SomaLogic 11k and Olink HT platforms. Mass spectrometry-based analysis precisely identifies immunoglobulin variable regions or specific protein fragments that are missed by affinity platforms. As a result, it was confirmed that the protein regions detected by each analysis platform are qualitatively different. Uncovering a New Autoimmune Mechanism of Graves' Disease The researchers combined the discovered pQTL data with genome-wide association study (GWAS) results and concluded that 21 proteins are involved in the development of 44 diseases. A representative example is the elucidation of the pathogenesis of Graves' disease, an autoimmune thyroid disease. They revealed that a genetic variant of the IGLV3-21 (Immunoglobulin Lambda Variable 3-21) protein, which has not been clearly elucidated, induces an autoimmune response in B cells and increases the risk of the disease. This is expected to be a useful guide for designing targeted therapies for Graves' disease in the future. Significance and Prospects This study has paved the way for addressing the blind spots in precision medicine by analyzing non-European populations, which have been neglected in genomic analysis. It is widely recognized that the study has refined the identification of customized drug targets by identifying unique protein genetic variants in specific populations. It also presented a practical standard for organically combining mass spectrometry and affinity-based analysis to create a complete plasma proteomic map. Of course, there are still challenges to be solved before it can be applied in clinical practice. Nanoparticle-enriched mass spectrometry has relatively high equipment operating costs and is slower compared to affinity-based platforms. In addition, further validation is required to generalize the results of the 1,400-person cohort to the entire South Asian population worldwide. A follow-up task is a large-scale clinical validation study that includes more non-European populations.
๐ก The pharmaceutical industry can actively utilize this database in the early stages of screening for customized drugs for non-European populations. For example, consider a bio company that is developing new drugs for autoimmune diseases that have a high prevalence in South Asians. In the past, target proteins were selected based on European-centric data, so there was a high probability of failure in clinical trials due to lack of efficacy or unexpected side effects in Asian populations. Now, it is possible to input the identified cis-pQTL information into a drug screening platform to precisely design candidate substances optimized for South Asian patients. Furthermore, in clinical practice, it can be quickly applied as a companion diagnostic technology to diagnose IGLV3-21 genetic variants in a patient's blood and prescribe customized immunomodulatory therapy. This is a shortcut to reduce unnecessary drug administration and increase treatment success rates.

Background Gene therapy has emerged as a next-generation medical technology that targets the root causes of diseases. However, precisely controlling the expression of exogenous genes within target cells is extremely challenging. Excessive gene expression in unintended tissues can trigger severe adverse effects or immune responses. Traditionally, gene switches that utilize antibiotic molecules, such as doxycycline, as regulators have been studied. However, long-term administration of antibiotics can lead to immune rejection or affect unintended tissues, resulting in safety issues. To address this, novel switch technologies that precisely control gene expression have long been of interest. Riboswitches are non-coding RNA molecules that change their structure upon binding to a specific molecule called a ligand, thereby regulating gene expression. This switch operates with RNA alone, without complex regulatory proteins, which efficiently saves the payload capacity of gene therapy vectors. However, redesigning naturally occurring switches found in bacteria for the mammalian cellular environment has been extremely difficult. Artificially engineered RNA structures have struggled to bind specifically to target molecules within complex human cells, making it difficult to achieve a clear on-off response. Key Findings The research team at the University of Navarra recently presented a new milestone in overcoming this impasse in a paper published in the journal 'Molecular Therapy - Nucleic Acids'. They combined artificial intelligence (AI)-based deep learning (DL) models, cutting-edge structural biology techniques such as cryo-electron microscopy (Cryo-EM), and a new high-throughput selection technology to create a platform that rapidly completes the design of customized riboswitches. This platform moves beyond traditional simple sequence analysis and predicts three-dimensional interactions between molecules through computer simulations, increasing the reliability of the switch. In fact, the research team used a DL model to screen billions of RNA library candidates and select the optimal aptamer structure that precisely binds to the target ligand. This structure is designed to operate with a clinically validated, safe ligand that is well-tolerated and easily permeates cell membranes. The research team elucidated the differences between natural switches found in bacteria and eukaryotes and designed a pathway to optimize the expression platform structure to prevent the loss of regulatory signals in mammalian cells. While previous riboswitch models had a difference of only 2-3 times in expression levels when turning genes on and off, the computer-designed next-generation switch showed a strong working efficiency, inducing more than 10 times the gene expression when the ligand was bound. In addition, the selectivity criteria were significantly improved, demonstrating high precision by not reacting at all to structurally very similar non-target substances. Significance and Prospects This research is considered an important milestone in advancing RNA-based therapeutic technologies, which have been limited to proof-of-concept studies at the laboratory level, towards actual clinical-grade therapies. In particular, it is expected to be useful in the next generation of immunotherapies that require precise control of the administered dose and in the treatment of rare diseases that require precise regulation of hormone secretion. However, there are still technical hurdles to overcome before it can be commercialized as a finished product that is directly administered to patients. First, it is necessary to carefully determine whether the customized ligand administered externally is safely degraded and excreted in the body and whether it causes toxicity with long-term use. Furthermore, minimizing the human immune response to the viral vector that delivers the riboswitch is also a challenge to be addressed in the future. Nevertheless, the research team is taking one step further and envisions a 'self-regulating switch' that detects disease-generated metabolites that spontaneously occur in the disease state and activates therapeutic genes without the patient having to take an inducing drug directly. There is growing anticipation that an era of intelligent, autonomous medicine, in which the body responds to internal danger signals and delivers only the necessary amount of therapeutic protein, will soon arrive.
๐ก The commercialization of riboswitch engineering technology will improve the convenience of treatment for patients with chronic diseases. For example, a diabetic patient can finely control the expression of the insulin gene in the body by simply swallowing a pill containing a specific food ingredient. This eliminates the need for daily injections. It also has great application potential in the field of cancer treatment. It is possible to inject a highly toxic anticancer gene and then perform precise controlled therapy by administering a safe inducing substance orally to activate it only in areas where cancer cells are concentrated. The biopharmaceutical industry is also welcoming this development. The complex protein regulatory factors no longer need to be forcibly incorporated into the gene delivery vector, which greatly helps to overcome the limitations of the vector's packaging capacity. It is also very beneficial for simplifying the manufacturing process and reducing production costs.

Background Reading the time of pet dogs with a biological clock Pet dogs share the same living space as humans and are exposed to similar stressors. As a result, they are considered valuable animal models for tracking human aging pathways and understanding disease mechanisms. In recent molecular biology, attention has been focused on biological age, which is measured by changes in the epigenome, rather than chronological age. A key indicator is DNA methylation (DNAm), which regulates gene expression. During the aging process, DNA methylation patterns may change in a specific direction, but at the same time, they also undergo 'epigenetic drift,' which is an increase in random variability. This leads to a phenomenon in which the methylation status of the genome becomes unstable and disordered, which is called epigenetic entropy. However, it has not been clearly elucidated how the rearing country or environmental differences are specifically related to the aging rate and genomic instability of pet dogs. Key Findings Accelerated aging and genomic disorder in Korean pet dogs A research team led by Professor Jae-min Kim of the Department of Animal Life and Biotechnology at Gyeongnam National University, and Professors Seong-rim Lee and Yong-ho Choi of the College of Veterinary Medicine, compared DNA methylation data from 824 pet dogs in Korea and the United States. The research team adopted an epigenetic clock, a pet dog-specific biological age measurement model, as an analysis tool. The biological age recorded a high correlation of r = 0.98 with the actual age. When the aging rate of dogs of the same breed was compared by country, an interesting difference was observed. The epigenetic age acceleration (EAA) of pet dogs in Korea was statistically significantly higher than that of pet dogs in the United States, indicating that Korean pet dogs are biologically aging faster. The research team tracked genomic disorder using a methylation entropy technique. The research team explained that the overall genomic disorder gradually increases. This epigenetic drift follows a different trajectory depending on the size of the pet dog. When comparing groups of small dogs (less than 10 kg), medium-sized dogs (10 to 25 kg), and large dogs (25 kg or more), the rate of entropy increase was steeper for larger dogs. This suggests that the characteristic of short lifespan in large dogs is directly related to the rapid disorder of the genome. In addition, differences in aging trajectories according to sex were also confirmed. The research team newly identified nine aging-related DNA methylation sites that reflect the domestic pet dog environment, and they expect that these sites will serve as indicators of environmental factors. Significance and Prospects Environmental factors leave a biological imprint This study biologically demonstrates that the genomic aging of pet dogs can be greatly affected not only by innate factors but also by acquired living environment. The rearing environment of pet dogs in Korea and the United States differs in various aspects, such as housing type, walking frequency, and feed composition. In particular, it is highly likely that the living characteristics of Korea, where the proportion of indoor living in apartments is high, have affected the epigenome. This is strong evidence supporting the hypothesis in the academic community that environmental stress factors leave traces on DNA methylation marks and regulate the aging rate. In addition, the nine aging-related DNA methylation sites discovered this time are expected to be used as tools for tracking epigenetic changes caused by environmental factors. By using these markers, it will be possible to verify the harmfulness of environmental stress on biological age. However, since it is a retrospective analysis using existing data, a prospective cohort study should follow in the future to observe changes in aging rate by controlling specific environmental factors.
๐ก This study provides practical tools that can be directly applied to the pet industry and clinical veterinary medicine. The most representative application scenario is the development of a 'personalized aging prediction kit' for pet dogs. By utilizing the nine Korean pet dog-specific DNA methylation sites discovered by the research team, it will be possible to accurately measure the biological aging rate of individual pet dogs through a simple blood or saliva test. This is expected to be used as a reference for prescribing customized functional feeds or designing exercise programs for pet dogs with fast aging rates. Furthermore, it can be used as an objective evaluation index to prove the efficacy of anti-aging candidate substances in veterinary clinical trials. This is a method of quantitatively demonstrating at the genome level whether a specific drug or dietary therapy actually slows down the biological age acceleration. In addition, the data from the pet dog aging model will serve as a bridge to accelerate research on preventing premature aging and chronic diseases in humans due to environmental factors.

Background CRISPR gene editing and base editors, core technologies for gene therapies, are considered humanity's weapons to fundamentally overcome rare genetic diseases and intractable inherited diseases. However, the off-target effect, where these therapies randomly edit unintended genomic sequences when acting in real-time within a patient's body, has consistently hindered their safety. Previous research efforts focused on precisely analyzing the 3D conformation of proteins to prevent these errors. These efforts primarily involved designing modifications to regulate the charge at the interface where the protein, guide RNA (gRNA), or DNA interact, aiming to reduce non-specific binding. However, this approach often led to a vicious cycle where both the binding affinity for the original target sequence and the editing activity were simultaneously reduced. Furthermore, the subtle conformational changes in protein structure caused by even a single nucleotide mismatch could not be adequately explained by static structural maps alone. There was a critical need for a new computational biology technique to quantitatively analyze the principles of genomic target recognition at the molecular level through dynamic contact information. Key Findings A joint research team from Peking University and East China Normal University developed ContactSeek, an AI-driven design analysis framework that easily improves the precision of gene editing tools, based on the powerful computational capabilities of AlphaFold3, a protein structure prediction AI. The researchers focused directly on the contact probability (CP) data, a key parameter outputted by AlphaFold3 during the simulation of various nucleic acid binding forms within protein complexes. CP is a metric that quantifies the probability of two constituent residues, amino acids and nucleotides, physically contacting each other in molecular space. The analysis revealed that this metric remarkably sensitively distinguishes between the binding specificity of the on-target site and off-target sites with only slight sequence differences. To apply this structural calculation ability to real-world experiments, the research team combined the deep learning-based contact information prediction results with experimental dI-profiling data, which represents off-target detection genomic sequencing. Through this process, they identified the amino acid groups that ultimately determine the non-specific binding of the editing tool and named them Consensus Contact Regions (CCRs). CCRs are functionally similar to a variable resistor switch that controls the precision of gene editing. By manipulating this switch, fine-tuning the residues within the CCRs to match the binding affinity of a specific genetic locus, it is possible to create optimal protein variants that preserve activity at the original target site while almost completely blocking unwanted off-target binding. In fact, the research team successfully constructed several improved models of adenine base editor (ABE) and Cas12a cytosine base editor (CBE), the most widely used gene editing tools, with significantly reduced error rates by applying this design system. Significance and Prospects The emergence of ContactSeek represents an evolution of AI, shifting its role from simply focusing on static structural analysis to a precise design tool for novel protein engineering. Previous protein modification experiments required a laborious screening process involving thousands of random mutations, followed by individual experimental validation of activity. In contrast, ContactSeek accurately identifies the favorable interaction sites between proteins and genetic sequences through computational science, reducing the experimental validation time for candidate materials by dozens of times. However, there are additional challenges to overcome before this AI-driven design tool can be successfully translated into clinical drug development. It is necessary to demonstrate in various ways whether the contact probabilities derived in molecular virtual space are consistently expressed in the complex ionic states and protein crowding environments within actual cell cytoplasm. The potential for immune responses induced by structural changes and compatibility with in vivo delivery systems are also challenges to be addressed. The research team plans to expand the target gene editing tools and focus on in vivo pharmacological activity validation in the future.
๐ก This research provides a practical roadmap to significantly accelerate the commercialization of gene therapies. In particular, it can dramatically improve safety in the development of personalized therapies for patients with inherited diseases. Specifically, in the case of diseases caused by a single nucleotide mutation, such as Huntington's disease or inherited retinal degeneration, it will be possible to design personalized base editors in real-time to minimize off-target effects. This involves acquiring off-target risk sequence information from patient-derived cells, inputting it into the ContactSeek analysis system, and inducing the creation of a custom Cas protein that is designed to respond specifically to that sequence. This is expected to significantly increase the likelihood of clinical approval for gene editing by eliminating the risk of the therapeutic agent targeting the wrong gene and causing cancer.

Background Staphylococcus aureus (S. aureus) is a major pathogenic bacterium that poses a significant global public health threat. The increasing prevalence of antibiotic-resistant strains, including Methicillin-Resistant Staphylococcus aureus (MRSA), has highlighted the limitations of existing treatments. Invasive S. aureus infections can lead to life-threatening complications such as sepsis and septic arthritis; however, no licensed vaccine is available to prevent these infections. Previous attempts to develop single-antigen vaccines have been unsuccessful due to the complex immune evasion mechanisms and diverse virulence factors of the bacterium. A new platform capable of simultaneously targeting multiple antigens to induce a multifaceted immune response is therefore needed. Digital design frameworks, which can dramatically accelerate antigen information analysis and vaccine design, are now attracting considerable attention. Key Findings The researchers constructed an integrated in silico pipeline to design a bivalent self-amplifying messenger ribonucleic acid (saRNA) vaccine cocktail targeting Clumping factor A (ClfA), Alpha-hemolysin (Hly), and Staphylococcus aureus receptor for platelets (SraP), which are key virulence antigens of S. aureus. This approach combines immunoinformatics, structural bioinformatics, and molecular simulation techniques to precisely analyze the immunodominant epitope regions of each antigen. Based on this data, two saRNA vaccine candidates, SaBVax807 and SaTVax876, were developed. The researchers used molecular docking and molecular dynamics (MD) simulations to verify the binding interactions between the candidate vaccines and human leukocyte antigen (HLA) alleles. To assess the global utility of SaTVax876, they also conducted an analysis of data from 16 geographically diverse regions. In addition to codon optimization to enhance translational efficiency in human host cells, the candidate vaccines were also evaluated for safety, including the potential for allergic and autoimmune reactions, and were found to be highly safe. Predictive models indicate that both candidate vaccines possess the physicochemical properties necessary to maintain structural stability in vivo and elicit a strong immune response. Significance and Future Directions This study is significant because it presents a new principle for the design of vaccines against bacterial infections using advanced computer modeling techniques. The developed saRNA vaccine replicates the genetic information of the antigen protein, thereby inducing a strong immune response even with a small dose. In the context of the increasing problem of multidrug resistance (MDR) in S. aureus, this study provides a new non-antibiotic prophylactic strategy. However, since the results are based on computer modeling, they need to be cross-validated with in vitro and in vivo preclinical animal studies. The complex dynamic changes within the actual immune system may deviate from the simulation predictions. The complex production process and formulation stability of a cocktail vaccine targeting multiple antigens are also challenges that need to be overcome before clinical application. Why It Matters This study provides a concrete pathway to significantly shorten the development time for vaccines against antibiotic-resistant bacteria. In clinical practice, it may be possible to administer the SaTVax876 cocktail prophylactically to chronic infection patients who cannot receive conventional antibiotics or to immunocompromised patients undergoing surgery, thereby reducing the incidence of surgical site infections. From a vaccine industry perspective, the adoption of a bivalent saRNA formulation that targets multiple variant antigens in a single dose can mitigate the production complexity associated with loading multiple antigens. Furthermore, the self-replicating mechanism reduces production costs and enables the establishment of a rapid large-scale production system, which is expected to be a practical solution for the widespread distribution of vaccines in developing countries with limited medical resources.
๐ก This study provides a concrete pathway to significantly shorten the development time for vaccines against antibiotic-resistant bacteria. In clinical practice, it may be possible to administer the SaTVax876 cocktail prophylactically to chronic infection patients who cannot receive conventional antibiotics or to immunocompromised patients undergoing surgery, thereby reducing the incidence of surgical site infections. From a vaccine industry perspective, the adoption of a bivalent saRNA formulation that targets multiple variant antigens in a single dose can mitigate the production complexity associated with loading multiple antigens. Furthermore, the self-replicating mechanism reduces production costs and enables the establishment of a rapid large-scale production system, which is expected to be a practical solution for the widespread distribution of vaccines in developing countries with limited medical resources.

Background Genome-Wide Association Studies (GWAS) have become a cornerstone for identifying genetic variants associated with diseases, particularly with the rise of large biobanks. However, most existing GWAS studies have been limited by their focus on European populations. Analyzing admixed populations, such as those of African or Hispanic descent, presents challenges due to population structure bias, which can lead to a high number of false positive signals. To address these limitations, the statistical algorithm 'Tractor,' which incorporates Local Ancestry Inference, was introduced in 2021. Tractor successfully classifies ancestry at specific chromosomal regions, enabling a more precise estimation of the impact of ancestry-specific variants. However, the original Tractor model assumes that there is no relatedness between the individuals being analyzed. Modern large biobanks, such as the Mexico City Prospective Study cohort and the UK Biobank, often contain samples with familial relationships. Excluding related individuals arbitrarily leads to significant loss of genomic data and reduced detection power, while including them without adjustment leads to statistical errors. This has been an ongoing dilemma. Key Findings A team of American researchers has developed a new algorithm, 'Tractor-Mix,' that performs precise analysis even in populations with complex familial relationships and mixed ancestry. The research findings were published in the latest issue of the international journal Nature Genetics. Tractor-Mix is a framework that integrates a Linear Mixed Model (LMM) with local ancestry inference. It incorporates a Genetic Relationship Matrix (GRM), which quantifies the degree of genetic relatedness between samples, into the mathematical model. This allows for the accurate calculation and correction of statistical biases arising from family relationships. Simultaneously, the algorithm tracks the ancestry of each chromosomal segment in parallel. The research team validated their hypothesis using data from the Mexico City cohort and the UK Biobank. The results showed that Tractor-Mix effectively controlled the false positive rate and significantly improved the detection power for variants enriched in specific ancestries, without excluding related samples. In particular, it accurately quantified the effect size of rare disease-related variants in non-European admixed populations without loss of genetic relative data. Significance and Future Directions This research establishes a foundation for expanding genomic research beyond its Eurocentric focus to include more diverse populations. It provides a statistical framework that allows for the full utilization of clinical genomic data from cohorts with complex familial relationships, such as those found in South America and Africa. The researchers plan to conduct follow-up work to improve the computational speed of the algorithm. Due to the simultaneous processing of local ancestry inference and large-scale genetic relationship matrix calculations, the computational burden increases when processing datasets of hundreds of thousands of individuals. Improving computational efficiency through the implementation of a cloud-based distributed computing system is a key priority.
๐ก The introduction of Tractor-Mix is expected to bring about direct changes in the development of personalized medicine and targeted therapies for diverse populations. Until now, pharmaceutical companies and research institutions have relied on data from European populations when identifying genetic variants associated with drug response and disease risk. This has often led to inaccuracies in target identification and the calculation of Polygenic Risk Scores (PRS) for diverse patient populations. In the future, it will be possible to analyze global population databases with diverse genetic backgrounds and familial relationships without loss of information. This will accelerate the discovery of disease-causing variants specific to certain ethnicities or admixed populations, and provide practical benefits for the selection of participants in global clinical trials and the development of personalized treatment strategies.

Background The limitations of gene editing tools that have evolved over hundreds of millions of years in nature. The CRISPR-Cas9 system has become a representative gene editing tool that precisely cleaves and corrects specific DNA within cells in gene therapy and agricultural biotechnology. However, directly utilizing naturally occurring Cas proteins as therapeutic agents or research tools has practical limitations. This is because their large molecular size makes it difficult to load them into adeno-associated viruses (AAV), which are used for in vivo delivery, or they may cause off-target effects. Researchers have attempted to improve the performance of naturally occurring gene editing tools by making minor modifications. Existing protein engineering techniques have been limited to sequence modifications that replace only 1-2% of amino acids in the native sequence. Excessive sequence changes could disrupt the three-dimensional structure of the protein, potentially leading to a complete loss of its catalytic activity. As a result, CRISPR technology has faced the limitation of being improved only within the range of natural gene resources. Key Findings SynTnpB, a protein redesigned based on artificial intelligence. A research team led by Professor Jennifer Doudna at the University of California, Berkeley, overcame this limitation by introducing generative artificial intelligence (AI) technology. The team utilized an inverse protein-folding AI model that takes a desired three-dimensional protein structure and function as input and calculates the optimal amino acid sequence in reverse. The target of this study was TnpB, an evolutionary precursor of the Cas12 protein and a small RNA-guided cleavage enzyme. The AI model designed a large number of synthetic variants (SynTnpBs) with completely new sequences based on the three-dimensional structure of natural TnpB. Surprisingly, the amino acid sequence of the synthetic enzymes designed by AI differed by up to 30% compared to natural TnpB. The research team applied the generated synthetic enzymes to bacteria, plants, and human cell lines to verify their actual gene editing ability. The results showed that SynTnpB exhibited higher cleavage efficiency than natural TnpB, and it also demonstrated excellent precision in accurately recognizing the target DNA sequence. This not only surpassed the 1-2% modification limit of existing protein engineering but also demonstrated that artificially redesigned proteins with large-scale sequence modifications can exhibit much better function in vivo. Significance and Prospects A new milestone in biotechnology and therapeutic development. This study shows that the focus of gene editing technology has shifted from the discovery of natural products to AI-driven custom design. The small and highly precise synthetic TnpB is easy to load into AAV viral vectors. This directly addresses the dose limitation problem that occurs when delivering therapeutic agents for intractable genetic diseases into the body. However, there are still challenges to overcome before it can be commercialized as an actual therapeutic agent. The immunogenicity of the non-natural amino acid sequences generated by AI when administered to the human immune system needs to be carefully evaluated. In addition, follow-up clinical studies are essential to ensure long-term stability in various tissues and biological environments. The research team plans to improve the AI model in the future to develop next-generation gene editing tools that minimize immune response and further enhance editing precision.
๐ก The small gene editing tool SynTnpB, designed by AI, provides concrete solutions in the fields of biopharmaceutical therapeutic development and the creation of gene-edited crops. The existing Cas9 enzyme consists of more than 4,000 amino acids, making it difficult to load into AAV for gene therapy. In contrast, SynTnpB has a significantly lower molecular weight, allowing it to load the gene editing tool, target guide RNA, and therapeutic gene fragment into a single viral vector. This characteristic greatly improves the therapeutic efficacy of intractable diseases that require in vivo gene delivery, such as retinitis pigmentosa and muscular dystrophy. In the field of plant biotechnology, it can also significantly shorten the breeding period by minimizing off-target mutations when inducing desired agricultural traits. Custom protein design technology that overcomes the limitations of natural proteins is expected to become a core platform in the next generation of biotechnology.