Three-Dimensional Mass Spectrometry Enables Comprehensive Analysis of RNA Therapeutics, Including Minor Impurities and Methylation Variants

Background
The rapid expansion of the RNA therapeutics market has increased the difficulty of quality control, which requires verifying the purity and sequence accuracy of the final product. Small interfering RNAs (siRNAs) and single guide RNAs (sgRNAs) used in CRISPR/Cas9 systems are frequently subject to chemical modifications. Subtle variations and trace amounts of impurities at the molecular level can reduce therapeutic efficacy and potentially trigger unexpected immune responses, necessitating precise analysis.
The existing bioindustry has used liquid chromatography-tandem mass spectrometry (LC-MS/MS) to verify RNA sequences. However, existing LC-MS/MS analysis is limited to confirming the presence or absence of a match with a pre-designed target sequence. It is impossible to perform de novo sequence determination, which decodes the sample information from beginning to end without prior knowledge. As a result, there is a limitation in that subtle variations or unexpected base modifications at unexpected locations are not detected during the analysis process and are simply overlooked.
Key Findings
Recently, the academic community has proposed a new type of sequencing platform to overcome these analytical limitations. The three-dimensional next-generation mass spectrometry-based sequencing (3D NGMS-Seq) technology developed by the researchers analyzes complex mixed RNA samples and directly reads the sequence with nearly 100% accuracy.
The researchers combined the mass value and retention time (tR), which are the existing two-dimensional data structure, with 'MS intensity' as the third variable. The RNA sample, which has undergone controlled acid hydrolysis, is broken down into ladder fragments of various lengths and dispersed into a three-dimensional layer of mass-intensity-tR. In the complex and mixed data, the researchers applied a nested algorithm that aligns the abundance of the parent RNA with the signal intensity of the cleaved fragments to computationally separate individual RNA components.
The fragments separated in each layer are subjected to base-calling, one base at a time, based on the precise mass difference between adjacent fragments. Standard nucleotides, as well as chemically modified methylated bases, are sequentially identified and then reassembled into the full-length RNA sequence. This technology has demonstrated excellent performance in validation experiments using synthetic siRNAs, microRNAs (miRNAs), and sgRNAs. In particular, it clearly resolved the positional ambiguity of subtle methylation isomers, Um and mU, and Am and mA, which were almost impossible to identify with existing analytical methods. Furthermore, it achieved the result of quantifying the relative amount of each RNA in the mixture and the ratio of modifications at specific locations as numerical data.
Significance and Prospects
This study is significant in that it lays the foundation for precise analysis of unknown RNA mixtures without prior input of designed sequence information. Drug developers can use this technology to transparently monitor trace amounts of structural isomers and impurities from the early stages of candidate substance selection. This is expected to be of great help in meeting the strict approval criteria of regulatory agencies, which require complete control of the chemical composition of complex oligonucleotide drugs.
However, some technical improvements are required before it can be fully introduced into clinical and commercial manufacturing sites. The current three-dimensional computational separation algorithm may be a bottleneck in terms of processing speed and analysis throughput for real-time analysis of large amounts of samples in large-scale production processes. There is also a challenge in improving the sequencing algorithm to maintain decoding accuracy when analyzing long-chain RNA molecules with a high density of complex chemical modifications.
The rapid growth of RNA-based therapeutics demands accurate sequencing of all RNA species, including minor and modified variants. Conventional LC-MS/MS typically confirms only a predefined target sequence rather than determining RNA sequences de novo from the analyzed sample, thereby overlooking coexisting impurities and modifications. Here, we present 3D NGMS-Seq, a three-dimensional next-generation mass spectrometry-based sequencing platform for de novo direct sequencing of mixed RNA samples with essentially 100% sequence accuracy. This method incorporates MS intensity into traditional 2D mass-retention time (tR) analysis and introduces a nested algorithm that aligns ladder fragment intensities with parent RNA abundances for computational separation. Controlled acid hydrolysis produces RNA ladder fragments, which are segregated into mass-intensity-tR layers. Within each layer, short reads are generated de novo by sequentially base-calling each nucleotide, canonical or modified, from mass differences between adjacent ladder fragments and subsequently assembled into full-length RNA sequences. Guided by hydrolysis kinetics and statistical modeling, 3D NGMS-Seq accurately sequences synthetic siRNA, miRNA, and CRISPR/Cas9 sgRNAs, reveals unexpected low-abundance RNA impurities, and resolves subtle methylation ambiguities (Um versus mU; Am versus mA), while providing a quantitative profile of each RNA's relative abundance and site-specific modifications. By enabling direct, unbiased sequencing of heterogeneous RNAs without prior sequence knowledge, 3D NGMS-Seq addresses key limitations of current RNA analysis and provides a powerful tool to aid small RNA drug development, quality control, and regulatory validation.
This technology can be immediately applied to automate the quality control (QC) steps of the bio-pharmaceutical manufacturing process and maximize safety. For example, in the production process of CRISPR gene therapy, if the distribution of chemical modifications of the guide sgRNA is uneven, a fatal adverse effect called off-target gene cleavage occurs. Pharmaceutical companies can use 3D NGMS-Seq to operate a process in which samples are collected immediately after the synthesis process and the base sequence errors and methylation modification patterns of the guide RNA are thoroughly investigated. Even without prior sequence information, it precisely identifies unexpected low-concentration impurities derived during the synthesis process, which functions as a powerful manufacturing process control tool to minimize batch-to-batch quality variations and prevent the release of products containing potentially adverse impurities.