Avoid Using "PacBio SMRT" and "HiFi" Interchangeably
Why Is This Important?
Short-read sequencing accurately reads many bases but struggles to link the phase of variants located far from repetitive sequences or large structural variants. PacBio SMRT sequencing generates long reads by observing the synthesis of a single DNA molecule in real time, and the current HiFi approach simultaneously pursues length and high consensus accuracy.
Using PacBio, long read, and HiFi interchangeably obscures data characteristics across generations. The past CLR technology and current HiFi differ in read generation, error profiles, and suitable analysis pipelines.
Events within the ZMW
SMRT sequencing observes a single DNA polymerase within a microstructure known as a zero-mode waveguide (ZMW). The nucleotide sequence is determined by detecting signals from fluorescently labeled nucleotides during synthesis. By constructing a circular template with hairpin adapters attached to both ends of a DNA fragment, the polymerase can traverse the same insert multiple times.
A single traversal is referred to as a subread, and a CCS (Circular Consensus Sequence) is generated by combining information from multiple traversals into a consensus sequence. A CCS with sufficient pass count and signal quality serves as the foundation for HiFi reads.
CLR and HiFi
CLR focuses on obtaining very long continuous reads but has a relatively high single-read error rate, requiring sufficient coverage or correction. HiFi improves per-read accuracy by repeatedly reading inserts to generate a consensus. However, if the insert length is excessively long for the same source material, the number of passes decreases, creating a trade-off between consensus quality and other factors.
Therefore, experimental selection is not based solely on maximum read length. Instead, one must consider the molecule length, per-read accuracy, yield, and coverage required by the research question.
What is more visible?
Long reads connect across repeats to unique flanking sequences, increasing the likelihood of observing deletions, insertions, inversions, and complex SV breakpoints on a single molecule. Linking multiple heterozygous variants within a single read also facilitates haplotype phasing. Long-range context offers advantages in de novo assembly and isoform sequencing as well.
However, not all repeats are shorter than the reads, and segmental duplications can still complicate mapping. Performance for low-frequency mosaic variants and small variants depends on depth, caller choice, and assay validation.
Small Example
Suppose a large insertion between two exons is suspected in short reads, but the breakpoint lies within repetitive sequences, preventing precise alignment. If sufficiently long HiFi reads span both flanking unique sequences and the entire insertion, the event structure and allele phase can be clearly resolved. Conversely, if the input DNA is highly fragmented, even long-read platforms may fail to obtain the necessary spanning connections.
Experimental Design and QC
High-molecular-weight DNA extraction, shearing, and size selection determine read length. Library yield, polymerase read length, insert-size distribution, HiFi yield, read quality, and coverage should be assessed concurrently. Sample type and storage history significantly impact DNA integrity.
The analysis must document the reference build, aligner/assembler, SV/small-variant caller, and their versions. Changes in chemistry and instrument software also affect comparability with historical data.
Common Misconceptions
- Long reads do not automatically imply high accuracy; HiFi generation conditions must be verified.
- High read N50 does not guarantee sufficient coverage at all loci.
- Consensus sequences do not eliminate both systematic biases and sample preparation artifacts.
- Events identified by long reads may require orthogonal confirmation and validation for clinical use.
Interpretation Boundaries
Utility depends on input molecule quality, insert size, coverage, chemistry, and software. Do not assume superiority over existing assays based solely on the platform name; instead, validate performance per question.
Reading in Context
The homologous region problem leads to the story of pseudogene mapping, assay selection as WGS, WES, and panels, and references and assemblies evolving from the Human Genome Project to pangenomes.