# Chapter 128. Long-Read, Full-Length, and Direct RNA Sequencing

## Scope Note

This chapter explains sequencing technologies that preserve transcript-scale linkage or native-RNA information lost in ordinary short-fragment RNA sequencing. It covers PacBio full-length transcriptomics, nanopore cDNA sequencing, nanopore direct RNA sequencing, modification-sensitive native signals, isoform- and allele-resolved analysis, platform-specific error models, validation, quantification, and short-read integration. Nascent RNA profiling, metabolic labeling, pulse-chase designs, RNA-lifetime estimation, and kinetic transcriptomics are owned by [Chapter 129](chapter1154.md); they appear here only as cross-chapter handoffs when a long-read experiment is combined with a kinetic assay.

## Executive Summary

Short-read RNA-seq counts short fragments very well, but fragmentation breaks the physical linkage among transcription start sites, splice junctions, retained introns, RNA editing or sequence variants, polyadenylation sites, and RNA damage or modification states. Long-read and full-length transcriptomics restore part of that linkage. PacBio full-length transcriptomics generally sequences cDNA molecules with high-accuracy circular consensus reads and is particularly strong for transcript annotation. Nanopore cDNA sequencing reads long cDNA molecules through nanopores and is flexible for isoform discovery, targeted panels, and samples where amplification is useful. Nanopore direct RNA sequencing reads native RNA molecules, usually with lower throughput and stronger protocol-specific biases, but it avoids reverse transcription for the sequenced strand and can retain information about native RNA chemistry, tails, damage, and viral RNA architecture.

The value of a long read is conditional on what entered the library and what survived analysis. A read can connect a transcription start site, splice pattern, sequence variant, editing event, and polyadenylation site on one molecule, yet still be truncated, internally primed, chimeric, misaligned, or drawn from a biased subset of the transcriptome. Long-read discovery therefore needs confidence tiers rather than a binary annotated-versus-novel label. Structural existence, regulated abundance, protein production, disease mechanism, and clinical actionability are separate claims with separate validation requirements.

The central practical lesson is that each platform and analysis layer has a distinct error model. Long-read cDNA isoforms can be truncated, internally primed, chimeric, under-sampled, or over-collapsed. Direct RNA reads can reveal native signal but can also be shaped by RNA degradation, adapter ligation, pore motor kinetics, sequence context, basecaller training, and modification-like artifacts. Allele assignment can be distorted by reference bias and paralogous mapping, while transcript quantification can become non-identifiable when isoforms share most sequence. Strong studies therefore match the assay to the question, distinguish discovery from quantification, preserve molecule-level uncertainty, and use orthogonal validation for high-impact claims.

## Concept Inventory

- **Long-read RNA sequencing:** sequencing of RNA-derived or native RNA molecules with reads long enough to span multiple transcript features. Long reads are not automatically full-length, complete, or quantitatively unbiased.
- **Full-length transcriptomics:** workflows designed to recover complete or near-complete transcript structures, commonly by sequencing cDNA molecules that show evidence of both 5′ and 3′ primer-derived ends. A full-length library molecule is not always proof of the complete native RNA.
- **Isoform:** a transcript variant from a locus, produced by alternative transcription initiation, splicing, intron retention, RNA editing, alternative 3′ end formation, or combinations of these features.
- **Direct RNA sequencing:** sequencing of native RNA without converting the sequenced strand into cDNA. In this chapter the term primarily refers to nanopore direct RNA sequencing.
- **Modification-sensitive direct RNA signal:** a change in raw current, dwell time, basecalling error, alignment residual, or mismatch pattern that may reflect a chemical modification but requires calibration before chemical identity is asserted.
- **Isoform-resolved analysis:** analysis that assigns features to molecule-level transcript structures rather than counting exons, junctions, or ends independently.
- **Allele-resolved analysis:** analysis that distinguishes transcripts from different alleles, haplotypes, parental chromosomes, strains, or viral variants.
- **Context-aware transcript quantification:** transcript discovery and abundance estimation against a catalog tailored to the sampled biological context; its accuracy depends on annotation completeness, discovery thresholds, and how ambiguous reads are assigned.
- **Short-read integration:** use of short-read RNA-seq or targeted assays with long-read and direct RNA data to improve depth, quantification, reproducibility, and evidence grading.

## What to Know Before Reading This Chapter

The prerequisite from [Chapter 125](chapter1116.md) is the basic RNA-seq workflow: RNA is extracted, selected or depleted, converted into a sequencing library, sequenced as reads, aligned or assigned to references, and quantified with statistical models. The special problem for [Chapter 128](chapter1117.md) is that many biological questions cannot be answered from short independent fragments. A 100-nucleotide read can support one splice junction, one variant, or one segment of an untranslated region, but it usually cannot prove that several distant features occur on the same molecule. A 5-kilobase long read may connect those features directly, although only for molecules successfully captured by the protocol.

The second prerequisite is the distinction between a biological transcript and a library molecule. A transcript is an RNA produced and processed in a cell. A library molecule is the physical nucleic acid that reaches the sequencer after extraction, selection, reverse transcription when applicable, amplification when applicable, adapter ligation, and size-dependent recovery. Long reads improve linkage within recovered molecules; they do not recover molecules that were lost or prove that a library molecule preserved both native ends.

This chapter uses three running examples. The first is a human gene with alternative first exons, cassette exons, a retained intron, a heterozygous variant, and alternative polyadenylation. The second is a viral infection in which copyback or defective viral RNAs contain rearranged junctions that must be validated on long molecules. The third is a direct RNA experiment in which a nanopore current shift might indicate an RNA modification but might also reflect sequence context or model error. These examples separate three questions that are often conflated: which features coexist on one transcript, which allele or viral lineage produced that transcript, and which native-molecule property generated an electrical signal.

## 128.1. PacBio full-length transcriptomics

PacBio full-length transcriptomics is a long-read cDNA strategy for recovering transcript structures molecule by molecule. The common logic is to reverse transcribe RNA into cDNA, identify cDNA molecules that carry evidence of 5′ and 3′ library adapters or primers, sequence those molecules with single-molecule real-time sequencing, and use repeated observations of the same circularized insert to build high-accuracy consensus reads. Current PacBio transcript workflows are often described with the vendor term Iso-Seq, but the generic scientific concept is full-length cDNA transcript sequencing.

The key term "full-length" needs an early boundary. In many workflows, a read is called full-length because the cDNA molecule contains both expected primer-derived ends and a poly(A)-associated end. That designation does not prove that the native RNA molecule was captured from the true transcription start site to the true cleavage and polyadenylation site. Reverse transcriptase can stop early at structured regions, damaged bases, long homopolymers, or modified nucleotides. Template switching can connect segments that were not contiguous in the original RNA. Internal priming at A-rich genomic regions can mimic a poly(A)-anchored 3′ end. Size selection can enrich mid-sized cDNAs while losing very short, very long, or fragile molecules. Therefore, the phrase full-length should be read as a library and analysis claim unless transcript-end evidence is independently established.

The biological payoff is linkage. Consider a neuronal gene with two promoters, three cassette exons, one retained intron, a disease-associated single-nucleotide variant, and two polyadenylation sites. Short-read RNA-seq may detect each junction and estimate local inclusion levels, but the short reads may not reveal whether promoter 1 is linked to exon skipping and distal polyadenylation on the same transcript. A PacBio full-length read can sometimes show the whole combination. That information changes interpretation: one isoform may encode a different protein domain, another may introduce a premature termination codon and become sensitive to nonsense-mediated decay, and a third may carry a long 3′ untranslated region with localization or microRNA regulatory elements.

![Figure 128.1. Long-read isoform linkage versus short-read fragmentation](../assets/figures/chapter1117_figure1.png)

**Figure 128.1. Long-read isoform linkage versus short-read fragmentation.** Short reads can support individual junctions, variants, and ends, whereas a long read can connect those features on one recovered molecule. Linkage does not guarantee native end completeness, unique transcript assignment, or population-level allelic imbalance.

PacBio full-length transcriptomics is especially useful for annotation refinement. It can discover unannotated exons, alternative first exons, alternative last exons, retained introns, transcript fusions, and combinations of splice junctions that were impossible to assemble confidently from short reads. In organisms with incomplete transcript annotations, long cDNA reads can provide the first coherent transcript models for many genes. In well annotated mammalian genomes, long reads are still valuable because reference annotations are not complete across cell types, developmental states, disease states, and perturbations.

The workflow has several computational steps that affect biological conclusions. Reads are demultiplexed, adapter and primer sequences are removed, poly(A) tails may be trimmed or measured, full-length non-chimeric reads are identified, and similar molecules are clustered or collapsed into transcript models. Alignment to the genome must tolerate splice junctions but reject implausible junctions caused by sequencing, mapping, or library artifacts. Transcript collapsing must decide whether two reads represent the same isoform despite minor end variation or incomplete coverage. These decisions are not purely technical; permissive settings can inflate the isoform catalog, while conservative settings can erase real cell-type-specific isoforms.

**Table 128.1. Evidence tiers for new isoforms.** Separate candidate structures from supported, end-validated, quantitative, and functional claims.

| Evidence tier | Typical evidence | Appropriate claim | Important next check |
| --- | --- | --- | --- |
| **Candidate** | One or a few long reads span a plausible feature combination. | Candidate transcript model observed. | Replication, junction inspection, and artifact screening. |
| **Structurally supported** | Independent reads and junction evidence support a common chain. | Supported isoform structure in the sampled material. | Independent library or targeted amplification. |
| **End-supported** | Long reads agree with cap-aware 5′ and independent 3′ end evidence. | End-supported transcript model. | Exclude degradation and internal priming. |
| **Quantitatively relevant** | Reproducible abundance is supported by adequate molecules or integrated data. | Isoform contributes measurably under stated conditions. | Report ambiguous assignment and confidence intervals. |
| **Functional** | Structural and quantitative evidence is paired with translation, localization, perturbation, disease, or rescue evidence. | Isoform has evidence for a specified role. | Match follow-up to the biological claim. |

Evidence for a new isoform should be graded. A low-abundance transcript model supported by one long read is a discovery candidate. A model supported by independent long reads, short-read junction coverage, reproducible samples, plausible splice motifs, cap-aware 5′ end evidence, 3′ end cleavage evidence, and absence of internal priming is stronger. A protein-coding consequence requires additional evidence, such as an intact open reading frame, ribosome profiling support, proteomic peptides, conservation, or perturbation data. A noncoding regulatory consequence requires still different evidence, such as localization, abundance, interaction partners, perturbation-rescue logic, or disease association with causality controls.

> **Box 128.1. Questions to ask before accepting a novel isoform**
>
> Ask what one molecule proves, whether independent molecules reproduce it, whether 5′ and 3′ ends are credible, whether abundance is estimable, and which additional evidence supports coding, regulatory, disease, or clinical claims. Use tiered language rather than calling every observed molecule a functional isoform.

Allele-specific transcript structure is a special strength when long reads span heterozygous variants and isoform-defining features. A read that carries a phased single-nucleotide variant and a splice pattern can show allele-specific splicing directly. In human genetics and cancer transcriptomics, that can connect a cis-regulatory variant, a splice-altering mutation, or a structural variant to an expressed isoform. Boundary cases remain common: reference bias can favor the reference allele, paralogous loci can attract misaligned reads, and low coverage can make allele imbalance look stronger than it is. Personalized or graph transcriptomes can reduce but not eliminate these issues.

Evidence status: Oikonomopoulos et al. (2020) provides a verified local review of long-read transcript-profiling methodologies, and Pardo-Palacios et al. (2024) provides a verified systematic assessment of transcript identification and quantification. A publication-ready history of PacBio Iso-Seq and HiFi development still needs a dedicated primary method anchor, and transcript-end claims need dedicated end-validation studies.

## 128.2. Nanopore cDNA and direct RNA sequencing

Oxford Nanopore sequencing reads nucleic acids by measuring changes in ionic current as a strand passes through a pore under the control of a motor enzyme. For RNA studies, two broad workflows matter. Nanopore cDNA sequencing first converts RNA to complementary DNA and sequences the cDNA. Nanopore direct RNA sequencing adapts native RNA and sequences the RNA strand itself. Both can produce long reads, but they answer different questions because they sequence different molecules.

Nanopore cDNA sequencing is useful when the primary aim is transcript structure, scalable discovery, or sample compatibility. Because cDNA can be amplified, cDNA workflows can work with lower inputs and targeted panels more readily than current direct RNA workflows. cDNA reads can connect distant exons, identify fusion transcripts, support long isoform models, and complement PacBio full-length transcriptomics. But cDNA sequencing inherits reverse-transcription artifacts. Structured RNA regions can be copied inefficiently; modified bases can cause misincorporation or stops; template switching can create artificial chimeras; polymerase chain reaction can distort molecule counts; and strand or end information depends on library design.

Direct RNA sequencing is conceptually different because the native RNA molecule is the sequenced strand. In common poly(A)-selected workflows, an adapter is ligated near the 3′ end and the molecule is pulled through the pore in a defined orientation. The read may include transcript sequence, poly(A) tail information, damage signatures, and modification-sensitive signal. Direct RNA sequencing therefore has distinctive use cases: native RNA modification discovery, poly(A) tail analysis, viral RNA genome reconstruction, RNA degradation studies, and comparisons between native RNA and cDNA-derived artifacts.

![Figure 128.2. Nanopore cDNA and direct RNA workflow comparison](../assets/figures/chapter1117_figure2.png)

**Figure 128.2. Nanopore cDNA and direct RNA workflow comparison.** Nanopore cDNA and direct RNA workflows both produce long reads but preserve different information. cDNA supports amplification; direct RNA retains native signal while introducing distinct loading, integrity, and throughput constraints.

The main direct RNA advantage is also its main interpretive trap. Avoiding reverse transcription does not make the experiment unbiased. Direct RNA sequencing depends on RNA integrity, adapter ligation, end accessibility, poly(A) selection or alternative capture design, pore chemistry, motor behavior, basecalling model, and alignment strategy. Many workflows show a 3′ coverage bias because molecules are loaded from the 3′ end and long or damaged RNAs may not be read completely to the 5′ end. Some RNA classes are missed because they lack poly(A) tails, carry blocked ends, are too short for the workflow, are highly structured, or remain tightly bound in RNPs during extraction.

Direct RNA sequencing can be especially powerful in viral systems. Viral RNA populations may include full-length genomes, subgenomic RNAs, defective interfering RNAs, copyback genomes, recombination products, and host-virus chimeric artifacts. Short reads can identify junctions, but they often cannot prove which termini and junctions belong to the same molecule. The local [Chapter 128](chapter1117.md) bibliography contains Pye et al. (2025), a direct RNA sequencing study validating diverse Sendai virus copyback viral genomes. That reference supports a narrow but important principle: native long reads can validate viral RNA structures that are difficult to reconstruct from fragmented evidence. The same logic applies more broadly to viral RNA biology, but the local bibliography does not yet support broad claims across all virus families.

**Table 128.2. Comparison of long-read RNA workflows.** Compare workflows by molecule, question, artifact, and validation.

| Workflow | Molecule sequenced | Strongest questions | Major limitations | Validation anchors |
| --- | --- | --- | --- | --- |
| **PacBio full-length cDNA** | High-accuracy consensus cDNA. | Transcript annotation, splice-chain linkage, alternative ends. | Reverse-transcription stops, template switching, internal priming, size selection, collapsing choices. | Junction support, end mapping, replicate libraries. |
| **Nanopore cDNA** | Long cDNA, sometimes amplified. | Flexible discovery, fusions, targeted panels. | Reverse-transcription and amplification bias, chimeras, basecalling and alignment error. | Independent library, short-read junctions, fusion controls. |
| **Nanopore direct RNA** | Native RNA, commonly loaded from the 3′ end. | Native signal, tails, damage, viral architecture, modification-sensitive candidates. | RNA integrity, adapter ligation, 3′ bias, lower throughput, pore and model variation. | Matched cDNA, synthetic controls, end and junction validation, orthogonal chemistry. |
| **Short-read integration** | Fragmented RNA-derived cDNA. | Deep abundance, replication, cohort statistics, junction support. | Fragmentation removes long-range phasing and leaves isoform ambiguity. | Use supported long-read models and report non-identifiable isoforms. |

Quantification is a boundary where long-read enthusiasm must be restrained. A long read that unambiguously identifies an isoform is structurally informative, but abundance estimates depend on sampling depth, molecule recovery, read length, alignment ambiguity, and transcript annotation quality. Direct RNA throughput is often lower than short-read RNA-seq, and long transcripts can be underrepresented if molecules break or fail to sequence end to end. Nanopore cDNA reads can be amplified, but amplification changes counting assumptions. For differential expression across many samples, short-read RNA-seq or targeted quantification may still be more precise, while long reads define the transcript models being counted.

Nanopore cDNA and direct RNA also differ in their relationship to RNA modifications. cDNA workflows usually erase native chemical information, although some modifications leave indirect traces through reverse transcriptase errors or stops. Direct RNA workflows preserve modification-sensitive signal but require special analysis of raw current, dwell time, or basecalling residuals. That makes direct RNA attractive for [Chapter 132](chapter1120.md) topics, but it also means direct RNA basecalling errors are not merely nuisance errors: some errors may contain biochemical information, while others imitate biochemical information.

Practical study design starts with the question. If the question is "Which isoforms exist in this tissue?", PacBio full-length or nanopore cDNA sequencing may be the most efficient discovery layer. If the question is "Does this viral RNA molecule contain these termini and rearranged junctions on the same strand?", direct RNA may be preferable. If the question is "How abundant is each annotated isoform across 300 samples?", long reads may need short-read integration. If the question is "Does this site carry a modification?", direct RNA signal alone is not enough; calibration and orthogonal chemistry are required.

Evidence status: Chen et al. (2025) and Pardo-Palacios et al. (2024) provide verified local benchmarking support for nanopore transcript-level analysis and cross-method assessment. The chapter still needs the original nanopore direct RNA method paper and a dedicated poly(A)-tail validation benchmark before historical priority or tail-accuracy claims are finalized.

## 128.3. Direct RNA modification-sensitive signals

RNA chemical modifications change bases, sugars, or termini without necessarily changing the genome-encoded sequence. Examples include N6-methyladenosine, 5-methylcytosine, pseudouridine, inosine, 2′-O-methylation, noncanonical caps, and many tRNA and rRNA modifications. Direct RNA sequencing is sensitive to some of these chemical states because the nanopore current is influenced by several adjacent nucleotides, local strand movement, and interactions between the RNA and pore. A modified base can shift current level, dwell time, or basecalling probabilities in a sequence-context-dependent way.

The measurement chain must be explicit. The instrument does not output "this nucleotide is N6-methyladenosine." The instrument records current over time. Basecalling software converts current into sequence using a statistical or machine-learning model trained on particular chemistries and reference assumptions. Modification callers then compare observed signal with an expected signal, compare native RNA with controls, or learn patterns from known modified and unmodified samples. A candidate site may be supported by a raw-current residual, a dwell-time change, a basecalling error, a mismatch pattern, or a model score. None of those observations alone proves chemical identity.

![Figure 128.3. Native RNA signal interpretation ladder](../assets/figures/chapter1117_figure3.png)

**Figure 128.3. Native RNA signal interpretation ladder.** Native RNA modifications can perturb nanopore current and dwell time, but sequence context, damage, structure, and model mismatch can produce similar residuals. Calibration determines the highest defensible claim.

There are several confounders. Local sequence context changes the expected current even in unmodified RNA. Neighboring modifications can distort a signal attributed to one site. RNA secondary structure or RNP remnants can alter motor kinetics. RNA damage from extraction, oxidation, heat, or storage can create unusual signals. Alignment errors near repeats, paralogs, splice junctions, edited sites, or homopolymers can create apparent modification signals. Basecallers trained on one pore chemistry, organism, or RNA population may not generalize to another. A sample with many unannotated isoforms can produce signal residuals because the wrong transcript model is being used.

> **Box 128.2. A signal shift is not a modification call**
>
> A current or dwell-time shift begins as a measurement difference. Wrong transcript models, neighboring chemistry, structure, damage, and basecaller mismatch can imitate a site. The claim should stop at the highest tier supported by calibrated controls.

Strong evidence designs are therefore comparative. Synthetic RNA standards with known sequence and known modification stoichiometry can define signal effects under controlled conditions. In vitro transcripts with and without a modification can isolate chemical effects from biological covariates. Genetic deletion, depletion, inhibition, or rescue of a writer enzyme can test whether candidate signals change as expected. Direct RNA data can be compared with matched cDNA data to separate native-signal effects from mapping or transcript-model errors. Orthogonal methods such as antibody-independent chemical mapping, site-specific biochemical assays, or liquid chromatography coupled to mass spectrometry can support identity or abundance, although many orthogonal methods have their own localization or antibody-bias limits.

Stoichiometry is particularly difficult. A transcript position may be modified on only a fraction of molecules. The observed current distribution then mixes modified molecules, unmodified molecules, molecules carrying neighboring modifications, damaged molecules, and technical noise. A caller may classify a site as modified, but estimating the percentage of molecules modified at that position requires calibration against mixtures and careful modeling of read-level uncertainty. Low-coverage isoforms and allele-specific transcripts make the problem harder because modification status may differ by isoform, cell type, allele, developmental stage, or stress state.

Direct RNA modification analysis also intersects with isoform biology. A signal assigned to an exon may actually belong to only one isoform that contains that exon. A modification-sensitive signal near a splice junction may be mislocalized if reads align ambiguously. A long direct RNA read can in principle connect modification-sensitive signal with isoform structure, poly(A) tail length, and allele, but only if enough high-quality reads span all features. This is a powerful reason to pursue direct RNA sequencing, not a reason to lower validation standards.

The conservative interpretation language is graded. Established: native RNA chemistry can alter direct RNA nanopore signal. Likely in a controlled system: a signal can be attributed to a particular modification when synthetic standards or writer perturbations support the assignment. Context-dependent: transcriptome-wide maps of modification-sensitive sites depend on caller, chemistry, coverage, and validation set. Disputed or incomplete: chemical identity and stoichiometry at many individual transcript sites. [Chapter 132](chapter1120.md) treats RNA modification detection more comprehensively, including chromatography and mass spectrometry.

**Table 128.3. Direct RNA modification-signal evidence ladder.** Prevent a signal difference from being reported prematurely as a chemically identified site.

| Evidence level | Observation | Supports | Does not establish |
| --- | --- | --- | --- |
| **Signal sensitivity** | Current, dwell time, basecall, or residual differs from expectation. | A native or technical difference warrants investigation. | Site, chemical identity, or stoichiometry. |
| **Candidate site** | Difference is reproducible and localized after isoform-aware alignment. | Prioritized modification-sensitive region. | Named chemistry or function. |
| **Chemically supported** | Synthetic standards or writer/eraser perturbation changes signal as predicted. | Stronger identity assignment in the tested context. | Universal portability across chemistries or contexts. |
| **Stoichiometric estimate** | Read distributions are calibrated against defined mixtures. | Approximate modified fraction with uncertainty. | Precision under low coverage or overlapping isoforms. |
| **Biological claim** | Chemical evidence is linked to condition, allele, isoform, phenotype, or perturbation. | A specific mechanistic hypothesis. | Causality without functional perturbation and rescue. |

A common misconception is that direct RNA sequencing is a direct modification detector in the same way that a mass spectrometer detects mass-to-charge features. It is better described as a native-molecule sequencing assay whose signal can be sensitive to modifications. The difference matters because overconfident modification calls can create false regulatory stories, especially in fields such as epitranscriptomics where antibody artifacts, stoichiometry errors, and cell-state confounders have already produced overinterpretation.

Evidence status: Zhao et al. (2022) supplies a verified systematic review of direct RNA modification detection, and Stephenson et al. (2022) supplies verified primary evidence that native nanopore signal is sensitive to RNA modification and structure. A broad cross-caller benchmark and additional synthetic-standard or perturbation panels remain necessary for transcriptome-wide site and stoichiometry claims.

## 128.4. Isoform-resolved and allele-resolved analyses

Isoform-resolved analysis asks which transcript features occur together on the same RNA molecule. This is different from asking whether individual features exist. A gene may have an alternative first exon, a skipped cassette exon, and a distal polyadenylation site. Short reads can detect each feature, but without phasing they may support several mutually incompatible transcript models. Long reads, targeted long-read panels, or carefully modeled short-read data can identify actual combinations. The result is an isoform catalog: a set of candidate transcript structures with read support, annotation status, and uncertainty.

An isoform catalog is not a functional catalog. Some isoforms are abundant, conserved, translated, localized, regulated, or disease-relevant. Others may be low-level splicing noise, incomplete processing intermediates, degradation products, or artifacts. A biologically meaningful isoform claim should state the evidence tier. Structural existence is supported by molecule-level reads. Quantitative relevance requires abundance and reproducibility. Coding consequence requires open reading frame and translation evidence. Regulatory consequence requires localization, binding, perturbation, or functional assays. Disease consequence requires comparison with controls, variant mechanism, and ideally rescue or independent cohorts.

Transcript ends create another phasing boundary. Alternative first and last exons often distinguish isoforms more strongly than internal splice junctions, yet 5′ truncation and internal priming make transcript ends especially vulnerable to artifacts. Cap-aware 5′ assays, independent cleavage-and-polyadenylation maps, or targeted amplification across the candidate molecule can convert a plausible end into a supported end. Without such evidence, analysts should preserve end uncertainty instead of collapsing a family of near-identical molecules into one falsely precise model.

Quantification becomes difficult when many isoforms share exons. A uniquely assigned long read is informative but often sparse; a truncated read may be compatible with several isoforms; and even a read called full-length by an analysis program may match more than one annotated transcript if the models differ only at their ends. In the datasets examined by Chen et al. (2023), 13.8–49.5% of reads were non-unique, illustrating that long reads reduce but do not abolish assignment ambiguity. The range is study-specific, not a universal platform rate. Quantification algorithms therefore distribute ambiguous reads according to explicit assumptions about transcript compatibility and abundance. Results should distinguish observed unique support, observed or estimated full-length support, and total abundance that includes model-assigned reads.

Bambu provides one concrete model of context-aware discovery and quantification. The method first corrects candidate splice-junction alignments, then groups reads that share a corrected splice chain into read classes. It ranks candidate transcript models across samples, extends the reference with candidates that pass discovery filters, assigns read classes to compatible transcripts, and uses expectation maximization to estimate transcript abundance in each sample against the shared extended catalog. Its outputs retain total abundance together with unique-read and full-length-read support. In this implementation, however, "full-length" means that a read class has the complete compatible splice-junction chain; it does not establish the biological 5′ and 3′ ends. The method uses splice junctions heavily, cannot reliably distinguish models that differ only in start or end coordinates, excludes subset transcript candidates by default to limit degradation-derived false positives, and gives less certain coordinates for single-exon transcripts.

The novel discovery rate (NDR) in Bambu also needs a scoped interpretation. For a well-annotated genome, the proportion of unannotated high-scoring models can provide a conservative upper bound on false discovery because most valid transcripts are presumed already annotated. In a poorly annotated genome, unannotated models include many valid missing transcripts, so the NDR instead mixes annotation incompleteness with false positives and is used as a conservative bound on the valid-discovery component. Without a suitable reference annotation, Bambu can rank candidates with a pre-trained transcript probability score, but that score is not calibrated like the NDR across analyses. Additional minimum-read and within-gene fraction filters still operate alongside the NDR. Thus, a single reported NDR value is interpretable only with the annotation, training model, sample set, and auxiliary filters.

Cross-condition comparisons require a common transcript definition. If each sample is collapsed independently, minor differences in read completeness can produce condition-specific models that look biological. A stronger workflow creates a harmonized candidate catalog, retains sample-level support, applies common end and junction rules, and then quantifies against that catalog. This does not eliminate condition-specific biology; it prevents the annotation procedure from becoming the condition effect. Designs whose primary question is RNA production, processing rate, or lifetime should hand long-read-defined isoforms to the kinetic framework in [Chapter 129](chapter1154.md) rather than treating this chapter as the owner of those inferences.

Context-aware catalogs can improve abundance estimates when a static reference contains inactive isoforms or omits isoforms expressed in the sample. Missing models can cause reads to be over-assigned to an annotated relative, whereas adding too many weak models enlarges equivalence classes and can destabilize quantification. In the Bambu spike-in benchmarks, discovery of deliberately omitted transcripts reduced abundance error, but increasingly permissive discovery also introduced false-positive models. In its human embryonic stem-cell example, the method assigned expression to 242 HERVH-derived gene models comprising 464 transcripts, with 64 loci accounting for 90% of the estimated HERVH RNA. That result is a context-specific demonstration in the tested cell line, not a general count of active HERVH loci in all pluripotent cells.

Allele-resolved transcriptomics adds haplotype information. The simplest case is allele-specific expression: one allele produces more RNA than the other. Long-read data can go further by linking allele-specific variants to splice choices, retained introns, editing sites, polyadenylation sites, or direct RNA signals. In imprinting, cis-regulatory variation, cancer, and viral quasispecies, this molecule-level phasing can be essential. However, allele-resolved analysis requires heterozygous or variant positions, adequate read coverage, accurate genotype or haplotype information, and bias-aware alignment. A reference genome can preferentially attract reads matching the reference allele, while repetitive genes, segmental duplications, and pseudogenes can produce false allele-specific assignments.

An allele-swap remapping test, exemplified by WASP, exposes reference-mapping bias at the read level. After initial alignment, each read overlapping known single-nucleotide variants is rewritten with the other possible alleles and remapped; the original read is retained only when all tested versions map uniquely to the same location. A personalized reference alone does not guarantee this symmetry because the reference and alternative haplotypes can have different uniquely mappable sequence. The remapping test trades bias for information: rejected reads reduce depth and can lower apparent locus abundance. Genotype errors, unknown variants, insertion-deletion variants, duplicate selection, phasing errors, paralogs, and overdispersed allele counts need separate controls. The 2015 WASP study established this principle with short-read RNA sequencing and chromatin data, so its exact software behavior should not be assumed to solve modern long-read error profiles without benchmarking.

Read-level phasing and population-level allelic imbalance answer different questions. One long read containing a heterozygous variant and an alternative exon establishes that those features co-occurred on one recovered molecule. It does not estimate the population frequency of that haplotype with useful precision. Conversely, thousands of short reads at a heterozygous site may estimate allelic imbalance but cannot necessarily assign that imbalance to a complete isoform. In the Geuvadis study, paired genome and short-read RNA sequencing of 462 lymphoblastoid cell lines supported population-scale analysis of both allele-specific expression and allele-specific transcript structure. Significant events occurred at a median of 6.5% and 5.6% of tested sites per individual, respectively, and their overlap showed that allelic abundance differences can be coupled to transcript-structure differences. Because the RNA reads were 75-nucleotide fragments from one cultured cell type, this evidence demonstrates population inference and its controls, not full-length allelic isoform phasing. A defensible long-read study therefore reports informative-molecule counts, biological replicates, remapping and duplicate policies, genotype uncertainty, the number and spacing of variants, and whether phasing was read-backed, statistically inferred, or imported from genomic data.

The viral example illustrates allele and variant resolution. In an RNA virus population, different genomes or defective RNAs may coexist. A long native read can connect a copyback junction with terminal sequence and internal viral variants, supporting a molecule-level structure. Short reads can quantify the junction across more samples but may not phase the entire viral RNA. Direct RNA can avoid some reverse-transcription artifacts, but RNA damage, template switching in any cDNA comparison, and alignment ambiguity near terminal repeats still require controls. The Pye et al. Sendai virus reference in the local bibliography is used here only for the narrow point that direct RNA sequencing can validate otherwise difficult viral copyback RNA structures.

Evidence status: Chen et al. (2023) directly supports the context-aware discovery and quantification example, including its annotation-regime and transcript-end limitations. Lappalainen et al. (2013) supplies population-scale evidence that allele-specific expression and transcript structure can overlap, while van de Geijn et al. (2015) supplies the allele-swap remapping principle and statistical artifact controls. Both allele papers used short-read data; a dedicated long-read read-backed phasing benchmark remains a specific gap.

## 128.5. Error models, validation, quantification, and short-read integration

An error model is an explicit account of how an assay can produce wrong or misleading observations. For long-read transcriptomics, major errors include incomplete cDNA synthesis, internal priming, template switching, chimeric molecules, sequencing errors, splice-junction misalignment, read truncation, length-dependent recovery, sample degradation, and over-collapsing or under-collapsing transcript models. For nanopore direct RNA sequencing, additional errors include adapter-ligation bias, 3′ loading bias, pore motor variability, current-level ambiguity, modification-sensitive basecalling, lower read depth, and RNA damage. For allele- and isoform-resolved analysis, reference bias, paralogous mapping, sparse informative variants, ambiguous read assignment, and annotation-dependent transcript definitions can dominate the biological signal.

Validation should match the biological claim. A novel splice junction can be checked by junction-spanning reverse-transcription polymerase chain reaction, independent long-read support, short-read junction coverage, and canonical splice-site evaluation. A novel full isoform requires evidence that its ends and internal junctions belong together; cap-aware 5′ end mapping and 3′ end sequencing help separate true ends from truncated cDNA and internal priming. A transcript fusion requires controls for readthrough transcription, mapping ambiguity, template switching, and genomic rearrangement. A direct RNA modification-sensitive site requires synthetic standards, writer or eraser perturbation, matched cDNA comparison, or orthogonal chemical evidence. An allele-specific isoform requires bias-aware alignment, sufficient informative molecules, independent genomic phasing when possible, and replication of the imbalance.

![Figure 128.5. Claim-specific validation routing](../assets/figures/chapter1117_figure5.png)

**Figure 128.5. Claim-specific validation routing.** A splice junction, complete transcript, fusion, allele-specific isoform, viral RNA structure, and modification-sensitive signal require different validation evidence. Allele-resolved claims additionally require mapping-symmetry, genotype, phasing, and molecule-count checks.

Short-read integration remains essential because long reads and short reads have complementary strengths. Short reads provide depth, mature-abundance estimates across large cohorts, compatibility with unique molecular identifiers in some protocols, and strong statistical power for differential expression. Long reads provide structural linkage, while direct RNA adds native-signal evidence. A good integrated analysis does not force all data types to agree; it asks whether disagreement comes from molecule selection, transcript definitions, mapping ambiguity, depth, or real biology.

A common integrated workflow begins by using long reads to discover transcript models. The analyst then filters obvious artifacts: low-complexity internal priming, noncanonical splice junctions without support, single-read models in noisy regions, implausible chimeras, and models inconsistent with independent end evidence. Next, short reads are assigned to the refined transcriptome to improve quantification, recognizing that short reads may still not distinguish all isoforms. Finally, high-impact claims receive targeted validation. In clinical or disease studies, this workflow prevents two opposite errors: dismissing real pathogenic isoforms because they were absent from the reference annotation, and overinterpreting rare technical molecules as disease mechanisms.

For multi-sample quantification, the refined transcriptome should be built once for the comparison, not rediscovered independently in each sample. A transparent report separates the discovery threshold from the quantification model, states whether the threshold was calibrated against a complete or incomplete annotation, and partitions abundance evidence into unique reads, ambiguous reads assigned by a model, and reads judged full-length under an explicit operational definition. Expectation maximization can produce useful estimates from equivalence classes, but it does not turn ambiguous molecules into observations of a particular isoform. Sensitivity analyses that remove weak candidate models, change end tolerances, or compare fixed-reference with context-aware catalogs reveal how much the conclusion depends on annotation complexity.

Error models should be reported by context, not hidden in a generic limitations paragraph. A plant transcriptome with many paralogs has different mapping risks from a compact viral genome. A degraded autopsy sample has different 3′ bias from fresh cultured cells. A direct RNA modification study in ribosomal-RNA-depleted total RNA has different background from a poly(A)-selected messenger-RNA study. A targeted rare-disease panel has different amplification and ascertainment biases from an untargeted population atlas. Good reporting states input material, selection strategy, read-length distribution, depth, strandedness, replicate structure, spike-ins or controls, pore and basecaller versions where relevant, alignment settings, transcript-collapsing criteria, and validation thresholds.

![Figure 128.6. Integrated long-read, short-read, and direct RNA evidence workflow](../assets/figures/chapter1117_figure6.png)

**Figure 128.6. Integrated long-read, short-read, and direct RNA evidence workflow.** Integrated transcriptomics separates structural discovery, artifact filtering, common-catalog construction, abundance estimation, native RNA interpretation, and validation. Discovery thresholds and assignment models create different uncertainties that must remain visible.

The most common overgeneralizations are predictable. Do not equate a long read with a complete native RNA molecule. Do not treat a transcript catalog as a list of functional isoforms. Do not describe direct RNA sequencing as unbiased simply because reverse transcription is omitted. Do not call a nanopore signal shift a chemical modification without calibration. Do not treat an allele-bearing read as a precise population-level estimate, and do not interpret an independently collapsed sample catalog as direct evidence of condition-specific isoforms.

The current consensus is conservative integration. Long-read technologies are indispensable for transcript structures, direct RNA is uniquely valuable for native-molecule questions, and short-read sequencing remains powerful for depth and cohort-scale statistics. The unsettled questions are how to define biologically meaningful isoforms, how to benchmark direct RNA modification callers across chemistries and organisms, how to quantify closely related isoforms without concealing ambiguity, and how to translate long-read discovery into clinical reporting standards.

## Experimental Foundations and Evidence Status

The chapter-local bibliography now includes verified reviews and benchmarks for long-read workflow design, transcript identification and quantification, nanopore transcript-level analysis, direct RNA modification-sensitive signal, targeted long-read sequencing, allele-specific population analysis, and reference-mapping-bias control. It also contains one verified published study and its preceding preprint on direct RNA validation of Sendai virus copyback genomes; the preprint is publication-history metadata rather than independent replication.

The newly available full texts sharpen two boundaries. WASP supplies a reference-bias control principle but is not a long-read phasing benchmark, and Bambu supplies an explicit context-aware discovery and quantification model but does not validate biological transcript ends. Remaining gaps are therefore specific: the original direct RNA and PacBio workflow papers, dedicated transcript-end and poly(A)-tail validation studies, a dedicated long-read read-backed phasing benchmark, a broad modification-caller benchmark, and a clinical reporting framework. No kinetic or nascent-transcriptomics source gap remains in this chapter because that evidence domain now belongs to [Chapter 129](chapter1154.md).

## Recent Consensus

The recent consensus is not that one sequencing platform has replaced the others. The consensus is that transcriptome questions must be matched to molecule type, time scale, and evidence level. PacBio full-length transcriptomics and nanopore cDNA sequencing are now accepted as central tools for transcript structure discovery because they can observe feature combinations directly. Short reads remain central for large cohorts, deep quantification, differential expression, and many statistical comparisons. Direct RNA sequencing occupies a distinct niche because native RNA signal can carry information that cDNA loses, but that advantage comes with throughput, chemistry, and interpretation costs.

For full-length cDNA transcriptomics, the most stable conclusion is that long-read transcript models improve annotation but require confidence tiers. A transcript supported by multiple full-length reads, independent junction evidence, plausible splice signals, and end-specific support is stronger than a transcript represented by one read in a noisy locus. Annotation systems increasingly need to distinguish candidate, supported, validated, and functional isoforms. That tiered view avoids two errors: treating every observed RNA molecule as a functional isoform, and treating reference annotations as complete descriptions of biology.

For nanopore direct RNA sequencing, the consensus is similarly balanced. Direct RNA sequencing is the clearest route to sequencing native RNA molecules at long range, and it is valuable for viral RNA structures, poly(A) tails, RNA damage, and modification-sensitive signal. At the same time, direct RNA sequencing is not a universal replacement for cDNA sequencing. Many studies need cDNA amplification, higher yield, broader sample compatibility, or short-read depth. Direct RNA should be chosen when native molecule state is part of the question, not merely because "direct" sounds more faithful.

For computational analysis, the consensus is that uncertainty should remain visible. Transcript-collapsing thresholds, splice-junction filters, annotation completeness, read-assignment models, basecalling versions, modification callers, allele-remapping filters, haplotype filters, and end-clustering tolerances can all change conclusions. A calibrated discovery threshold controls only the error quantity defined by that method and annotation regime; it does not certify every model or make ambiguous assignments direct counts. Reproducible reporting should include software and chemistry versions, read-length distributions, alignment parameters, transcript-model filters, replicate structure, unique and ambiguous support, allele-remapping and duplicate policies, and validation thresholds. This is not bookkeeping; it is the difference between a reusable transcriptome result and a catalog that cannot be interpreted outside the original pipeline.

## Open Questions, Controversies, Deprecated Models, and Common Misconceptions

Open questions:

- How can researchers define a biologically meaningful isoform? A molecule can be real but rare, unstable, nonfunctional, or produced by splicing noise. Conversely, a rare isoform can be critical in a specific cell type, developmental stage, infection state, or genetic disease. The useful distinction is not simply whether the isoform is real, but what claim is being made about the isoform. Structural existence, regulated production, translation, localization, disease mechanism, and therapeutic relevance are different claims requiring different evidence.

Controversies:

- Direct RNA modification calling remains a fast-moving and unsettled area. The field agrees that native RNA modifications can affect nanopore signal. The controversy lies in chemical identity, site localization, stoichiometry, cross-platform comparability, and benchmarking. A caller trained on one pore chemistry or modification standard may not generalize to all organisms, transcript classes, or neighboring sequence contexts. Claims that a transcriptome-wide direct RNA map definitively identifies thousands of modification sites should be read with attention to controls, calibration, and orthogonal evidence.
- Long-read quantification remains methodologically unsettled. Read truncation and isoform overlap create ambiguous assignments, while transcript discovery and quantification performed on the same data can overfit noise. Benchmark conclusions depend on truth-set completeness, expression range, RNA integrity, platform chemistry, and whether evaluation is performed at gene, junction, transcript, or haplotype level.

Common misconceptions:

- "Long reads automatically solve quantification." Long reads improve connectivity, but quantification still depends on depth, bias, molecule recovery, error correction, and model assumptions.
- "Direct RNA sequencing is automatically unbiased." Direct RNA sequencing avoids reverse transcription but still has pore, motor, modification, length, input, and basecalling biases.
- "A full-length cDNA read is automatically a complete native RNA." cDNA reads can reflect template switching, truncation, priming, degradation, or library artifacts.
- "A transcript model is automatically functional." A transcript model is an annotation hypothesis; function requires expression, processing, localization, perturbation, and consequence evidence.
- "One allele-bearing read proves allele-specific expression." A single phased molecule establishes co-occurrence, not a reproducible population imbalance.
- "A personalized reference eliminates allele-mapping bias." The two haplotypes can still differ in unique mappability; allele-swap remapping or an equivalent symmetry test is needed, with the resulting loss of reads reported.
- "Condition-specific transcript catalogs prove condition-specific splicing." Independent collapsing can turn technical differences in completeness and depth into apparently condition-specific isoforms.
- "A full-length support count proves complete transcript ends." Some quantifiers define full-length by splice-chain compatibility, which does not validate the native transcription start or cleavage site.

Deprecated or weakened claims:

- Several older simplifications should be treated as deprecated. It is no longer adequate to describe a gene as having one representative transcript when long-read and end-mapping data show regulated isoform structure. It is not adequate to regard every multi-exon read as full-length, every novel transcript model as functional, or every direct RNA mismatch as a modification. These simplifications can still be useful in introductory diagrams, but they should not drive mechanistic conclusions.
