# Chapter 127. RNA End, Tail, and Cleavage Profiling: TSSs, 3' Ends, Poly(A), and Degradomes

## Scope Note

This chapter explains transcriptome-scale methods that turn an RNA terminus into a mapped coordinate, a tail measurement, or a cleavage profile. Its primary ownership is library chemistry and the first inferential step: which terminal chemical groups are eligible, which molecules are selected, what coordinate a read reports, and how peaks, tails, or cleavage sites are called with uncertainty. Cap-selected transcription-start-site (TSS) profiling, 3' cleavage and alternative-polyadenylation assays, poly(A)-tail sequencing, non-A tail detection, degradome and parallel analysis of RNA ends (PARE), and 5'-phosphate profiles are treated as related measurement systems rather than as interchangeable forms of RNA sequencing. [Chapter 123](chapter1155.md) owns targeted primer-extension and rapid amplification of cDNA ends (RACE) confirmation; [Chapter 126](chapter1157.md) owns specialized small- and stable-RNA profiling; [Chapter 128](chapter1117.md) and [Chapter 129](chapter1154.md) own long-read transcript reconstruction and kinetic labeling; [Chapter 141](chapter1128.md) owns reusable annotation and isoform informatics. Chapters on transcription, processing, translation, and decay own the underlying biology.

## Executive Summary

An RNA end is both a coordinate and a chemical object. A 5' end may carry an N7-methylguanosine cap, a triphosphate, a monophosphate, or a hydroxyl. A 3' end may carry a hydroxyl, a phosphate, or a 2',3'-cyclic phosphate, and it may be followed by an encoded or untemplated tail. These states determine whether a ligase, phosphatase, pyrophosphatase, cap-binding reagent, reverse transcriptase, or nanopore motor can admit the molecule to a library. Therefore, an end-sequencing result describes an eligible molecular population after a defined sequence of reactions. It is not a census of all ends in the sample.

Cap-selected methods enrich 5'-complete capped RNAs and map candidate transcription starts. CAGE obtains a short tag from the cap-proximal end; RAMPAGE combines cap enrichment with paired-end information that can improve transcript or promoter assignment. The exact measured entity remains a captured capped 5' end, not initiation itself. Recapping, post-transcriptional cleavage, RNA degradation, incomplete cap selection, reverse-transcription truncation, and mapping ambiguity can all affect interpretation. Promoter clusters are analysis objects assembled from nearby end coordinates, and broad initiation regions should not be forced into a single-base model merely because the sequencer reports single-base positions.

3' end libraries map cleavage and polyadenylation sites by capturing terminal sequence or by priming through a poly(A) tract. Oligo(dT) priming is efficient but can initiate within genomically encoded A-rich sequence, producing internal-priming artifacts. Direct ligation to the terminal hydroxyl avoids that particular mechanism but introduces ligation and end-chemistry selection. A read ending near a polyadenylation signal supports a cleavage-site candidate; assignment to an isoform additionally requires strand, transcript structure, and compatible upstream evidence. Alternative polyadenylation (APA) is a biological interpretation of changes among validated sites, not a synonym for every cluster of 3' reads.

Poly(A)-tail methods face a different measurement problem because homopolymers are difficult for standard sequencing and amplification. PAL-seq relates a tail-dependent biochemical signal to calibrated standards; TAIL-seq reads terminal sequence with specialized signal analysis and can detect non-A residues; mTAIL-seq increases sensitivity for limited input. Direct RNA nanopore measurements can link a tail estimate to a long individual RNA, but tail calling depends on signal segmentation, motor behavior, basecalling, and platform-specific calibration. Method comparisons must distinguish molecule-level tail distributions from gene-level averages and must specify whether non-A residues, very short tails, or non-polyadenylated molecules were eligible.

Degradome and PARE workflows usually select uncapped RNAs with ligatable 5' monophosphates. A sharp 5' end can identify the downstream product of endonucleolytic cleavage, including plant microRNA-directed cuts, but decapping, exonuclease pauses, nonspecific breakage, and secondary decay can create the same chemistry. 5PSeq uses phased 5'-phosphorylated decay intermediates to infer ribosome-associated protection patterns. These ends are footprints of an RNA-decay process coupled to translation; they are not conventional ribosome-protected fragments and do not directly measure occupancy in the same way as ribosome profiling.

Analysis requires explicit coordinate conventions, peak definitions, denominators, and uncertainty. A 5' read coordinate may denote the first aligned nucleotide of a downstream fragment; a cleavage event lies on the phosphodiester bond adjacent to that nucleotide. A 3' coordinate may denote the last templated base, the cleavage boundary, or the first untemplated nucleotide. One-nucleotide differences can therefore be notational rather than biological. Spike-ins with known end chemistry and tail length, synthetic sequence diversity, negative controls, independent library chemistries, and targeted confirmation are needed to separate chemistry, recovery, sequencing, and inference errors.

## Concept Inventory

- **Molecular eligibility:** the set of RNA molecules capable of completing the stated end-capture reactions.
- **Terminal chemical state:** the chemical group at an RNA 5' or 3' terminus, including cap, phosphate multiplicity, hydroxyl, phosphate, or cyclic phosphate.
- **End conversion:** an enzymatic treatment that changes terminal chemistry to admit or exclude a molecular class.
- **End coordinate:** the reported reference position associated with the terminal aligned nucleotide under a stated convention.
- **Cleavage bond:** the phosphodiester bond broken to create an upstream and downstream product; it is conceptually between coordinates.
- **Cap selection:** enrichment for capped RNA through cap binding, oxidation/biotinylation, enzymatic discrimination, or related chemistry.
- **TSS cluster:** a set of nearby cap-associated 5' ends summarized as a promoter-associated initiation region.
- **Internal priming:** reverse transcription initiated within an internal A-rich RNA sequence rather than at a terminal poly(A) tail.
- **Cleavage-and-polyadenylation site:** a transcript boundary supported as a pre-mRNA cleavage site and usually followed by tail addition.
- **Tail-length distribution:** the molecule-level distribution of terminal tail lengths, not merely its mean or median.
- **Non-A tail:** one or more untemplated terminal nucleotides other than adenosine within or after an oligo(A) tract.
- **Degradome:** the sampled population of RNA decay or cleavage products selected by a stated terminal chemistry.
- **PARE:** parallel analysis of RNA ends, commonly selecting 5'-monophosphorylated uncapped products to map cleavage.
- **Protected-end periodicity:** repeated end positions phased relative to codons or a bound complex, interpreted through an explicit protection model.
- **Process control:** a synthetic or biological RNA that experiences selected workflow stages and tests their combined performance.
- **Isoform assignment ambiguity:** uncertainty about which transcript model generated an end-compatible read or peak.

## What to Know Before Reading This Chapter

RNA is synthesized and processed with directional chemistry. Polymerases extend the 3' hydroxyl, leaving chemically distinct termini depending on the polymerase and subsequent enzymes. Eukaryotic RNA polymerase II transcripts acquire a 5' cap early, and many mature mRNAs receive a poly(A) tail after endonucleolytic cleavage. Bacterial primary transcripts commonly begin with a triphosphate. Endonucleases can leave 5' phosphate/3' hydroxyl, 5' hydroxyl/2',3'-cyclic phosphate, or other combinations. Exonucleases alter ends again. The same genomic coordinate can consequently occur in primary, processed, and degraded molecules with different chemistries.

Three running examples expose the inferential steps. First, a mammalian gene uses a broad promoter and two 3' cleavage sites. A cap-selected library must summarize distributed TSSs without inventing a single universal start, while a 3' library must distinguish the distal cleavage site from internal priming in an A-rich terminal exon. Second, a plant mRNA is sliced by a microRNA. PARE can enrich the 5'-monophosphorylated downstream product, but a cleavage assignment also depends on the expected guide-target register, replication, background decay, and independent validation. Third, a maternal mRNA changes poly(A)-tail length during embryonic activation. A population median can shift because every molecule changes or because the mixture of short- and long-tailed states changes; the biological models are different.

A sequencing read is not the original RNA. Library preparation may ligate an adapter, copy RNA into cDNA, amplify it, trim a homopolymer computationally, and align the remaining bases to a reference. At each transformation, some molecules are lost and others are duplicated. Unique molecular identifiers can reduce amplification distortion but cannot recover molecules that never ligated or copied. Random fragmentation controls sequence and mapping behavior but does not reproduce terminal chemistry. Standards must therefore be placed before the stage they are intended to evaluate.

This chapter teaches primary end inference. It provides enough promoter, polyadenylation, translation, and decay biology to interpret the measurements, but deeper mechanisms are handed to [Chapter 26](chapter1025.md), [Chapter 29](chapter1028.md), [Chapter 120](chapter1114.md), and [Chapter 35](chapter1033.md). Targeted primer extension and RACE in [Chapter 123](chapter1155.md) are valuable confirmation tools precisely because they use different selection and amplification chains. Reusable transcript assembly, feature annotation, and comparative isoform analysis remain with [Chapter 141](chapter1128.md).

## 127.1. RNA-end chemistry, molecular eligibility, library design, and process controls

The first design question is not “which sequencer?” but “which end chemistry is the analyte?” A ligation reaction is a chemical gate. Common RNA ligases join a donor and acceptor only when the required phosphate and hydroxyl are presented in the appropriate geometry. A 5'-monophosphorylated degradation product can be directly ligatable in a workflow that excludes capped and triphosphorylated primary transcripts. Conversely, cap enrichment selects a molecular class that a direct 5'-phosphate ligation omits. Reporting “5' ends” without the gate makes the result irreproducible and biologically ambiguous.

Five-prime states need explicit separation. A cap is a modified guanosine attached through an unusual 5'-to-5' triphosphate bridge and is not equivalent to a simple triphosphate. Primary bacterial or organellar RNAs may carry 5' triphosphates; processing products may carry 5' monophosphates or hydroxyls. RNA pyrophosphohydrolases can convert triphosphate to monophosphate, while phosphatases remove accessible phosphates and polynucleotide kinases add phosphate to hydroxyl ends. Tobacco acid pyrophosphatase and newer decapping reagents have been used to convert capped ends, but treatment efficiency and off-target activity must be measured rather than assumed. A subtractive design—untreated, phosphatase-treated, and conversion-treated aliquots—can identify chemistry classes only if each conversion is efficient and controls reveal cross-reaction.

Three-prime states are equally consequential. Many intact biological RNAs and hydrolytic cleavage products have a 3' hydroxyl suitable for adapter ligation. Some endonucleases generate 3' phosphate or 2',3'-cyclic phosphate, which blocks standard ligation until opened or exchanged. Periodate oxidation can discriminate terminal cis-diols, while enzymatic repair can collapse several states into a common ligatable form. Such conversion increases coverage but destroys the information that originally distinguished the states. Parallel untreated and repaired libraries preserve more interpretive power than a single universal-repair library.

The complete eligibility chain includes more than the terminal group. RNA length, structure, bound protein, nearby modification, terminal sequence, and compartment-specific extraction can affect ligation, reverse transcription, purification, and mapping. An adapter ligase can prefer some terminal bases and structures; reverse transcriptase can stop at modifications or stable folds; size selection can remove long precursors and short cleavage products. Poly(A) selection excludes non-polyadenylated ends and often underrepresents deadenylated intermediates. Ribosomal depletion retains a broader molecular space but changes background and input requirements.

Library designs should be expressed as truth tables. Rows list cap, triphosphate, monophosphate, hydroxyl, 3' hydroxyl, 3' phosphate, cyclic phosphate, and tail classes. Columns list untreated, enzyme-treated, affinity-selected, ligated, reverse-transcribed, and retained. Each cell states expected admission, exclusion, or uncertain efficiency. This makes it possible to see, for example, that a “degradome” library may contain decapped full-length mRNAs as well as endonucleolytic products because both expose 5' monophosphate.

Process controls must challenge the chemistry. A useful synthetic panel varies terminal state, adjacent sequence, structure, length, and tail length at known molar ratios. Controls added before extraction measure more stages than controls added before ligation, but synthetic naked RNA may not mimic ribonucleoprotein accessibility or endogenous decay. Negative controls include adapter-only reactions, no-enzyme aliquots, conversion-minus aliquots, cap-depleted or phosphatase-treated material, and target-free matrices when applicable. A single spike-in cannot simultaneously calibrate recovery of capped long mRNA, a short cyclic-phosphate fragment, and a heavily structured rRNA end.

![Figure 127.1. Terminal chemistry determines molecular eligibility](../assets/figures/chapter1158_figure1.png)

**Figure 127.1. Terminal chemistry determines molecular eligibility.** Make the physical end state the first branch in assay selection.

**Table 127.1. End-state eligibility and diagnostic conversions.** Provide a reusable assay truth table.

| Original state | Directly eligible example | Useful conversion | Information preserved or lost | Essential control |
| --- | --- | --- | --- | --- |
| **5' cap** | Cap capture | Decap to 5'P for ligation | Conversion loses cap identity unless paired aliquot retained | Conversion-minus and cap-depleted controls |
| **5' triphosphate** | Triphosphate-selective workflow | Pyrophosphate removal to 5'P | Conversion merges primary and other 5'P ends | Defined tri- and monophosphate spike-ins |
| **5' monophosphate** | Direct 5' adapter ligation | Phosphatase exclusion | Reports uncapped ligatable products, not their mechanism | Phosphatase-treated aliquot |
| **5' hydroxyl** | Hydroxyl-aware chemistry | Kinase to 5'P | Conversion merges original hydroxyl with existing phosphate unless separated | Hydroxyl and phosphate standards |
| **3' hydroxyl** | Common 3' adapter ligation | None | Retains compatible terminal class | Sequence-diverse ligation panel |
| **3' phosphate/cyclic phosphate** | Blocked in common ligation | Phosphatase or cyclic-phosphate repair | Original state lost after universal repair | Untreated/repaired pair |

The reporting unit is the eligible molecule after stated processing. Raw molecules, ligated molecules, unique cDNAs, aligned reads, and called sites are different denominators. Library complexity and duplication diagnose only molecules that entered the observable chain. The broader sampling and extraction framework is treated in [Chapter 122](chapter1115.md), while the specialized modification and stable-RNA barriers of short molecules are treated in [Chapter 126](chapter1157.md).

## 127.2. 5' end and transcription-start-site mapping by cap capture and enzymatic selection

A transcription start site is the genomic position at which a polymerase began an RNA chain, but a cap-associated 5' end is the measurable proxy in many eukaryotic systems. Cap analysis of gene expression (CAGE) enriches capped, 5'-complete cDNAs and sequences short tags adjacent to their 5' ends. Counts can quantify promoter-associated capped ends, while clusters of nearby tags describe focused or broad initiation patterns. The method helped establish promoter atlases because it combines nucleotide-level end information with scalable expression profiling.

Cap trapping exploits chemical or affinity features of the cap and full-length RNA-cDNA hybrids. Implementations differ, so “CAGE” does not name one immutable selection chain. RNA integrity, reverse-transcription reach, cap oxidation or capture, restriction or tagmentation, PCR, and tag mapping all affect recovery. The CAGE protocol literature therefore emphasizes full-length cDNA selection and quality controls. A TSS observed in one protocol but absent in another may reflect biology, input depth, cap selection, reverse-transcription behavior, or mapping filters.

RAMPAGE—RNA annotation and mapping of promoters for the analysis of gene expression—combines cap-associated enrichment with paired-end sequencing of 5'-complete cDNAs. One end identifies the cap-proximal coordinate; the mate provides downstream sequence that can improve assignment to a gene or transcript in repetitive or promoter-dense regions. The mate does not necessarily reconstruct the full RNA, and RAMPAGE remains primarily a promoter-activity assay. Long-read approaches in [Chapter 128](chapter1117.md) can connect more distant transcript features but have their own completeness and end-validation problems.

Enzymatic selection provides complementary views. A workflow may dephosphorylate existing 5' monophosphates, convert caps or triphosphates to monophosphates, and ligate an adapter only after conversion. The difference between treated and untreated aliquots can enrich primary ends. However, incomplete phosphatase treatment leaks processed products into the selected set, incomplete conversion loses true starts, and any endogenous recapping can produce capped ends not created by transcription initiation. In bacteria, differential RNA sequencing and related strategies distinguish triphosphorylated primary transcripts from monophosphorylated processed products through analogous conversion logic.

Peak calling begins with strand-specific 5' coordinates. Technical duplicates are collapsed when identifiers permit; low-mappability positions and internal artifacts are filtered; nearby positions may be clustered. A focused promoter can have a dominant base, while a broad promoter distributes starts over tens of nucleotides. Summarizing both as the “highest base” erases a meaningful architectural distinction. Cluster widths, dominant-start fractions, expression thresholds, replicate concordance, and the distance rule used to merge tags should be reported.

Annotation is not proof. A cap-associated cluster near an annotated promoter gains credibility from promoter chromatin, polymerase occupancy, nascent transcription, sequence motifs, and reproducibility, yet no one feature is mandatory across all promoter types. An intragenic cluster may mark an alternative promoter, a recapped cleavage product, a transposable-element promoter, or mapping error. Changes in CAGE signal can reflect initiation, cap stability, or the abundance of capped products. Perturbing initiation factors and observing nascent RNA provides stronger causal evidence than a static cap map.

![Figure 127.2. From capped molecules to TSS clusters](../assets/figures/chapter1158_figure2.png)

**Figure 127.2. From capped molecules to TSS clusters.** Distinguish cap-associated capture, read mapping, clustering, and initiation interpretation.

**Table 127.2. TSS evidence and alternative explanations.** Match promoter claims to evidence.

| Observation | Supported inference | Main alternative | Strong next evidence |
| --- | --- | --- | --- |
| **Replicated CAGE cluster** | Stable capped 5' region exists | Recapping or stable processed end | Nascent transcription and promoter perturbation |
| **RAMPAGE pair maps into a gene** | Promoter-to-gene assignment strengthened | Shared downstream sequence or incomplete cDNA | Splice/long-read linkage |
| **Narrow dominant start** | Focused captured-start distribution | Caller/depth compression | Replicate cluster-width analysis |
| **Broad cluster** | Dispersed captured starts | Mapping or degradation background | Cap-selection control and promoter chromatin |
| **Condition-specific cluster** | Capped-end usage changes | Cell composition or differential stability | Matched cell types and kinetics |

Orthogonal validation should match the claim. Targeted primer extension or 5' RACE in [Chapter 123](chapter1155.md) can confirm a boundary with a different selection chain. Nascent RNA and kinetic assays in [Chapter 129](chapter1154.md) can test whether a site is newly transcribed. Promoter perturbation, reporter analysis, and chromatin evidence can support initiation function. The mechanistic interpretation of focused and dispersed promoters belongs to [Chapter 26](chapter1025.md); this chapter owns how mapped ends become candidate TSSs and clusters.

## 127.3. 3' end mapping, cleavage sites, and alternative polyadenylation assays

Most mature eukaryotic mRNAs are cleaved and then polyadenylated. A 3' end assay seeks the junction between the last genome-templated nucleotide and the untemplated tail. This is not merely ordinary RNA-seq coverage near a transcript end: the library must preserve or specifically select the natural terminus. Protocol families use oligo(dT)-primed reverse transcription, adapter ligation to a 3' hydroxyl, splinted capture of a poly(A) junction, or combinations of fragmentation and terminal selection. Each family admits a different population and has a distinct artifact model.

Oligo(dT) priming captures polyadenylated RNA efficiently but can anneal within internal A-rich sequence. Such internal priming produces a cDNA whose apparent 3' boundary lies upstream of a genomic run of adenosines. Filters based on downstream A content reduce this artifact, but aggressive filtering can remove genuine cleavage sites that happen to occur near A-rich sequence. Stronger evidence includes an untemplated poly(A) junction, terminal-ligation support, reproducibility across chemistries, nearby polyadenylation signals, and targeted validation. None of these features alone is universal.

Direct 3' adapter ligation avoids oligo(dT) internal priming at the capture step, but ligation requires a compatible terminal chemistry and can show sequence or structure preferences. Random fragmentation before selection changes which side of the junction is retained. PCR and size selection can distort abundance. Short-read protocols usually identify the local end but not the complete upstream exon chain, so a cleavage peak shared by several isoforms cannot automatically be assigned to one transcript model. Paired-end sequence, long reads, splice-junction evidence, and condition-specific annotation improve assignment.

Site calling groups nearby terminal reads because cleavage is often heterogeneous over a small window. A cluster can be summarized by its modal base, weighted mean, median, or interval. The summary choice matters when comparing conditions: two clusters can have equal centers but different dispersion, or an apparent shift can result from changing mixtures of adjacent sites. The site definition should specify strand, coordinate convention, clustering radius, minimum support, replicate rule, and treatment of multimapping and PCR duplicates.

Alternative polyadenylation is the regulated use of alternative cleavage and polyadenylation sites. It can alter the terminal exon, coding sequence, or 3' untranslated region (UTR). A proximal-versus-distal usage ratio is interpretable only after site identity and compatible transcript context are established. Changes in total gene expression, RNA stability, cell composition, library depth, or priming efficiency can affect raw site counts. Within-gene proportions reduce some global effects but create compositional dependence: increasing one site forces the proportions of others downward.

Biological interpretation requires handoff to processing mechanisms. A shift toward a proximal site can remove regulatory sequence from the mature RNA, but the phenotypic consequence depends on which binding sites, localization elements, or translation features are actually changed. A cleavage peak does not identify the responsible cleavage factor. [Chapter 29](chapter1028.md) treats cleavage and polyadenylation machinery, regulation, and APA biology; [Chapter 72](chapter1067.md) treats UTR consequences. This chapter retains ownership of capture, site definition, usage estimation, and artifacts.

![Figure 127.3. Genuine 3' cleavage versus internal priming](../assets/figures/chapter1158_figure3.png)

**Figure 127.3. Genuine 3' cleavage versus internal priming.** Show how two molecular routes create similar mapped 3' tags.

**Table 127.3. 3' end assay artifacts and controls.** Diagnose site and usage errors.

| Failure mode | Observable signature | Control or orthogonal method | Reporting requirement |
| --- | --- | --- | --- |
| **Internal oligo(dT) priming** | Downstream genomic A-rich tract | Terminal ligation or untemplated-junction evidence | Filter definition and retained exceptions |
| **Ligation bias** | End-sequence-dependent recovery | Randomized terminal standards | Ligase, adapter, and bias estimate |
| **Adjacent sites merged** | Broad or bimodal cluster | Parameter sensitivity and targeted assay | Cluster radius and representative coordinate |
| **Site split by noise** | Unstable low-count peaks | Replicate-aware caller | Minimum support and replicate rule |
| **Wrong isoform assigned** | End shared by transcripts | Long-read or splice evidence | Alternative compatible transcripts |
| **Compositional usage shift** | One site rises as another denominator changes | Per-cell spike-in and total RNA | Absolute and within-gene denominators |

Validation combines different evidence classes. A targeted 3' RACE product can confirm a local junction but is itself sensitive to internal priming and amplification. A Northern blot can reveal whether the inferred isoforms have the expected size difference. Long reads can link a 3' end to upstream splicing if terminal completeness is independently established. Cleavage-factor perturbation can test mechanism, and genomic editing of a polyadenylation signal can test site dependence. Agreement among terminal chemistry, transcript architecture, and perturbation is stronger than deeper sequencing of one biased library.

## 127.4. Poly(A)-tail length, non-A tail composition, and molecule-resolved tail profiling

A poly(A) tail is an untemplated chain of adenosines added to an RNA 3' end, but it is neither chemically nor statistically simple. Tails can span from a few residues to hundreds, can contain terminal or internal non-A nucleotides, and can vary among molecules of the same transcript. The correct object is therefore a joint distribution of transcript identity, cleavage site, tail length, and tail composition. A gene-level median is a useful compression, not a complete representation.

Standard short-read sequencing performs poorly across long homopolymers because cluster generation, phasing, base discrimination, and alignment provide little sequence complexity. PAL-seq and TAIL-seq addressed this limitation with different measurement strategies. PAL-seq couples transcript identification to a calibrated tail-dependent signal. TAIL-seq uses paired-end sequencing and signal-level analysis of the tail-containing read to infer the boundary and composition of the terminal tract. The distinction matters: calibration, eligible terminal sequence, and non-A sensitivity differ, so numerical tail lengths from different methods should not be merged without benchmark evidence.

TAIL-seq revealed that mammalian mRNA tails can carry terminal uridines or guanosines in characteristic tail-length contexts. This finding illustrates why a tail is not necessarily a pure run of A. However, a detected non-A base can also arise from sequencing error, misidentified tail boundary, adapter sequence, internal priming, or genomic variation. Confidence increases when the signal is terminal, reproducible, biochemically dependent on a tailing enzyme, and observed with a chemistry that preserves terminal composition. mTAIL-seq improved sensitivity for lower-input developmental material, illustrating how a protocol revision can change the sampled transcript population even when its conceptual output is similar.

PAL-seq demonstrated that the relationship between tail length and translation differs across biological contexts, including a strong developmental coupling in early embryos and weaker coupling in many post-embryonic cells. The measurement lesson is that correlation is conditional on organism, developmental state, transcript population, and denominator. Tail length does not possess one universal functional mapping. Combining tail measurements with ribosome profiling, abundance, and kinetic data can discriminate whether a tail change precedes altered translation, stability, or both.

Nanopore direct RNA sequencing reads native polyadenylated molecules from the tail toward the body and can estimate poly(A) length from ionic-current segments. Its attraction is molecule-level linkage between a tail estimate and a long aligned RNA. Its limitations include selection for sufficiently polyadenylated molecules, incomplete reads, current segmentation, motor-speed variation, basecalling error, and transcript-assignment ambiguity. A direct RNA read is physically direct in the sense that RNA traverses the pore; the tail length remains a model-derived estimate. Synthetic standards of defined length and composition are required for calibration.

Molecule-resolved analysis should display distributions rather than only averages. A bimodal population of 20- and 100-nucleotide tails can have the same mean as a unimodal population near 60. Sampling depth varies by transcript, and very short tails may be systematically depleted. Hierarchical models can share information across replicates while preserving molecule-level uncertainty, but their assumptions must be reported. Tail-length thresholds used to label “short” and “long” should follow assay resolution and biology rather than round-number convention.

![Figure 127.4. Four views of a poly(A)-tail distribution](../assets/figures/chapter1158_figure4.png)

**Figure 127.4. Four views of a poly(A)-tail distribution.** Compare PAL-seq, TAIL-seq/mTAIL-seq, and nanopore signal-derived estimates while preserving molecule-level distributions.

**Table 127.4. Tail-profiling output is method-conditional.** Compare measurements without pretending equivalence.

| Method family | Tail observable | Principal strength | Principal limitation | Minimum calibration |
| --- | --- | --- | --- | --- |
| **PAL-seq** | Tail-dependent calibrated signal | Large-scale length profiling | Assay-specific signal conversion and terminal selection | Defined length standards across range |
| **TAIL-seq** | Tail-read signal and composition | Length plus terminal non-A information | Homopolymer boundary and input sensitivity | Length/composition standards |
| **mTAIL-seq** | Enhanced mRNA-focused tail profiling | Lower-input sensitivity | Changed transcript sampling relative to original workflow | Input titration and standards |
| **Nanopore direct RNA** | Signal dwell and molecule linkage | Native long molecule context | Segmentation, motor behavior, and poly(A) eligibility | Platform-matched tails of defined length |
| **Targeted PAT/RACE** | Locus-specific product size | Accessible validation | Priming and amplification selection | Known-size target standards |

The biology of polyadenylation, deadenylation, terminal uridylation, guanylation, and quality control belongs to [Chapter 29](chapter1028.md) and [Chapter 35](chapter1033.md). Long-read isoform reconstruction and direct RNA platform behavior are developed in [Chapter 128](chapter1117.md). This chapter owns the tail-specific calibration problem and the rule that eligible molecules, boundary calling, tail composition, and aggregation level must accompany every reported tail length.

## 127.5. Degradome, PARE, ribosome-protected ends, and targeted cleavage mapping

The degradome is not a homogeneous waste pool. It contains products of decapping, endonucleolytic cleavage, exonuclease progression, quality control, co-translational decay, and sample handling. PARE and related degradome methods make this space observable by selecting a terminal chemistry, commonly a free 5' monophosphate on an uncapped RNA, ligating an adapter, and sequencing the adjacent downstream fragment. The measured coordinate is the first retained nucleotide of that fragment, not the broken bond itself.

Plant microRNAs often direct Argonaute-catalyzed cleavage at a predictable register within a highly complementary target. PARE made it possible to match guide-target alignments with abundant 5' ends at the expected site across the transcriptome. The inference is strongest when the end is replicated, enriched over local background, dependent on the small RNA or silencing machinery, and compatible with guide pairing. A high PARE peak without these features may reflect ordinary decay. Conversely, a true cleavage can be missed if the product is rapidly degraded, has an ineligible end, is rare in the sampled tissue, or maps ambiguously.

Computational pipelines such as PAREsnip formalize pairing and category rules, but a category is not a universal probability of cleavage. Plant datasets differ in transcript abundance, degradome depth, guide expression, target complementarity, and background structure. Randomized controls, expression-aware nulls, replicate consistency, and validation of representative calls are necessary. Animal Argonaute cleavage is less general because imperfect guide pairing often leads to repression without slicing; method expectations must follow the biological system.

Other endonucleases create characteristic products but not always the chemistry selected by PARE. RNase L, tRNA anticodon nucleases, ribosome quality-control nucleases, and RNA-processing enzymes may leave hydroxyl, phosphate, or cyclic-phosphate termini. Enzyme repair or alternative adapter chemistry can expand access, but it also changes background. Paired-end strategies that recover both cleavage products or methods that map complementary terminal chemistries can strengthen bond-level assignment. Targeted Northern, primer extension, RACE, or ligation assays remain important orthogonal tests.

5PSeq selects 5'-phosphorylated degradation intermediates and uses their codon-phased positions to infer ribosome-associated protection during 5'-to-3' co-translational decay. A translating ribosome can impede exonuclease progression, leaving a downstream fragment whose 5' end occurs at a characteristic distance from the protected codon. Three-nucleotide periodicity and metagene offsets support this model. The assay does not isolate nuclease-protected ribosome footprints in vitro, and the signal depends jointly on translation, decapping, exonuclease activity, cleavage, and fragment lifetime.

This distinction prevents a common category error. Conventional ribosome profiling in [Chapter 120](chapter1114.md) sequences fragments protected from exogenous nuclease digestion in a ribosome preparation. 5PSeq sequences endogenous 5'-phosphorylated decay intermediates shaped by ribosome barriers in vivo. Both can show codon phase, but their denominators and artifacts differ. Drug treatment, stress, nuclease genotype, and RNA-decay rates can alter 5PSeq without an equivalent change in steady-state ribosome occupancy.

![Figure 127.5. Cleavage evidence ladder for PARE and 5PSeq](../assets/figures/chapter1158_figure5.png)

**Figure 127.5. Cleavage evidence ladder for PARE and 5PSeq.** Separate a chemistry-selected decay end from guide-directed cleavage and translation-coupled protection.

For any cleavage map, the causal ladder should be explicit: eligible end, reproducible coordinate, sequence or structural compatibility, genetic dependence, complementary product when detectable, and functional consequence. The cleavage and decay mechanisms themselves are covered in [Chapter 35](chapter1033.md), [Chapter 56](chapter1051.md), and enzyme-specific chapters. This chapter owns how the products are selected, mapped, scored, and kept distinct from mechanistic proof.

## 127.6. Peak calling, coordinate systems, isoform assignment, normalization, and uncertainty

End analysis begins with a coordinate convention. On the plus strand, the first aligned base of a downstream 5' fragment has a familiar increasing coordinate; on the minus strand, the analogous biological direction runs oppositely. A cleavage bond lies between the last nucleotide of the upstream product and the first nucleotide of the downstream product. A 3' end can be reported as the last templated base, the position immediately after it, or a half-open interval boundary. Software that silently mixes one-based closed and zero-based half-open systems creates apparent one-nucleotide disagreements.

Every dataset should document reference assembly, annotation version, strand convention, read-end definition, soft-clipping policy, adapter trimming, tail trimming, and the conversion from alignment coordinates to biological sites. Small-RNA and cleavage studies sometimes name a site by guide-relative position, while genome browsers display the retained nucleotide. A worked example on both strands is more reliable than a prose statement alone. Coordinate uncertainty should include not only mapping but also microheterogeneous cleavage and tail-boundary ambiguity.

Peak callers compress observed end coordinates into biological candidates. Parameters include minimum count, local-background model, clustering distance, replicate rule, relative dominance, mappability, and blacklist filters. Broad TSS clusters, tight polyadenylation clusters, and sparse degradome peaks require different models. Applying one generic peak caller can fragment a broad promoter, merge adjacent cleavage sites, or overcall isolated degradation ends. Sensitivity analyses across reasonable parameters reveal which conclusions depend on arbitrary thresholds.

Multimapping is a biological issue, not merely a nuisance. Repeated promoters, paralogous genes, transposable elements, rRNA arrays, and duplicated terminal exons can generate identical end tags. Discarding multimappers loses real biology; assigning them proportionally imports model assumptions. Reports should separate uniquely supported sites, probabilistically allocated signal, and unresolved families. Sequence variants and incomplete genomes can also shift or eliminate apparent ends.

Isoform assignment requires compatible upstream structure. A short 3' tag may map to a terminal exon shared by several transcripts; a cap tag may lie near overlapping genes on the same strand. Nearest-gene rules are convenient but fail in dense loci. Paired-end evidence, splice junctions, long reads, promoter-enhancer context, and perturbation can narrow the possibilities. [Chapter 141](chapter1128.md) owns reusable transcript models and isoform-level computational workflows; the end caller should export site-level evidence and uncertainty rather than conceal it behind one transcript label.

Normalization depends on the question. Counts per million mapped end tags estimate composition within the eligible library, not ends per cell. Within-gene site fractions estimate relative usage but are compositional. External spike-ins can reveal global changes if added per cell or per sample before the relevant loss, but different end chemistries require matched controls. Tail distributions need molecule-count and calibration-aware normalization rather than read-depth scaling alone. Degradome peaks may require comparison with total transcript abundance because more substrate can produce more decay ends without a higher cleavage fraction.

Replicates estimate biological and technical variability only if they are independent at the relevant level. Split libraries reveal library-preparation noise; independent cultures or organisms reveal biological variability. Confidence intervals should propagate count uncertainty, peak-boundary variability, tail-calling error, and model choice when they matter. Batch effects can alter cap recovery, ligation, read length, and signal calibration; randomization and balanced processing remain essential.

![Figure 127.6. Coordinate conventions on both genomic strands](../assets/figures/chapter1158_figure6.png)

**Figure 127.6. Coordinate conventions on both genomic strands.** Prevent one-base and strand-orientation errors.

**Table 127.5. Analysis decisions that change end calls.** Make computational uncertainty inspectable.

| Decision | Possible effect | Sensitivity analysis | Retained output |
| --- | --- | --- | --- |
| **Coordinate convention** | One-base or strand inversion | Worked plus/minus examples | Raw alignment and converted site |
| **Clustering radius** | Merge adjacent sites or split broad regions | Parameter sweep | Member coordinates and peak interval |
| **Multimapper treatment** | Lose repeats or import allocation model | Unique-only versus family-level result | Mapping class flag |
| **Annotation version** | Change assigned gene/isoform | Reannotation against second release | Site-level record independent of transcript label |
| **Count normalization** | Conceal global or compositional shifts | Spike-in, per-cell, library, and within-gene views | Numerator and denominator |
| **Tail threshold** | Relabel continuous distribution | Distribution-first analysis | Molecule-level calls and callability |

A useful output retains raw end counts, molecule counts after deduplication, site definitions, alternative assignments, quality flags, and the exact software version. This layered record allows future annotation changes without recreating the library. A single spreadsheet of “final genes” discards the most valuable property of end profiling: its direct, inspectable connection between terminal chemistry and genomic position.

## 127.7. Artifacts, orthogonal validation, reporting, and processing-pathway handoffs

Artifacts are best organized by stage. Before library construction, degradation creates new ends, deadenylation removes eligible tails, and extraction changes ribonucleoprotein accessibility. During chemistry, incomplete conversion, ligation bias, internal priming, reverse-transcription stops, and size selection distort eligibility. During sequencing, homopolymer error, index hopping, limited complexity, and platform calibration affect reads. During analysis, trimming errors, coordinate shifts, multimapping, clustering, annotation, and threshold choice affect sites. A control should be assigned to each plausible stage rather than presented as an undifferentiated checklist.

Internal priming deserves direct measurement. Candidate 3' sites should be annotated for downstream A-rich sequence, but sequence filtering should not be treated as truth. A terminal-ligation library, evidence of untemplated A, or targeted junction assay provides orthogonal chemistry. Cap-selected TSSs need conversion-minus controls and comparison with uncapped-end libraries. Degradome sites need total-RNA context and genetic or guide-sequence evidence. Tail measurements need standards spanning the relevant lengths and non-A compositions.

Orthogonal validation means changing the dominant artifact. Repeating one oligo(dT)-primed 3' protocol at greater depth does not address internal priming. Targeted RACE changes scale but may retain oligo(dT) and amplification artifacts; a Northern blot adds size information; direct ligation changes end selection; long-read evidence changes linkage; perturbing the processing factor tests causality. The strongest validation set is designed around the leading alternative explanation.

Reporting should include input amount and integrity; enrichment or depletion; every terminal conversion enzyme; adapter and ligase; size selection; reverse transcriptase; PCR cycles and molecular identifiers; sequencer; read structure; trimming; mapping; coordinate convention; site caller; normalization; replicate definition; spike-in composition and addition stage; and exclusion filters. For tails, report calibration standards, callable range, treatment of non-A residues, and whether values are molecule-level or aggregated. For degradomes, report eligible 5' chemistry and how substrate abundance enters the score.

The handoff to biology must remain disciplined. A cap cluster is a candidate initiation region; transcription-factor and promoter mechanisms belong to [Chapter 26](chapter1025.md). A 3' cluster is a candidate cleavage/polyadenylation site; processing-factor mechanisms belong to [Chapter 29](chapter1028.md). A short tail is an observed molecular state; deadenylation and translation consequences belong to [Chapter 35](chapter1033.md) and [Chapter 72](chapter1067.md). A degradome peak is a candidate cleavage or decay intermediate; nuclease and small-RNA mechanisms belong to [Chapter 56](chapter1051.md) and decay chapters. This separation permits cross-chapter integration without converting measurement into mechanism by wording alone.

> **Box 127.1. Audit an RNA-end claim from molecule to mechanism**
>
> - Questions: Which terminal state was eligible? Which conversions occurred? What coordinate is reported? Could priming, ligation, degradation, or mapping create the same peak? What denominator was used? Which molecules were not callable? Does the evidence establish an end, an isoform, or a mechanism? Which orthogonal method changes the dominant artifact?
> - Output: A six-row audit form: analyte; chemistry; coordinate; alternative; uncertainty; biological handoff.
> - Misconception prevented: A high, precise peak automatically identifies its generating pathway.

## Experimental Foundations and Evidence

The evidentiary history of end profiling is a history of replacing indirect abundance with molecular boundaries while discovering new selection biases. CAGE established that short cap-proximal tags could produce promoter-level expression maps at scale; protocol refinements and comparisons with RAMPAGE, chromatin assays, and transcript annotation then showed where cap maps agree and where their promoter assignments remain conditional. This evidence class is strongest for the existence and relative use of capped terminal regions. It is weaker for polymerase recruitment kinetics or the downstream fate of each initiated molecule, which require nascent transcription, perturbation, and decay measurements.

Three-prime methods followed a similar trajectory. Conventional RNA-seq inferred terminal exons from coverage, but dedicated junction capture localized candidate cleavage sites. Synthetic constructs and sequence-context analysis demonstrated internal oligo(dT) priming; cross-method benchmarks showed that protocols and callers recover overlapping but nonidentical site sets. A credible site therefore requires more than database overlap. Replicate terminal reads, an untemplated junction or independent end chemistry, plausible processing context, and targeted confirmation form an evidence ladder. Perturbing a cleavage factor or polyadenylation signal adds mechanism but may also change transcription and stability, so downstream consequences need separate measurement.

Tail profiling required calibration because ordinary sequence calls cannot reliably count long homopolymers. PAL-seq and TAIL-seq used different physical observables and synthetic standards, while biological perturbations of developmental state or tailing enzymes connected the measurements to function. Cross-platform reviews show that methods differ in input requirement, callable range, transcript linkage, amplification, and sensitivity to non-A residues. Agreement at a broad distributional level is reassuring; disagreement is diagnostic when standards reveal length-dependent recovery. Tail claims should therefore report both biological replicates and calibration performance.

PARE evidence is unusually intuitive—a sharp downstream end at the predicted slicing register—but also vulnerable to mechanistic overreach. Small-RNA loss, Argonaute perturbation, guide-target complementarity, and complementary product detection strengthen cleavage assignments because they test independent consequences of the proposed mechanism. 5PSeq added an aggregate evidence pattern: codon periodicity and offsets expected from ribosome protection. Random-fragmentation controls, nuclease mutants, and comparisons with ribosome profiling help determine whether periodicity arises from translation-coupled decay rather than sequence or alignment structure.

## Biological Contexts Across Organisms, Cell Types, and Perturbations

The same method name can observe different biology across organisms because terminal states and processing pathways differ. In mammals, cap-selected maps resolve focused and broad promoters, while tissue mixtures can make an apparent alternative TSS reflect changing cell composition. A sorted population or single-cell assay can narrow that explanation, but lower molecule counts then increase sampling uncertainty. In bacteria, primary transcripts commonly carry 5' triphosphates rather than eukaryotic caps, so differential enzymatic conversion rather than cap trapping is the relevant selection logic. Organellar and viral transcripts add their own caps, triphosphates, protein-linked ends, or processing states and require organism-specific chemistry controls.

Alternative polyadenylation is prominent in development, differentiation, neuronal activation, immunity, and cancer, but the same proximal-site shift can have different causes. A proliferating cell population may change cleavage-factor availability, transcription elongation, or cell-cycle composition. Neurons can use long distal UTRs with localization elements, whereas a short-read 3' tag alone cannot demonstrate localization. Plant polyadenylation signals and site architecture differ from mammalian conventions, so applying one motif filter across kingdoms can suppress genuine sites.

Tail biology is especially context-sensitive. In early vertebrate embryos and oocytes, cytoplasmic polyadenylation can couple tail length strongly to translation before zygotic transcription dominates. In many somatic contexts, tail length relates more directly to decay state and shows a weaker simple correlation with translation. Stress can redistribute RNA between translation, storage, and decay pathways, changing both tail distributions and which molecules survive extraction. Sampling time is therefore part of the analyte, and kinetic methods in [Chapter 129](chapter1154.md) are required when temporal ordering matters.

Plant PARE benefits from frequent near-perfect guide-target pairing and slicing, whereas many animal microRNAs act without endonucleolytic cleavage. Fungi provided a tractable system for 5PSeq because genetic perturbation of decay enzymes and strong codon-phase signals could test the co-translational-decay model. In every organism, a null result is chemistry-conditional: absence of a PARE peak can mean no cleavage, rapid product removal, an incompatible terminal state, low expression, or inadequate depth.

## Technology, Computational, Clinical, and Engineering Links

End profiles support genome annotation, promoter design, therapeutic RNA engineering, and biomarker development. Promoter atlases can refine transcriptional regulatory elements and expose cell-type-specific starts. Three-prime maps identify UTR boundaries needed to design perturbations and reporters. Tail profiling informs synthetic mRNA manufacturing because tail length and composition affect translation, stability, and product heterogeneity. These applications raise the standard of quantification: a discovery assay that ranks sites may be insufficient for release testing, where calibrated accuracy, predefined acceptance criteria, and traceable standards are required.

Clinical specimens magnify preanalytical effects. Ischemia, fixation delay, freeze-thaw cycles, endogenous nucleases, and variable cell composition can create or remove ends before the assay begins. A disease-associated cleavage peak is not clinically useful until its stability, matrix dependence, specificity, and denominator are established in independent cohorts. Liquid-biopsy RNA fragments can be informative precisely because they are protected or processed, but that enrichment also makes their relation to tissue abundance indirect.

Computationally, end data are natural candidates for layered, versioned representations. A site record should retain chemistry, coordinate convention, read evidence, replicate support, peak interval, alternative transcript assignments, and caller version. Knowledge graphs can connect the site to promoters, polyadenylation signals, guides, enzymes, and phenotypes without collapsing observation into mechanism. Training and benchmark datasets need hard negatives such as internal-priming sites, recapped ends, paralogous mappings, and cleavage-like decay peaks; otherwise a model learns sequence proximity rather than evidence discrimination.

Engineering a new assay should begin with standards that span the intended failure space. A terminal-chemistry panel tests conversion; randomized adjacent bases test ligation; structured constructs test accessibility; defined tails test length and composition; mixed isoforms test assignment; and dilution series test quantitative response. An assay can then be compared with an established method on the same molecules rather than on unrelated biological samples. This design makes technology improvement measurable and clarifies whether added yield comes from genuine broader eligibility or from increased background.

## Cross-Method Synthesis: Select the Assay from the Molecular Claim

The most useful comparison begins with the molecule. To ask where capped polymerase-II products begin, use cap-selected 5' profiling with controls for cap selection and recapping. To ask where polyadenylated transcripts were cleaved, use terminal 3' capture with internal-priming controls. To ask how long individual tails are, use calibrated tail-specific or molecule-resolved measurements. To ask where 5'-monophosphorylated decay products accumulate, use degradome chemistry and interpret peaks within substrate and decay context. No universal “end-seq” library can answer all four questions without destroying the distinctions that make the answers meaningful.

Method combinations are valuable when their eligibility sets overlap in a known way. Cap-selected and 5'-phosphate libraries can contrast capped and uncapped populations. 3' site maps and long reads can link a boundary to an isoform. Tail measurements and kinetic labeling can relate tail state to RNA age. PARE and small-RNA profiles can connect candidate cleavage products to guides. The joint analysis must preserve the selection rule of each input rather than treating every read as an equivalent observation.

## Recent Consensus

End profiling is now understood as chemistry-conditioned measurement. Cap capture provides high-resolution promoter evidence; terminal 3' libraries provide more direct cleavage-site evidence than ordinary coverage; dedicated tail methods reveal distributions and non-A composition that bulk oligo(dT) assays miss; and degradome methods identify chemically eligible decay products rather than all degradation events. Across these areas, current practice favors process-matched spike-ins, replicate-aware peak calling, explicit coordinate conventions, and validation by a method with a different dominant artifact.

Alternative polyadenylation is widespread and context-dependent, but site usage must be separated from site identity and from total transcript abundance. Tail length can regulate translation and stability, yet the relationship differs strongly across developmental and cellular contexts. A precise end coordinate is not automatically a mechanistic event: cap status, cleavage chemistry, and downstream processing are needed to distinguish initiation, maturation, and decay.

## Open Questions, Controversies, Deprecated Models, and Common Misconceptions

Open questions:

- How accurately can different end-conversion reactions be calibrated across the diversity of native RNA structures and ribonucleoprotein states?
- Which combinations of short-read end maps and long direct RNA reads provide unbiased molecule-level linkage among TSS, splicing, cleavage site, and tail state?
- How much apparent cell-type-specific promoter or APA usage reflects true regulation versus RNA stability, cell composition, and protocol-specific eligibility?
- Can tail-length and non-A-tail callers be standardized across sequencing platforms without erasing method-specific uncertainty?
- How completely do currently sampled 5'-phosphate degradomes represent endonucleolytic cleavage when products have short lifetimes or alternative terminal chemistries?

Controversies:

- The extent to which single capped intragenic ends represent alternative initiation, cytoplasmic recapping, or stable processed products remains locus- and context-dependent.
- Gene-level tail-length summaries are biologically useful but can obscure multimodal molecular populations; the appropriate default summary remains analysis-dependent.
- Sequence-context filters for internal priming improve specificity but can reduce sensitivity at genuine A-rich cleavage sites, so no universal filter threshold is accepted.

Deprecated or weakened claims:

- A unique annotated TSS or 3' end should not be assumed to represent every transcript from a gene; distributed initiation and alternative cleavage are common.
- Poly(A)-tail length is not a universal proxy for translational efficiency across all cells and developmental stages.
- A degradome peak is not by itself proof of a specific nuclease reaction.

Common misconceptions:

- “Every sequenced 5' end is a transcription start site.” A sequenced end is an eligible molecule after selection and may result from processing, cleavage, decapping, recapping, or damage.
- “Oligo(dT)-primed reads prove a natural 3' end.” Oligo(dT) can prime within internal A-rich sequence, and terminal support or orthogonal chemistry is needed.
- “A direct RNA tail length is read without a model.” Nanopore measures ionic current from native RNA, but segmentation and calibration infer tail length.
- “PARE measures all RNA degradation.” Standard PARE enriches a chemically defined subset, commonly uncapped 5'-monophosphorylated products.
- “One-nucleotide disagreements always indicate biological heterogeneity.” They can arise from whether a pipeline reports the retained nucleotide, cleavage bond, or interval boundary.
- “More reads eliminate ligation bias.” Depth reduces sampling error among admitted molecules but does not recover molecules excluded by end chemistry or library bias.
