# Chapter 26. 5' Capping, Cap Diversity, Cap-Binding Complexes, and Decapping

## Scope Note

The 5′ end of an RNA molecule is a small chemical region with large biological consequences. It can identify the polymerase that made the RNA, recruit processing and export factors, control translation initiation, alter RNA stability, mark self versus nonself RNA, or expose the RNA to decay enzymes. The canonical eukaryotic N7-methylguanosine cap is the central example in this chapter, but it is not the only relevant 5′ end. RNAs can carry cap0, cap1, cap2, cap-adjacent m6Am, metabolite-linked caps such as NAD-RNA and dpCoA-RNA, viral protein-linked or cap-snatched ends, triphosphate ends, diphosphate ends, monophosphate ends, or incompletely capped quality-control intermediates.

This chapter explains how 5′ caps are installed, read, modified, removed, measured, and engineered. It gives enough context about transcription, export, translation, innate immunity, viral replication, RNA decay, and mRNA therapeutics to make cap biology understandable, while leaving complete treatment of those neighboring subjects to Chapters [21](chapter1020.md), [25](chapter1024.md), [30](chapter1029.md), [32](chapter1031.md), [35](chapter1033.md), [66](chapter1061.md)-[72](chapter1067.md), [108](chapter1103.md), [115](chapter1109.md)-[120](chapter1114.md), [132](chapter1120.md), [153](chapter1137.md), [156](chapter1139.md), and [159](chapter1142.md).

## Executive Summary

RNA 5′ ends are molecular identity tags. In canonical eukaryotic messenger RNA biogenesis, a nascent RNA polymerase II transcript emerges with a 5′ triphosphate. A capping apparatus converts that raw end into a 7-methylguanosine cap joined to the first transcribed nucleotide through an unusual 5′-to-5′ triphosphate bridge. The standard cap0 pathway has three core chemical steps: an RNA 5′ triphosphatase removes the gamma phosphate, an RNA guanylyltransferase transfers GMP through a covalent enzyme-GMP intermediate, and a guanine-N7 methyltransferase methylates the added guanine. Additional ribose methylation on the first and second transcribed nucleotides produces cap1 and cap2.

Capping is also a timing event. In many nuclear RNA polymerase II systems, capping begins early during transcription and is coordinated with the transcription machinery, including recruitment to phosphorylated forms of the polymerase C-terminal domain in well-studied eukaryotes. The cap therefore becomes one of the first durable marks of a pre-mRNA. It helps define the transcript as a particular kind of ribonucleoprotein particle, or mRNP, before splicing, 3′-end formation, export, translation, or decay are complete.

The canonical m7G cap is not the whole field. Noncanonical cap-like structures include NAD-RNA, FAD-RNA, ADPR-RNA, dpCoA-RNA, UDP-GlcNAc-RNA, and other metabolite-linked 5′ ends. These structures are chemically distinct from cap0, cap1, and cap2. Some can be installed when an RNA polymerase initiates transcription with a metabolite such as NAD or NADH instead of a standard nucleoside triphosphate; other cases imply post-transcriptional chemistry or processing-dependent formation. Their functions are unevenly established, so abundance, installation route, enzyme sensitivity, and biological consequence must be evaluated separately for each organism and RNA class.

Cap-binding proteins convert cap chemistry into RNA fate. In the nucleus, the cap-binding complex links the m7G cap to pre-mRNA processing, export, surveillance, and early mRNP transitions rather than merely shielding the RNA end. In the cytoplasm, eIF4E and the eIF4F-centered initiation apparatus connect the cap to ribosome recruitment, scanning, and regulated translation. Cap structure also influences innate immune discrimination: cap1 and related ribose methylation states can help distinguish many cellular RNAs from viral, triphosphorylated, or incompletely matured RNAs.

Decapping is the removal of a cap or cap-like 5′ structure. The DCP2 enzyme is the main catalytic subunit for canonical cytoplasmic mRNA decapping in eukaryotes, and DCP1 plus additional cofactors regulate DCP2 in species-specific assemblies. Decapping often commits an mRNA to 5′-to-3′ degradation by exposing an end that exonucleases can attack, but decapping is regulated by deadenylation, translation repression, nonsense-mediated decay, miRNA-associated repression, stress pathways, and cytoplasmic RNP organization. Other enzymes, including DXO/Rai1-family enzymes and selected Nudix hydrolases, participate in cap quality control and noncanonical cap removal.

Cap analysis and cap engineering require methodological caution. A method that maps 5′ ends does not necessarily identify cap chemistry, and a method that identifies cap chemistry may lose transcript identity. Noncanonical cap analysis therefore requires orthogonal evidence from chemistry, enzymology, sequencing, and mass spectrometry. In synthetic mRNA and therapeutic RNA production, capping route and cap structure affect cap orientation, translation, stability, immune recognition, impurity control, and manufacturing analytics.

## Concept Inventory

- **5′ cap:** a covalent RNA-end structure that can protect RNA, recruit proteins, influence processing or translation, and mark RNA identity. The term includes canonical m7G caps and several cap-like noncanonical ends. A cap is not automatically evidence that an RNA is translated.
- **M7G cap:** a canonical eukaryotic cap in which N7-methylguanosine is linked to the first transcribed nucleotide through a 5′-to-5′ triphosphate bridge. It is often written as m7GpppN. The related structures GpppN, cap1, cap2, and trimethylguanosine caps must be distinguished from m7G cap0.
- **Cap0:** m7GpppN without ribose 2′-O methylation on the first transcribed nucleotide. It can support many cap-binding functions, but in some immune contexts cap0 differs from cap1.
- **Cap1:** an m7G cap in which the first transcribed nucleotide carries a ribose 2′-O methyl group. Cap1 is a structural self-RNA mark in many animal and viral contexts, not merely a translation-enhancing decoration.
- **Cap2:** an m7G cap in which the first and second transcribed nucleotides carry ribose 2′-O methyl groups. Cap2 biology is less broadly characterized than cap1 biology and should not be generalized beyond the evidence.
- **M6Am:** N6,2′-O-dimethyladenosine at the first transcribed nucleotide when that nucleotide is adenosine. It is adjacent to the cap but is not the 5′ cap bridge itself, and it should not be confused with internal m6A.
- **Capping enzyme or capping apparatus:** converts a nascent RNA 5′ triphosphate or diphosphate into a capped RNA end. Eukaryotic nuclear, viral, organellar, and phage-associated capping systems are not all homologous.
- **Nuclear cap-binding complex:** often described as CBP20/CBP80 or NCBP2/NCBP1, binds the m7G cap and helps couple the 5′ end to processing, export, surveillance, and mRNP handoff.
- **EIF4E:** a cytoplasmic cap-binding translation initiation factor. In many eukaryotic mRNAs, eIF4E binds the m7G cap and functions with eIF4G and eIF4A in the eIF4F initiation assembly.
- **NAD-RNA:** RNA with nicotinamide adenine dinucleotide covalently linked at the 5′ end. NAD-RNA claims require direct RNA-end evidence; general NAD metabolism is not evidence for NAD-capped RNA.
- **Decapping:** enzymatic removal or hydrolysis of a 5′ cap or cap-like structure. It often creates an RNA end that is vulnerable to 5′-to-3′ decay, but cap removal can also function in quality control or cap recycling.
- **DCP2:** the conserved Nudix-family catalytic subunit of the major eukaryotic cytoplasmic mRNA decapping enzyme.
- **Nudix hydrolases:** enzymes with a conserved Nudix motif that hydrolyze nucleoside diphosphates linked to other chemical groups. Nudix-family membership alone does not prove an enzyme acts on RNA caps in vivo.
- **DXO/Rai1-family proteins:** cap quality-control and exonuclease enzymes that can act on incompletely capped or noncanonically capped RNAs, with substrate range varying by organism and enzyme.
- **Cap-snatching:** a viral strategy in which capped fragments from host RNAs are used as primers for viral transcription. It is distinct from viral capping and from viral decapping.
- **Synthetic mRNA capping:** uses enzymatic, co-transcriptional, chemical, or chemoenzymatic strategies to produce capped RNA in vitro. It is one component of therapeutic RNA design, not the whole platform.
- **Cap analysis:** measures RNA 5′ cap structure, abundance, sequence assignment, or product quality. Transcription start-site mapping, chemical cap identification, and stoichiometric cap quantification are related but different tasks.

## What to Know Before Reading This Chapter

RNA chains have polarity. Nucleotides are joined inside RNA by phosphodiester bonds that connect the 3′ carbon of one ribose to the 5′ phosphate of the next nucleotide. RNA polymerases extend RNA by adding new nucleotides to the 3′ end, so the first nucleotide incorporated during transcription remains at the 5′ end. That 5′ end is chemically exposed before the rest of the RNA has been synthesized.

Many newly synthesized RNAs begin with a 5′ triphosphate. The first nucleotide keeps three phosphates because it entered the polymerase as a nucleoside triphosphate. By contrast, internal nucleotides lose pyrophosphate during chain elongation and become part of ordinary phosphodiester linkages. This difference gives enzymes a way to identify the beginning of a transcript. A 5′ triphosphate end, a 5′ diphosphate end, a 5′ monophosphate end, and a capped end can be recognized by different proteins and decay enzymes.

The word "cap" should be used chemically, not vaguely. The canonical eukaryotic m7G cap is a methylated guanosine linked backward to the first transcribed nucleotide through a 5′-to-5′ triphosphate bridge. The notation m7GpppN means N7-methylguanosine, three bridging phosphates, and the first transcribed nucleotide. If the first nucleotide has a ribose 2′-O methyl group, the structure is cap1; if the second nucleotide also has that ribose methylation, the structure is cap2. If the first nucleotide is adenosine and also carries N6 methylation, the cap-adjacent nucleotide can be m6Am.

Three verbs organize the chapter. Capping installs or incorporates a cap. Cap binding reads a cap through proteins such as the nuclear cap-binding complex or eIF4E. Decapping removes a cap or cap-like structure. These phases can occur on the same RNA molecule, but they are not the same process and they do not use the same evidence.

The main running example is a metazoan RNA polymerase II mRNA. Its 5′ end is capped early, bound by nuclear cap-binding proteins, coupled to processing and export, handed off to cytoplasmic translation factors, and eventually decapped during decay. Boundary examples include bacterial and archaeal NAD-RNAs, viral RNAs that synthesize or steal caps, snRNAs with specialized cap handling, and synthetic mRNAs whose caps are engineered during manufacturing.

## 26.1. Capping enzyme chemistry and timing

**Table 26.1. Cap Structures and Evidence Confidence.** Known and proposed RNA 5′ cap structures, their chemical signatures, installation routes, biological contexts, decapping enzymes, and functional evidence.

| Cap structure | Chemical signature | Installation route | Organism / RNA class | Decapping enzymes | Functional evidence |
| --- | --- | --- | --- | --- | --- |
| **Cap0 (m7GpppN)** | N7-methylguanosine via 5′-to-5′ triphosphate bridge; no ribose methylation | Triphosphatase, guanylyltransferase, N7 methyltransferase | Eukaryotic Pol II mRNA; broad eukaryotic distribution | DCP2/DCP1; scavenger decappers | CBC and eIF4E binding; processing; export; translation initiation |
| **Cap1 (m7GpppNm)** | Cap0 plus 2′-O-methyl on first transcribed nucleotide | CMTR1-family 2′-O-methyltransferase after cap0 formation | Metazoan mRNA; viral mRNAs mimicking host | DCP2/DCP1 | Self-RNA mark; reduced IFIT-family and innate immune recognition versus cap0 |
| **Cap2** | Cap1 plus 2′-O-methyl on second transcribed nucleotide | CMTR2-family 2′-O-methyltransferase | Metazoan mRNA | DCP2/DCP1 | Additional ribose methylation marker; less broadly characterized than cap1 |
| **m6Am (cap-adjacent)** | N6,2′-O-dimethyladenosine at first position; not part of 5′-to-5′ bridge | PCIF1 methyltransferase when first nucleotide is adenosine | Metazoan Pol II mRNA with adenosine at +1 | DCP2/DCP1 acts on cap; m6Am is cap-adjacent | May modulate mRNA stability; distinct from internal m6A; requires positional resolution |
| **NAD-RNA** | NAD covalently at 5′ end via adenosine moiety | Transcription initiation with NAD or NADH instead of ATP; post-transcriptional routes also possible | Bacteria, eukaryotes, mitochondria, archaea | NudC, DXO/Rai1-family, NUDT-family hydrolases | Context-dependent; deNADding characterized; protection or decay marking varies by organism |
| **ADPR-RNA** | ADP-ribose at 5′ end | Transcription initiation or post-transcriptional; route not fully resolved | Archaea; bacteria | NUDT-family; Rai1-like enzymes | Detected in archaea; linked to RNA turnover; functional model developing |
| **FAD-RNA** | FAD covalently at 5′ end | Transcription initiation with FAD in place of ATP | Bacteria; possible eukaryotic contexts | Nudix hydrolases | Chemical identity established; cellular regulatory function under investigation |
| **dpCoA-RNA** | Coenzyme A at 5′ end via diphosphate linkage | Transcription initiation with CoA precursor | Bacteria | NudC-like proteins | NudC cleaves dpCoA-RNA in vitro; cellular function developing; cleavage position matters for product identity |
| **UDP-GlcNAc-RNA** | UDP-GlcNAc sugar at 5′ end | Route under active study | Human and possibly other eukaryotes | hNudt5 | Chemical detection and in vitro enzyme activity demonstrated; cellular regulatory function remains under active study |
| **Trimethylguanosine cap (TMG)** | 2,2,7-trimethylguanosine at 5′ end; built on earlier m7G cap | Post-transcriptional hypermethylation by TGS1 methyltransferase | Metazoan snRNAs, snoRNAs; trypanosome mRNA (cap4 variant) | Specialized or context-dependent | snRNP assembly and nuclear reimport signal; cap4 in trypanosomes is distinct from metazoan TMG |

Capping begins with a problem created by transcription itself. A new RNA polymerase II transcript has a chemically reactive 5′ triphosphate, but a mature eukaryotic mRNA normally needs a protected 5′ end that can recruit mRNP factors and avoid being mistaken for aberrant RNA. The canonical capping pathway solves this by converting the first nucleotide into the base of an unusual terminal structure.

![Figure 26.1. Canonical Capping Chemistry and Timing](../assets/figures/chapter1025_figure1.png)

**Figure 26.1. Canonical Capping Chemistry and Timing.** A nascent RNA polymerase II transcript emerges from the elongation complex carrying a 5′ triphosphate end. The canonical cap0 pathway converts this end in three sequential chemical steps: an RNA 5′ triphosphatase removes the gamma phosphate to yield a 5′ diphosphate RNA; an RNA guanylyltransferase transfers GMP through a covalent enzyme-GMP intermediate, creating the unusual 5′-to-5′ triphosphate bridge; and a guanine-N7 methyltransferase uses S-adenosylmethionine to methylate the added guanine, producing m7GpppN or cap0. Subsequent ribose 2′-O methylation of the first transcribed nucleotide by a CMTR1-like enzyme produces cap1, and additional methylation at the second transcribed nucleotide produces cap2. Transcripts that fail to acquire a complete or properly methylated cap are recognized by cap quality-control enzymes such as DXO/Rai1-family proteins and directed to surveillance pathways.

![Figure 26.2. Cap-Binding Handoffs Across the mRNA Life Cycle](../assets/figures/chapter1025_figure2.png)

**Figure 26.2. Cap-Binding Handoffs Across the mRNA Life Cycle.** A capped pre-mRNA is first bound in the nucleus by the nuclear cap-binding complex, which couples the 5′ end to pre-mRNA processing, export factor recruitment, surveillance, and early mRNP identity. During and after export to the cytoplasm, eIF4E within the eIF4F initiation assembly takes over cap recognition and promotes 40S ribosomal subunit recruitment, scanning, and start-codon selection. Throughout the mRNA life cycle the cap represents a competed binding site: as translation declines, deadenylation progresses, or decay factors gain access, DCP2-centered decapping becomes more likely and commits the mRNA to 5′-to-3′ exonucleolytic degradation. The figure emphasizes that cap binding is a series of regulated handoffs rather than a static protective interaction.

![Figure 26.3. Canonical and Noncanonical Decapping Routes](../assets/figures/chapter1025_figure3.png)

**Figure 26.3. Canonical and Noncanonical Decapping Routes.** Multiple enzymatic routes remove the cap or cap-like structure from an RNA 5′ end. DCP2, the catalytic Nudix subunit of the major eukaryotic decapping enzyme, is activated by DCP1 and additional cofactors including EDC proteins, Pat1, LSM complexes, and DDX6-like helicases; its activation is coupled to translational repression, deadenylation, and regulatory RNA-protein assemblies. DXO/Rai1-family enzymes target incompletely capped or noncanonically capped RNAs as part of cap quality control, with exonuclease and decapping activities that vary between metazoan DXO and yeast Rai1. Nudix-family enzymes other than DCP2, including NudC-like proteins and selected NUDT proteins, remove metabolite-linked caps such as NAD, FAD, dpCoA, and UDP-GlcNAc from the RNA 5′ end. Large DNA viruses such as orf virus encode their own Nudix-family decappers, exemplified by OV71, which can decap capped RNA substrates and support viral replication.

The cap0 pathway has three core reactions. First, an RNA 5′ triphosphatase removes the gamma phosphate from the transcript's 5′ triphosphate, leaving a 5′ diphosphate RNA. Second, an RNA guanylyltransferase reacts with GTP to form a covalent enzyme-GMP intermediate and then transfers GMP to the diphosphate RNA. This transfer produces the 5′-to-5′ triphosphate bridge, a linkage that is chemically different from the 3′-to-5′ phosphodiester bonds inside the RNA chain. Third, a guanine-N7 methyltransferase uses S-adenosylmethionine as a methyl donor to methylate the added guanine base, producing m7GpppN, or cap0.

This sequence of reactions explains why "capping enzyme" can mean different things in different organisms. In metazoans, the RNA 5′ triphosphatase and guanylyltransferase activities are combined in a bifunctional enzyme, commonly referred to as RNGTT in humans, while the guanine-N7 methyltransferase activity is provided by RNMT and its associated regulatory proteins. Fungi, protists, plants, and viruses can arrange related activities in different domain architectures or different proteins. The chemistry required to create cap0 is conserved at the level of transformations, but the protein architecture is not universal.

Capping is usually early. In well-studied nuclear RNA polymerase II systems, capping begins after a short nascent RNA has emerged from the polymerase and before the full pre-mRNA has been synthesized. This early timing depends partly on the transcription complex. The C-terminal domain of the largest Pol II subunit is phosphorylated in changing patterns during initiation and elongation, and capping enzymes interact with early elongation states of Pol II in several model systems. The result is a kinetic checkpoint: a transcript that is being productively elongated receives an early 5′ identity mark.

The biological logic of early capping is not limited to protection from exonucleases. The cap creates a binding site for nuclear cap readers, helps define an RNA as a Pol II product, and can influence downstream processing. A newly capped pre-mRNA is more than a naked RNA with a protected end; it is beginning to become an mRNP. [Chapter 25](chapter1024.md) develops the broader idea that RNA processing begins during transcription. Here the point is narrower: capping is one of the first stable chemical differences between a productive Pol II transcript and an exposed triphosphorylated RNA end.

Incomplete capping creates quality-control substrates. A transcript could fail triphosphatase action, receive guanylate but fail N7 methylation, carry an unusual 5′ end, or lose cap-binding protection. Such RNAs can be recognized by quality-control enzymes, including DXO/Rai1-family proteins in some eukaryotic systems. This is why cap state must be specified. GpppN, m7GpppN, cap1, cap2, NAD-RNA, and 5′ triphosphate RNA have different chemical signatures and can have different biological fates.

Viruses reveal that capping chemistry is a recurrent evolutionary problem. Many viruses that make mRNA-like transcripts must either synthesize a cap, mimic host cap methylation, steal capped primers from host RNAs, or use a translation strategy that bypasses ordinary cap dependence. Viral capping enzymes may be multifunctional and embedded within replication proteins, and viral cap methyltransferases can be antiviral drug targets in selected systems. The same chemical endpoint, a host-like cap, can therefore arise by different routes.

The main misconception in this section is that capping is a late finishing step. For many Pol II mRNAs, cap installation is an early, co-transcriptional event with consequences for the subsequent mRNP pathway. Another common overgeneralization is to assume that all capping enzymes are equivalent because they make a cap. Domain organization, recruitment, substrate specificity, and inhibitor sensitivity vary enough that capping systems must be compared explicitly.

## 26.2. Cap0, Cap1, Cap2, NAD caps, and noncanonical 5-prime ends

Cap0, cap1, and cap2 describe a chemical series built on the canonical m7G cap. Cap0 is m7GpppN. Cap1 is m7GpppNm, where the first transcribed nucleotide has a ribose 2′-O methyl group. Cap2 carries ribose 2′-O methylation on both the first and second transcribed nucleotides. These are not interchangeable terms for a mature cap. They mark specific positions and methylation states, and those positions can be read by proteins, viral immune evasion enzymes, innate immune sensors, or analytical assays.

![Figure 26.4. Noncanonical Cap Detection Decision Tree](../assets/figures/chapter1025_figure4.png)

**Figure 26.4. Noncanonical Cap Detection Decision Tree.** No single method simultaneously resolves cap chemical identity, transcript assignment, and stoichiometry. Liquid chromatography-tandem mass spectrometry after nuclease digestion provides the highest chemical specificity and can distinguish cap0, cap1, cap2, NAD, FAD, dpCoA, and UDP-GlcNAc caps, but nuclease digestion destroys most sequence context. NAD captureSeq, NAD tagSeq, and SPAAC-NAD-Seq connect cap chemistry to specific transcripts at genome scale but each depends on reaction completeness, enzyme specificity, or labeling efficiency. APB gel electrophoresis and CapZyme-style decapping-to-ligation assays provide independent enzymatic validation orthogonal to sequencing. Artifact checkpoints at each decision node include incomplete enrichment, enzyme promiscuity, ligation bias, RNA fragmentation, and abundance confounding; a robust noncanonical cap study uses at least two independent methods and includes controls for uncapped, triphosphorylated, and chemically related RNA species.

Ribose methylation adds another layer to the 5′-end code. Cap1 methylation in metazoans is associated with enzymes such as CMTR1, and cap2 methylation with CMTR2. Direct biochemical, structural, and genetic sources now support CMTR1, CMTR2, and cap-proximal ribose methylation rather than leaving them as generic capping-literature examples. The functional point is that first- and second-nucleotide ribose methylation change how the RNA is interpreted. In antiviral contexts, a cap that looks sufficiently host-like can reduce recognition by self-nonself discrimination systems, whereas cap0 or triphosphate ends can be more immunostimulatory.

Cap-adjacent m6Am is related but distinct. If the first transcribed nucleotide is adenosine and the ribose is already 2′-O methylated as part of cap1, N6 methylation of the adenine base produces m6Am. This mark is adjacent to the m7G cap rather than part of the 5′-to-5′ bridge. It should also be distinguished from internal m6A elsewhere in the transcript. A method that detects N6-methyladenosine without positional resolution can confuse cap-adjacent and internal marks; a method that maps the first nucleotide can separate them.

Lineage-specific caps prevent a mammalian default model from becoming too narrow. Some small nuclear RNAs acquire trimethylguanosine caps. Trypanosome splice-leader RNAs and trans-spliced mRNAs can carry highly methylated cap4 structures. Viral RNAs can carry host-like caps, unusual methylation states, protein-linked ends, or cap-snatched fragments depending on virus family. The phrase "mature cap" therefore needs a species, RNA class, and chemical description.

Noncanonical caps broaden the definition of a cap-like RNA end. NAD-RNA, FAD-RNA, dpCoA-RNA, UDP-GlcNAc-RNA, ADPR-RNA, and related metabolite-linked RNAs place a metabolic cofactor or nucleotide-sugar-derived group at the 5′ end. These structures are not rare spelling variants of m7G caps. They have different chemistry, different likely installation routes, different detection strategies, and different candidate decapping enzymes.

![Figure 26.5. Canonical and Noncanonical RNA 5′-End Topologies](../assets/figures/chapter1025_figure5.png)

**Figure 26.5. Canonical and Noncanonical RNA 5′-End Topologies.** Three comparison bands distinguish a shared canonical m7G-cap core with cap0, cap1, and cap2 ribose-methylation states; specialized trimethylguanosine and metabolite-linked NAD-RNA ends; and uncapped triphosphate, monophosphate, and hydroxyl termini. Each end is paired with the proteins that can access it and the enzyme class that removes, converts, repairs, or degrades it. The figure explicitly limits IFIT and RIG-I recognition to compatible molecular contexts, treats metabolite-cap readers and trimethylguanosine turnover as context-dependent, and shows that XRN-family 5′-to-3′ exonucleases require a monophosphate rather than simply any uncapped end.

NAD-RNA is the best-developed noncanonical cap example. In several systems, RNA polymerases can initiate transcription with NAD or NADH in place of ATP when the first transcribed position is compatible, creating a transcript whose 5′ end carries NAD. Bacterial, eukaryotic, mitochondrial, and archaeal contexts have all contributed to the field. However, transcriptional initiation cannot explain every proposed noncanonical cap on every processed RNA, so post-transcriptional routes remain plausible or likely in selected settings.

The function of NAD-RNA and related caps is not a single rule. In some contexts NAD-related caps can protect RNA from particular 5′-end-dependent decay routes; in other contexts they can recruit deNADding enzymes or mark RNA for turnover. Archaeal studies have identified NAD-RNA and ADPR-RNA species and have connected these states to specific detection and turnover models. The appropriate claim is context-dependent. A bacterial sRNA, an archaeal RNA, a mitochondrial RNA, and a synthetic NAD-capped model RNA may not have the same fate.

UDP-GlcNAc-RNA illustrates a field still moving from chemical discovery to biological interpretation. Recent work identified UDP-GlcNAc-capped RNA species and characterized enzymes that can modify or remove such caps in vitro, including hNudt5 activity on UDP-GlcNAc-RNA. That establishes chemical and enzymological plausibility, but it is not the same as a complete model for transcript-specific regulation in living cells. The chapter therefore treats UDP-sugar caps as real cap-like 5′ structures with developing functional evidence.

The main caution for this section is that cap detection is easier than cap function. A noncanonical cap can be chemically present, enriched by a method, or hydrolyzed by an enzyme in vitro without necessarily acting as a regulatory signal in vivo. Strong biological claims require information about abundance, transcript identity, enzyme specificity, perturbation, and phenotype.

## 26.3. Cap-binding complexes in processing and export

A cap becomes biologically meaningful when proteins read it. The cap-binding proteins discussed in this chapter are not generic RNA-binding proteins; they recognize the special chemistry and geometry of capped 5′ ends. By binding the cap, they can protect the end, recruit processing factors, promote export, license translation, or compete with decay enzymes.

The nuclear cap-binding complex is the main cap reader for many newly made Pol II transcripts. In the classical metazoan description, CBC contains a small cap-binding subunit, CBP20 or NCBP2, and a larger subunit, CBP80 or NCBP1. The complex binds the m7G cap and helps couple the 5′ end to early pre-mRNA processing, export, surveillance, and mRNP remodeling. Structural work explains why CBC is not simply a generic RNA clamp: cap recognition involves a defined cap-binding pocket and conformational accommodation, while newer effector-complex structures show how productive and degradative co-transcriptional pathways can compete for a CBC-bound transcript.

CBC should not be imagined as only a cap cover. It is a platform that helps specify a nuclear mRNP state. A capped pre-mRNA can be spliced more efficiently in some contexts, can interact with 3′-end formation and export factors, and can be routed away from immediate decay. The cap also helps distinguish properly initiated Pol II products from cryptic, uncapped, or incompletely matured RNAs. This does not mean CBC alone determines RNA fate; it means the cap contributes to a network of processing and surveillance decisions.

Export illustrates the difference between an input and a complete pathway. Mature mRNA export requires adaptors, TREX-related factors, nuclear pore interactions, remodeling enzymes, and quality-control checkpoints. The cap and CBC are important inputs into export competence, but they are not the whole export machinery. NCBP3-containing complexes, CBC-ARS2 pathways, and CBC-ALYREF/TREX contacts show that cap-linked export coupling is modular rather than a single fixed route. Mutations or disruptions in export factors can produce strong developmental or disease phenotypes, yet those phenotypes should not be attributed to cap binding unless the experiment specifically tests the cap-CBC-export connection.

Small nuclear RNA export and maturation provide a useful boundary case. Some snRNAs are transcribed by Pol II, capped, bound by cap-associated export factors such as PHAX, exported to the cytoplasm, assembled with Sm proteins, hypermethylated at the cap, and then reimported into the nucleus as spliceosomal small nuclear RNPs. The detailed spliceosomal pathway belongs in [Chapter 27](chapter1026.md), but the example teaches a general principle: cap binding can start a compartmental itinerary rather than simply stabilize an RNA.

Cap-binding handoff is a recurring theme. A cap may first be bound by nuclear CBC, later by factors involved in export or surveillance, and eventually by cytoplasmic translation factors or decapping machinery. These transitions are regulated and incomplete. Some RNAs remain nuclear, some are degraded before export, some use specialized cap-binding proteins, and some viral RNAs manipulate or bypass the normal sequence of handoffs. A cap is therefore a binding site whose meaning changes over the RNA life cycle.

**Table 26.2. Cap-Binding and Decapping Factors.** Key proteins and complexes that recognize, bind, or remove RNA 5′ caps, with their main substrates, cellular compartments, functions, and organism-specific caveats.

| Factor or complex | Main substrate | Compartment | Main function | Organism caveats |
| --- | --- | --- | --- | --- |
| **CBC (NCBP2/NCBP1; CBP20/CBP80)** | m7G cap (cap0 or cap1) | Nucleus | Couples cap to processing, export, surveillance, and mRNP assembly | NCBP3 and ARS2/SRRT extend CBC function in some contexts; metazoan and fungal compositions differ |
| **eIF4E within eIF4F (eIF4E, eIF4G, eIF4A)** | m7G cap | Cytoplasm | Recruits 40S ribosome; links cap to scanning and start-codon selection | eIF4E-binding proteins (4E-BPs) compete with eIF4G; regulated by mTOR and stress signaling |
| **DCP2/DCP1 complex** | m7G-capped mRNA | Cytoplasm; P-bodies | Catalytic decapping committing mRNA to 5′-to-3′ decay | Yeast and metazoan cofactor sets differ; metazoan assemblies include EDC proteins, Pat1, LSM complexes |
| **DcpS (DCPS)** | m7GpppN cap remnants from 3′-to-5′ decay | Cytoplasm; nucleus | Scavenger decapping; prevents accumulation of cap fragments | Substrate length specificity differs between yeast and mammals |
| **DXO/Rai1-family** | Incompletely capped RNA; some noncanonical caps | Nucleus (primarily) | Cap quality control; removes aberrant caps; has exonuclease activity | Metazoan DXO and yeast Rai1 differ in substrate range and exonuclease coupling |
| **NudC-like proteins** | dpCoA-RNA; possibly NAD-RNA | Primarily bacterial; context-dependent | Removes CoA-linked caps (deCoAping) | Primarily bacteria; eukaryotic homologs under active investigation |
| **NUDT-family proteins (e.g., hNudt5, NUDT16)** | m7G caps; UDP-GlcNAc-RNA; other metabolite caps | Cytoplasm; nucleus | Metabolite-cap removal; selective m7G decapping activities | Family membership alone is insufficient evidence; direct substrate proof required per member |
| **Xrn1/Rat1 (downstream of decapping)** | RNA with exposed 5′-monophosphate after cap removal | Cytoplasm (Xrn1); nucleus (Rat1) | 5′-to-3′ exonuclease; acts after cap or metabolite-cap removal | Xrn1/Rat1 are not decapping enzymes; they act on RNA ends exposed by upstream enzymes |
| **Viral D9/D10-like Nudix decappers (poxvirus)** | m7G-capped host mRNA | Cytoplasmic viral replication site | Host mRNA decapping as part of viral host-shutoff | Poxvirus family; not universal among large DNA viruses |
| **OV71 (orf virus Nudix decapper)** | Capped RNA substrates | Cytoplasm | Viral decapping; supports efficient viral replication | Orf poxvirus; recently characterized; replication role demonstrated biochemically and virologically |
| **Cap-snatching endonuclease (e.g., influenza PA)** | Capped host mRNA 5′ fragments | Nucleus (influenza and related viruses) | Generates capped primers for viral transcription by cleaving host mRNA | Influenza and selected segmented RNA viruses; mechanistically distinct from viral de novo capping |

## 26.4. Cap-dependent translation and surveillance

Translation initiation is the process that positions a ribosome at an initiation codon and prepares the first peptide bond. In many eukaryotic mRNAs, the 5′ cap promotes initiation by recruiting eIF4E, a cap-binding protein that functions within the eIF4F initiation assembly. eIF4F is typically described as containing eIF4E, the scaffold protein eIF4G, and the RNA helicase eIF4A. Together with other initiation factors, the 40S ribosomal subunit, initiator tRNA, the poly(A)-binding protein, and the 5′ untranslated region, the eIF4F-centered apparatus helps ribosomes find a start codon.

Cap-dependent translation should not be reduced to "cap equals protein synthesis." The cap is a privileged recruitment handle, but the transcript's 5′ UTR structure, upstream open reading frames, RNA-binding proteins, poly(A) tail, coding-region features, cellular stress state, nutrient signaling, and competition among mRNAs all influence translation output. A capped RNA can be poorly translated, and an uncapped or circular RNA can be translated through cap-independent mechanisms in selected contexts. [Chapter 66](chapter1061.md) and Chapters [67](chapter1062.md)-[72](chapter1067.md) develop the full translation machinery.

The cap-binding translation apparatus is also a regulatory target. eIF4E availability can be limited by 4E-binding proteins that compete with eIF4G, signaling pathways can alter eIF4F assembly, and viral proteins can redirect or inhibit cap-dependent initiation. Structured 5′ UTRs can increase dependence on helicase activity, while short and unstructured leaders can behave differently. The cap therefore sets up an initiation opportunity that must be interpreted together with the rest of the mRNA.

Surveillance pathways read cap state indirectly and directly. One widely taught example is nonsense-mediated decay, which connects translation termination behavior to mRNP history and exon junction information. Older "pioneer round" models emphasized a CBC-bound early translation phase, but this should not be treated as a rigid universal chronology for every mRNA. Cap-binding state, translation initiation, exon junction complexes, poly(A)-binding proteins, and decay factors form a network whose exact timing can differ across transcripts and organisms.

Cap structure also contributes to innate immune surveillance. Many antiviral sensors are tuned to features common on viral or aberrant RNAs, including exposed 5′ triphosphates, double-stranded RNA, uncapped or incompletely capped RNAs, and absent ribose methylation. Cap1 formation and related 2′-O methylation states can help cellular RNA avoid some antiviral recognition pathways, while viral RNAs often need capping or cap mimicry to translate efficiently and evade immunity. The receptor-level mechanisms, including RIG-I-like receptors, IFIT-family proteins, TLR7/8, PKR, OAS/RNase L, and TRIM25-linked signaling, are treated in [Chapter 108](chapter1103.md).

Circular RNA translation is a boundary case that prevents overgeneralization. Some circular RNAs can be translated through internal ribosome entry, RNA modifications, or engineered translation elements, but they do not have a conventional 5′ cap because they lack a free 5′ end. Circular-RNA translation papers therefore help define what cap-dependent translation is not. They should not be used as evidence for eIF4E-mediated cap-dependent initiation.

> **Box 26.1. Common Model Confusions**
>
> - Not all RNA caps are m7G caps. Many RNAs have triphosphate, monophosphate, protein-linked, metabolite-linked, or lineage-specific 5′-end structures; the m7G cap is the canonical eukaryotic Pol II case, not a universal RNA feature.
> - NAD metabolism is not evidence for NAD-RNA cap biology. Claims about NAD-capped RNA require direct RNA-end evidence; measurements of cellular NAD levels or general NAD-pathway activity do not establish that specific RNAs carry NAD at their 5′ ends.
> - A capped RNA is not necessarily translated. Cap-dependent translation also requires compatible UTR structure, active initiation factors, successful ribosome recruitment, and the absence of dominant translational repression; cap presence is necessary but not sufficient.
> - Circular RNA translation is not cap-dependent translation. Circular RNAs lack a free 5′ end and rely on internal ribosome entry elements, RNA modifications, or engineered sequences; they do not use eIF4E-mediated m7G cap recognition.
> - P-body localization is not proof of decay mechanism. Observing a transcript or enzyme in a P-body does not establish where decapping occurs, whether P-body assembly is causally required for decay, or whether an mRNA is being degraded rather than stored or transiently repressed.
> - Viral capping, cap-snatching, and viral decapping are distinct mechanisms. Viral capping synthesizes a new cap on viral RNA; cap-snatching acquires a capped fragment from a host mRNA for use as a transcription primer; viral decapping removes caps from host or viral RNAs as part of host-shutoff or replication strategy.

## 26.5. Decapping enzymes and RNA decay entry points

Decapping removes the protective and regulatory structure from an RNA 5′ end. For many eukaryotic mRNAs, decapping is a decisive step because it exposes the RNA to 5′-to-3′ exonucleases such as Xrn1 in the cytoplasm or related nuclear enzymes in other contexts. The cap and poly(A) tail can be viewed as protective end features; deadenylation weakens the 3′ end, while decapping opens the 5′ end. Many decay pathways begin with deadenylation and then proceed through decapping, but the order and coupling vary.

DCP2 is the major catalytic subunit for canonical cytoplasmic mRNA decapping in eukaryotes. DCP2 belongs to the Nudix hydrolase family and hydrolyzes the cap linkage to generate products that can be further processed. DCP1 is a regulatory factor rather than the main catalytic subunit. Additional cofactors, including EDC proteins, Pat1, LSM complexes, DDX6-like helicases, and metazoan scaffold proteins, influence activation, substrate access, localization, and coupling to other decay or repression events.

DCP2 regulation is not identical in every eukaryote. Yeast studies have provided much of the mechanistic and structural logic for decapping activation, including autoinhibited states and cofactor-stabilized active conformations. Metazoan cells contain additional paralogs, scaffolds, and regulatory connections, including links to miRNA-mediated repression and developmentally controlled mRNA decay. A diagram of "the decapping complex" should therefore mark conserved catalytic logic while avoiding a false one-size-fits-all holoenzyme.

Decapping is coupled to translation and repression. A translating mRNA is often protected from decapping because translation factors and the closed-loop mRNP state can oppose decay-factor access. As translation decreases, deadenylation progresses, or regulatory proteins bind, decapping can become more likely. Nonsense-mediated decay, AU-rich element-mediated decay, miRNA-associated repression, maternal mRNA clearance, stress responses, and developmental transitions can all converge on decapping in particular contexts.

P-bodies are important but often overinterpreted. Processing bodies, or P-bodies, are cytoplasmic RNP granules enriched for translationally repressed mRNAs and decay factors. They can influence mRNA storage, repression, and decay, and recent work supports direct roles for P-body organization in selected decay pathways. However, observing a transcript or enzyme in a P-body does not prove that decapping occurs there, that P-body assembly is required for decay, or that every P-body-associated mRNA is being degraded. Localization, biochemical activity, and causal requirement are separate claims.

The DCP2 pathway is only one branch of cap removal. Scavenger decapping enzymes act on cap remnants after 3′-to-5′ decay has produced small capped fragments. DXO/Rai1-family enzymes can participate in cap quality control by acting on incompletely capped RNAs or selected noncanonical caps. Nudix-family proteins other than DCP2 can remove m7G caps or metabolite-linked caps depending on substrate and organism. Family membership is only a clue; direct biochemical and cellular evidence are needed to assign a specific decapping role.

Noncanonical caps have their own decapping vocabulary. Removal of an NAD cap is often called deNADding; analogous terms such as deFADding or deCoAping are used for other metabolite caps. Enzymes such as NudC-like proteins, selected NUDT proteins, DXO/Rai1-family enzymes, and Xrn/Rat1-linked activities have been implicated in removal or turnover of different cap-like ends. The cleavage position matters. Removing a cap-like group intact, cleaving within a pyrophosphate linkage, or degrading from a processed end can produce different chemical products and different biological interpretations.

Viral decapping enzymes show that cap removal can be a host-shutoff strategy. Some large DNA viruses encode Nudix-family decapping enzymes that alter host and viral RNA pools. Orf virus OV71 is a recent example: biochemical and functional evidence indicates that OV71 can bind RNA, decap capped RNA substrates, and support efficient viral replication. Viral decapping must be distinguished from viral capping and from cap-snatching. Viral capping builds a cap on viral RNA; cap-snatching steals a host capped fragment for use as a primer; viral decapping removes caps from RNAs.

The main teaching point is that decapping is a regulated commitment point, not a passive reversal of capping. A decapped RNA has a new biochemical identity. It may be rapidly degraded, processed by surveillance, or used as evidence of an upstream regulatory decision. To interpret decapping, one must ask which cap was removed, which enzyme acted, what product was generated, what exonuclease or processing step followed, and whether the event was causally required for the observed RNA fate.

## 26.6. Cap analysis methods and therapeutic cap engineering

Cap analysis asks at least three different questions. The first is chemical identity: what 5′ structure is present? The second is transcript assignment: which RNA molecules carry that structure? The third is stoichiometry: what fraction of a given RNA class or product molecule carries each cap state? No single common method answers all three perfectly.

Cap-sensitive transcription start-site methods, including CAGE-like and RAMPAGE-like approaches, are powerful for mapping capped 5′ ends or promoter usage, but they often infer cap presence from capture behavior rather than directly identifying cap chemistry. They can distinguish transcript starts and promoter architecture better than they distinguish cap0 from cap1, m6Am, NAD-RNA, or other cap-like structures. Cap-trapping and newer full-length sequencing approaches improve recovery of complete 5′ ends, but they still need orthogonal chemistry when the biological claim concerns exact cap identity rather than transcript boundaries.

Mass spectrometry answers a different question. Liquid chromatography-tandem mass spectrometry can identify cap structures after nuclease digestion and can quantify cap species in a purified RNA sample. Its strength is chemical specificity. Its weakness is that digestion destroys most sequence context unless the experiment uses careful enrichment, fractionation, or targeted oligonucleotide analysis. A mass spectrum can establish that a cap exists in a sample without proving which transcript carried it.

**Table 26.3. Cap-Analysis Methods.** Methods for measuring RNA 5′ cap identity, transcript assignment, and stoichiometry, with their outputs, strengths, main limitations, and recommended orthogonal validation.

| Method | Cap / 5′ end detected | Output | Strength | Main limitation | Orthogonal validation |
| --- | --- | --- | --- | --- | --- |
| **CAGE (Cap Analysis of Gene Expression)** | Capped 5′ ends; primarily m7G and related | Transcription start-site map; promoter usage | Genome-scale TSS mapping; high resolution | Infers cap from capture behavior; cannot distinguish cap0/cap1 or detect metabolite caps | LC-MS/MS for cap chemistry; RAMPAGE |
| **Cap-trapping** | m7G-capped mRNA (chemical enrichment) | Enriched capped RNA library | Broad enrichment; identifies capped transcripts genome-wide | Enrichment efficiency varies; does not identify cap methylation state | Mass spectrometry; enzyme-sensitivity assay |
| **RAMPAGE** | Capped 5′ ends (nascent or mature) | TSS map with strand specificity | Lower background than bulk CAGE; active promoter identification | Infers cap from enrichment; not designed for metabolite-cap chemistry | CAGE; LC-MS/MS; direct cap enzyme assays |
| **LC-MS/MS** | Any cap after nuclease digestion (cap0, cap1, cap2, NAD, FAD, dpCoA, UDP-GlcNAc) | Cap species identity and quantification | Direct chemical identification; distinguishes closely related structures | Nuclease digestion destroys transcript context; enrichment needed for low-abundance caps | Transcript sequencing; APB gel; enzymatic decapping assay |
| **NAD captureSeq** | NAD-capped RNA | Transcript-level NAD-RNA inventory | Links NAD cap to specific transcripts at genome scale | Incomplete capture efficiency; possible side reactions with NAD analog | LC-MS/MS; SPAAC-NAD-Seq; NudC sensitivity assay |
| **NAD tagSeq** | NAD-capped RNA | Transcript map with single-nucleotide 5′ resolution | Direct sequence readout of NAD-RNA 5′ end position | Enzyme-dependent; possible off-target ligation events | LC-MS/MS; NAD captureSeq; APB gel |
| **SPAAC-NAD-Seq** | NAD-capped RNA via click-chemistry label | Transcript-level NAD-RNA map | Chemical selectivity without enzyme dependence; strain-promoted azide-alkyne cycloaddition | Labeling efficiency and background; requires intact NAD group on RNA | LC-MS/MS; NAD captureSeq |
| **NAD-capQ** | NAD-RNA (bulk quantification) | Stoichiometric estimate of NAD-RNA fraction in a sample | Quantitative; assesses bulk NAD-RNA levels without sequencing | No transcript identity provided; enzyme-dependent quantification | LC-MS/MS; captureSeq for transcript-level resolution |
| **APB gel (acryloylaminophenylboronic acid)** | NAD-RNA detected by boronate affinity mobility shift | Gel mobility shift indicating NAD-capped RNA fraction | Simple; independent of sequencing; orthogonal to chemical methods | Low throughput; no transcript identity; requires abundant input RNA | Mass spectrometry; NAD captureSeq |
| **CapZyme-Seq** | Metabolite caps recognized by specific decapping enzyme | Transcript map after enzyme-specific cap removal and ligation | Enzyme selectivity provides cap-type specificity | Enzyme promiscuity; ligation bias; requires well-characterized decapping enzyme | LC-MS/MS; APB gel; chemical enrichment methods |
| **Manufacturing QC assays (HPLC, CE, RP-HPLC)** | Cap0, cap1, uncapped RNA, dsRNA impurities in therapeutic product | Capping efficiency; cap species ratios; impurity profile | Direct product characterization; quantitative; designed for release testing | Batch-specific; method development required; may not resolve all cap isomers | Mass spectrometry for cap identity; functional translation assay |

Enzymatic and chemical enrichment methods connect cap chemistry to sequence, but they introduce their own assumptions. NAD captureSeq, NAD tagSeq, SPAAC-NAD-Seq, NAD-capQ-like assays, APB electrophoresis, and CapZyme-style workflows have made noncanonical cap biology experimentally tractable. Their limitations include incomplete reaction efficiency, enzyme promiscuity, side reactions, transcript abundance bias, RNA fragmentation, ligation bias, and difficulty distinguishing closely related cap-like states. Strong noncanonical cap studies use orthogonal validation rather than relying on a single enrichment signal.

Cap analysis is especially important for synthetic mRNA. In vitro transcribed RNA can be capped after transcription by enzymes, capped during transcription with cap analogs, or produced through chemical or chemoenzymatic routes. Post-transcriptional enzymatic capping can produce high capping efficiency but adds process steps and requires control of enzyme performance. Co-transcriptional capping is operationally attractive but depends on cap analog design, polymerase initiation behavior, cap orientation, and competition with ordinary nucleotides. Older cap analogs could incorporate in reverse orientation; anti-reverse cap analogs, or ARCAs, were designed to reduce that problem. Modern trinucleotide and cap1-oriented systems, including CleanCap-like workflows, are designed to improve orientation and cap1 installation, but they still require product-specific analytics rather than assumptions from reagent identity.

Therapeutic cap engineering is not only about translation. Cap structure can influence translation initiation, RNA half-life, innate immune recognition, product heterogeneity, release testing, and interaction with other design variables. A cap0 product, a cap1 product, an uncapped RNA impurity, a 5′ triphosphate impurity, and double-stranded RNA contaminants can have different biological effects. Those effects also depend on modified nucleosides, UTRs, coding sequence, poly(A) tail design, purification, lipid nanoparticle formulation, delivery route, target cell, and intended immune context.

The cap can also be an antiviral target. Viral methyltransferases, viral polymerase-associated capping enzymes, cap-snatching endonucleases, and host factors required for viral cap maturation can be vulnerable in particular viral systems. However, "targeting viral capping" is not a single therapeutic strategy. A drug that blocks influenza cap acquisition, a drug that inhibits a flaviviral methyltransferase, and a host-directed perturbation of a methyltransferase dependency would have different selectivity, resistance, toxicity, and immune consequences. This chapter marks antiviral cap targeting as a mechanistic area that requires target-by-target references before therapeutic ranking claims.

The practical rule for cap methods is to match the assay to the claim. Use chemical methods for chemical identity, sequence-readable methods for transcript assignment, quantitative product analytics for stoichiometry, and perturbation experiments for biological function. A cap peak, a cap-enriched sequencing read, and a phenotype after enzyme depletion are related pieces of evidence, not interchangeable proof.

## Experimental Foundations and Evidence Standards

Cap biology spans chemistry, enzymology, cell biology, virology, immunology, and manufacturing. Each evidence type has a specific reach.

Biochemical reconstitution can establish mechanism. If a purified capping enzyme converts a defined triphosphate RNA to a capped product, or a purified decapping enzyme hydrolyzes a defined cap substrate, the experiment can identify substrates, products, cofactors, kinetics, and inhibitor sensitivity. The limitation is context. An enzyme that acts on a model RNA in vitro may not encounter the same substrate in a cell, may require cofactors, or may be regulated by localization and RNP state.

Structural biology explains recognition and catalysis. Structures of capping enzymes, cap-binding proteins, eIF4E-like domains, DCP2-family decapping enzymes, or Nudix hydrolases can reveal cap pockets, catalytic residues, conformational switches, and cofactor contacts. Structural data do not by themselves establish transcript targets or cellular flux, but they are essential for understanding specificity and druggability.

Genetics and perturbation connect cap enzymes to cellular function. Depleting or mutating a capping enzyme, methyltransferase, cap-binding protein, or decapping factor can change RNA abundance, processing, export, translation, immune activation, or viability. Interpretation requires care because cap enzymes often affect many RNAs. A phenotype after RNMT, CMTR1, DCP2, or NUDT depletion may reflect direct cap chemistry, indirect stress responses, altered transcription, global translation changes, or decay of specific targets.

Sequencing gives transcript-level resolution but often indirect chemistry. Cap-enriched libraries can identify candidate capped transcripts or transcription start sites, but capture efficiency depends on cap structure, RNA integrity, ligation behavior, and enzymatic pretreatment. Noncanonical cap sequencing is especially sensitive to chemistry-specific bias. Strong studies validate the same cap class with independent methods and include controls for uncapped, triphosphorylated, and chemically related RNAs.

Mass spectrometry gives chemical confidence but may lose transcript context. Nuclease digestion followed by LC-MS/MS can distinguish cap structures that sequencing cannot. However, the measured cap pool can come from multiple RNA classes, and sample preparation can enrich or deplete specific RNAs. The strongest cap-identity claims combine mass spectrometry with transcript mapping and enzymatic sensitivity.

Clinical and manufacturing evidence answers product-specific questions. For therapeutic mRNA, cap analytics must measure capping efficiency, cap structure, impurity profile, batch consistency, and functional output. Translation in cultured cells, innate immune readouts, animal pharmacology, and clinical immune responses are all relevant, but none replaces direct product characterization. Cap engineering should be interpreted as one element in a whole-RNA and delivery system.

## Biological Contexts Across Organisms and RNA Classes

Metazoan Pol II mRNAs are the standard teaching case because they connect early capping, CBC binding, splicing, export, eIF4E-dependent translation, and DCP2-mediated decay. In these transcripts, the m7G cap is part of a broader mRNP identity system. It cooperates with the poly(A) tail, exon junction complexes, UTR-bound proteins, codon-level features, and cytoplasmic signaling to determine RNA fate.

Fungi, plants, and protists share major themes but differ in details. Fungal capping enzyme architecture, plant RNA decay pathways, and protist cap maturation can diverge from mammalian defaults. Trypanosomes provide an important reminder because trans-splicing and cap4 structures create a 5′-end biology that cannot be reduced to mammalian cap0-cap1 terminology. A claim about "eukaryotic capping" should therefore state which organisms support the generalization.

Small nuclear RNAs and related stable RNPs use cap biology for routing as well as protection. Some snRNAs begin as capped Pol II products, undergo export, RNP assembly, cap hypermethylation, and nuclear reimport. snoRNAs and scaRNAs have their own processing and RNP assembly paths. These RNAs show that cap recognition is not only about protein-coding mRNA translation.

Bacteria and archaea broaden the chemical universe of RNA 5′ ends. Many bacterial RNAs are not m7G-capped, but they can carry triphosphate, diphosphate, monophosphate, or metabolite-linked ends. NAD-RNA and related caps connect transcription initiation, nucleotide metabolism, and RNA decay in ways that differ from canonical eukaryotic capping. Archaeal NAD-RNA and ADPR-RNA findings further show that noncanonical cap biology is not restricted to bacteria or eukaryotic organelles.

> **Box 26.2. Open Questions in Noncanonical Capping**
>
> - Which noncanonical cap species are abundant enough in specific cell types to act as regulatory signals rather than trace metabolic byproducts?
> - Which individual transcripts carry each noncanonical cap in particular organisms, and does this transcript selectivity reflect a biological function?
> - Which noncanonical caps are installed by transcription initiation using a metabolite in place of a standard nucleoside triphosphate, and which arise by post-transcriptional chemistry, RNA processing, or sample-handling artifacts?
> - Which decapping enzymes act selectively on noncanonical caps in living cells rather than only on purified RNA substrates in vitro?
> - How often do metabolite-linked caps function as regulatory signals rather than as protection marks, decay marks, or byproducts of metabolic fluctuations?
> - How can cap sequencing approaches account for enrichment bias, enzyme promiscuity, ligation artifacts, and RNA fragmentation when reporting the identity and abundance of noncanonical cap species?

Organelles provide additional boundary cases. Mitochondrial and chloroplast RNA 5′ ends reflect organelle-specific polymerases, processing enzymes, and evolutionary histories. Some mitochondrial systems have evidence for NAD-related 5′ ends, while many organellar RNAs are processed by endonucleases and exonucleases rather than by canonical nuclear capping. Organellar cap claims should therefore be made from direct organelle-specific evidence, not inferred from nuclear mRNA.

Viruses treat the 5′ end as a replication, translation, and immune-evasion problem. Some viruses synthesize caps that mimic host mRNA caps. Some steal capped host fragments through cap-snatching. Some encode decapping enzymes that destabilize host RNAs. Some use protein-linked genome ends or internal initiation strategies. These mechanisms have different enzymes and different implications for immune sensing and therapy.

Synthetic and therapeutic RNAs are engineered biological contexts. Their cap structures are chosen during manufacturing rather than inherited from a cellular polymerase pathway. A synthetic mRNA intended for protein replacement, genome editing, vaccination, or cell therapy must balance translation, stability, innate immune tone, manufacturability, analytical release specifications, and compatibility with delivery. Cap chemistry is central, but it is never the only determinant.

## Technology, Computational, Clinical, and Engineering Links

Cap biology is a design layer for RNA therapeutics. mRNA vaccines and therapeutic mRNAs generally require efficient capping, controlled cap methylation, low uncapped RNA content, low double-stranded RNA impurities, appropriate UTRs, optimized coding sequence, a poly(A) tail, and a delivery formulation. Cap0, cap1, cap analog choice, and capping efficiency can alter translation and immune recognition, but the final biological behavior is a system property of the whole RNA and formulation.

**Table 26.4. Therapeutic mRNA Cap Engineering Choices.** Comparison of capping routes used in therapeutic mRNA production, including their representative cap structures, advantages, risks, immune relevance, and manufacturing implications.

| Capping route | Representative cap | Main advantage | Main risk | Immune relevance | Manufacturing implication |
| --- | --- | --- | --- | --- | --- |
| **Post-transcriptional enzymatic capping** | Cap0; or cap1 if CMTR1-like step is added | High capping efficiency; flexible cap state control | Added enzymatic steps; process complexity; enzyme lot-to-lot consistency | Can yield cap1 with sequential methylation; reduced innate stimulation versus cap0 or triphosphate RNA | Sequential enzymatic steps; additional purification; cap state assay required |
| **Co-transcriptional with m7GpppG cap analog** | m7GpppG (cap0) | Single step combined with transcription; operationally simple | Approximately 50% reverse incorporation; reverse-oriented caps are not translated | Cap0 product; more immunostimulatory than cap1; uncapped impurities possible | Requires analog-to-GTP ratio optimization; product heterogeneity must be assessed |
| **ARCA (anti-reverse cap analog)** | ARCA cap0 (3′-OMe blocking on N7-methylguanosine moiety) | Eliminates reverse incorporation; all caps in correct orientation | Higher cost; analog-to-GTP ratio must be validated; still produces cap0 | Cap0; no innate immune advantage over standard cap0 | Requires validated ARCA supply and GTP ratio; standard analytical release methods |
| **Cap1 enzymatic methylation (CMTR1-like step)** | m7GpppNm (cap1) | 2′-O-methyl at +1 position; reduces innate immune activation relative to cap0 | Incomplete methylation yields mixed cap0/cap1 product; enzyme specificity and consistency | Reduces IFIT-family and interferon-related recognition; self-RNA-like profile | Additional enzymatic step after capping; cap1 content ratio required in release assay |
| **Trinucleotide cap analog co-transcriptional** | Cap1 trinucleotide (m7GpppNmpN) | Single co-transcriptional step; cap1 orientation built in; high efficiency | Analog cost and availability; limited direct manufacturing literature currently available | Cap1 structure from the start; favorable innate immune profile | Current commercial platforms (CleanCap-type); cap1 ratio must be verified by release assay |
| **Modified phosphate cap analogs** | Phosphorothioate or other phosphate-modified cap | Potential nuclease resistance; engineered chemical stability | May alter cap protein binding or recognition; limited clinical precedent | Must be tested empirically; phosphate modification can shift innate sensor recognition | Custom synthesis; specialized analytical characterization required |
| **Chemoenzymatic or chemical synthesis routes** | Custom-defined cap structure | Chemical flexibility; access to defined cap states not achievable enzymatically | Scale-up challenges; purity; uncertain compatibility with polymerases and cap readers | Must be tested empirically for each structure | Early development stage; analytical and purification methods may need independent development |

Cap chemistry is also a quality attribute. Manufacturing assays must determine whether the intended cap is present, whether the cap is in the correct orientation, how much uncapped RNA remains, whether cap0 and cap1 states are controlled, and whether impurities could stimulate innate immunity. For this reason, current therapeutic cap discussions should separate reagent choice, enzymatic or co-transcriptional installation, analytical measurement, and biological output.

Antiviral development can target cap pathways. Viral capping enzymes and cap methyltransferases can be attractive because host-like capping is often required for efficient viral mRNA translation and immune evasion. Host-directed cap dependencies can also matter, as illustrated by influenza work linking viral capping or replication to a cellular RNA methyltransferase dependency. A clinical claim must still specify virus family, target enzyme, selectivity window, resistance risk, and host toxicity.

Computational analysis enters cap biology through transcript annotation, cap-enriched sequencing interpretation, cap-aware RNA-seq workflows, and product analytics. Models that infer promoter usage from cap-associated reads must account for RNA processing, degradation intermediates, template switching, reverse-transcription bias, and library construction artifacts. For therapeutic products, computational pipelines must link raw analytical data to cap identity, impurity estimates, batch comparability, and acceptance criteria. These are measurement problems before they are biological interpretation problems.

> **Box 26.3. Cap Engineering for Therapeutic RNA**
>
> - Cap orientation must be controlled. Older cap analogs can incorporate in the reverse orientation and produce untranslatable mRNA; ARCA and trinucleotide cap analogs were designed to enforce correct 5′-to-5′ orientation.
> - Capping efficiency is a quality attribute. Uncapped RNA and 5′-triphosphate RNA impurities can stimulate innate immune sensors and reduce translational output; their levels must be measured and controlled.
> - Cap methylation state must be specified and verified. cap0 and cap1 products differ in innate immune recognition; cap1 content and cap0/cap1 ratios require quantitative assays as part of release testing.
> - Translation, stability, and immune response are jointly determined. Cap structure interacts with modified nucleosides, UTR sequence, coding-sequence optimization, poly(A) tail length, purification method, lipid nanoparticle formulation, delivery route, and target cell type; cap engineering is one design variable in a whole-RNA system.
> - Analytical characterization is required for each engineering choice. Product-specific cap chemistry, impurity profile, batch consistency, and functional output must be measured directly; claims about cap state cannot substitute for measurement.

## Recent Consensus

Current consensus supports several stable conclusions.

The canonical m7G cap is an early, information-rich RNA identity mark. It protects RNA, recruits readers, and participates in processing, export, translation, immune discrimination, and decay decisions. It should not be described only as an exonuclease shield.

Cap0, cap1, cap2, and m6Am are chemically distinct. They occupy different positions in the cap-proximal region and can have different protein, enzyme, immune, and analytical consequences.

Noncanonical metabolite caps are real RNA-end states, but their functions are context-dependent. NAD-RNA has the strongest functional and methodological base, while several other metabolite caps are still closer to chemical discovery and enzymology than to complete cellular regulatory models.

Cap binding is a life-cycle handoff. Nuclear CBC-linked events and cytoplasmic eIF4E/eIF4F-linked events read the same general cap chemistry in different RNP contexts.

Decapping is regulated and mechanistically diverse. DCP2-centered mRNA decapping is central for cytoplasmic mRNA decay, but DXO/Rai1-family enzymes, scavenger decapping enzymes, Nudix hydrolases, and viral decappers expand the cap-removal landscape.

Cap analysis requires orthogonal validation. Sequencing, chemistry, enzymology, and mass spectrometry each answer different questions and each can introduce artifacts.

## Open Questions, Controversies, Deprecated Models, and Common Misconceptions

Open questions:

- Which noncanonical caps are abundant enough to matter biologically in specific cells?
- Which RNAs carry noncanonical caps stoichiometrically rather than at trace levels?
- Which noncanonical caps are installed by transcription initiation?
- Which noncanonical caps arise by post-transcriptional chemistry?
- Which cap-like ends arise during decay or sample handling?
- Which enzymes remove noncanonical caps in vivo?
- Which cap-like ends act as regulatory signals rather than protection marks, decay marks, or metabolic byproducts?
- How should cap-binding handoffs among CBC, export factors, surveillance factors, eIF4E-family proteins, and decapping factors be resolved transcript by transcript?
- How should therapeutic cap-engineering platforms be compared when cap state is a measured product attribute rather than a generic capped-or-uncapped label?

Common misconceptions:

- "All RNA caps are m7G caps." Many RNAs have triphosphate, monophosphate, protein-linked, metabolite-linked, or lineage-specific cap-like ends.
- "Cap0, cap1, cap2, and m6Am are synonyms for cap maturity." They name different chemical states at defined positions.
- "A capped RNA is necessarily translated." Cap-dependent translation also requires compatible UTRs, initiation factors, ribosome recruitment, cellular state, and absence of dominant repression.
- "NAD metabolism papers prove NAD-RNA cap biology." NAD-RNA claims require direct RNA-end evidence.
- "P-bodies are simply sites where all decay occurs." P-body localization, translational repression, storage, and decay require separate evidence.
- "Viral capping, cap-snatching, and viral decapping are the same mechanism." Viral capping builds caps, cap-snatching acquires capped primers, and viral decapping removes caps.
- "A cap-enriched sequencing read identifies exact cap chemistry." Enrichment often reports capture behavior, not complete chemical identity.
