RNA biology begins with a small set of chemical facts that recur in nearly every later chapter. An RNA molecule is a directional polymer of ribonucleotide residues. Each residue contains a nucleobase, a ribose sugar, and a phosphate-containing backbone connection. The bases carry the information-bearing and recognition surfaces; the ribose-phosphate backbone gives RNA polarity, charge, geometry, and chemical reactivity; the 2′-hydroxyl group distinguishes RNA from DNA and helps explain both RNA’s structural versatility and its vulnerability to cleavage. This chapter teaches the chemical vocabulary needed to understand RNA folding, RNA catalysis, transcription, splicing, decay, RNA modifications, RNA measurement, and RNA therapeutics. It does not attempt to replace the later chapters on thermodynamics, three-dimensional structure, modifications, or clinical chemistry; instead it supplies the chemical grammar that those chapters assume.
RNA is built from ribonucleotides, but the word nucleotide is often used imprecisely. A nucleoside is a base attached to a sugar. A nucleotide is a nucleoside with one or more phosphate groups. A nucleotide residue is the part of a nucleotide that remains in a polymer after incorporation. This vocabulary matters because a free adenosine triphosphate, an adenosine residue in an mRNA, a pseudouridine residue in a tRNA, and an antiviral nucleoside analog all have related chemistry but different biological meanings.
RNA polymers have chemical directionality. The 5′ and 3′ labels name positions on the ribose sugar, not arbitrary ends of a line. Polymerases extend RNA by adding nucleotides to a 3′-hydroxyl, so RNA sequences are conventionally written from 5′ to 3′. A mechanistic claim about complementarity, guide targeting, ligation, exonuclease digestion, capping, or sequencing is incomplete if the strand and polarity are unclear.
Three chemical features dominate RNA behavior. First, phosphate groups make the backbone negatively charged under most biological conditions. This charge makes RNA soluble and helps separate nucleic acids from hydrophobic biomolecules, but it also creates electrostatic repulsion during folding and complicates delivery of synthetic RNA into cells. Second, nucleobases can hydrogen-bond, stack, tautomerize, become protonated, and undergo chemical modification. These properties allow RNA to encode sequence, pair with complementary strands, form structured pockets, and respond to chemical probes. Third, the ribose 2′-hydroxyl group influences RNA shape, recognition, and phosphodiester cleavage.
Base pairing is not limited to the familiar A:U and G:C Watson-Crick pairs. RNA also uses G:U wobble pairs, Hoogsteen-edge contacts, sugar-edge contacts, base triples, protonated pairs, modified-base pairs, and larger noncanonical interaction networks. The distinction matters because a secondary-structure diagram shows only part of an RNA’s molecular logic. A base may be unpaired in a helix diagram but stacked, ligand-contacting, metal-coordinating, protein-protected, or part of a tertiary contact.
RNA phosphodiester chemistry links stability, processing, catalysis, and decay. Hydrolysis cleaves a bond using water as the nucleophile. Transesterification exchanges ester linkages, often through attack by an alcohol group. In RNA, the 2′-hydroxyl can be positioned for intramolecular attack on an adjacent phosphate, producing cleavage products such as a 2′,3′-cyclic phosphate and a 5′-hydroxyl. Ribozymes and protein enzymes accelerate phosphodiester reactions by controlling geometry, charge, nucleophile activation, leaving-group stabilization, metal-ion binding, and acid-base chemistry.
Modified nucleotides and synthetic analogs are chemically specific features, not a generic category of improved RNA. A modified base, modified sugar, altered backbone, terminal cap, fluorescent label, clickable handle, photo-crosslinker, or therapeutic analog can change hydrogen bonding, stacking, charge distribution, nuclease sensitivity, immune recognition, protein binding, polymerase compatibility, and manufacturing behavior. The same chemical change can be beneficial in one context and disruptive in another.
This chapter assumes that DNA and RNA are nucleic acids and that molecular structures are built from atoms connected by covalent bonds and stabilized by noncovalent interactions. It does not assume prior expertise in organic chemistry. The most important habit is to read chemical notation as biological information. The prime marks in 2′-hydroxyl and 5′ end refer to ribose atom numbers. The colon in A:U or G:C indicates a base-pairing relationship, not a covalent bond. The phrase “phosphodiester linkage” specifies the covalent backbone connection. The phrase “chemical reactivity” can refer to spontaneous reactions, enzyme-catalyzed reactions, experimental probes, chemical damage, or intentionally installed synthetic handles; the context must be stated.
The running examples are a short mRNA segment, a tRNA, a riboswitch aptamer, a self-cleaving ribozyme, and a chemically modified siRNA. The mRNA segment illustrates sequence polarity, capping, codons, and modification effects. The tRNA illustrates dense modification, wobble pairing, and structured folding. The riboswitch illustrates how bases can create a ligand-binding pocket. The self-cleaving ribozyme illustrates intramolecular phosphodiester chemistry. The siRNA illustrates how synthetic chemistry tunes stability, specificity, and pharmacology.
Several distinctions should be kept separate from the start. A base is not a nucleoside; a nucleoside is not a nucleotide; a nucleotide substrate is not the same object as a residue in a chain. A base pair is not the same as base stacking. A natural modification is not the same as sample damage or a synthetic label. A chemically stable RNA is not automatically functional, and a chemically reactive position is not automatically regulatory. These distinctions prevent many later errors in interpreting RNA experiments.
An RNA polymer is easiest to understand by building it from the inside out. The sugar in RNA is ribose, a five-carbon sugar. Ribose atoms are numbered 1′, 2′, 3′, 4′, and 5′. The prime marks distinguish sugar atoms from atoms in the base. The base attaches to the 1′ carbon through a glycosidic bond. The 2′ carbon carries the 2′-hydroxyl group. The 3′ carbon carries a hydroxyl group that can be used to extend the chain. The 5′ carbon connects to phosphate.

Figure 2.1. RNA Nucleotide Anatomy and Polymer Polarity. RNA is built from ribonucleotide residues, each composed of a nucleobase attached to ribose through a glycosidic bond, with the backbone connected by a phosphodiester linkage between the 3′ oxygen of one residue and the 5′ phosphate of the next. This figure illustrates ribose atom numbering (1′ through 5′), the position of the 2′-hydroxyl group that distinguishes RNA from DNA, and the chemical distinction among a free nucleoside, a nucleoside triphosphate, and a nucleotide residue in a polymer. A notation panel shows how the 5′-to-3′ directionality of RNA chains determines the correct orientation for reading sequences, describing complementarity, and interpreting enzymatic activities.
A free ribonucleoside triphosphate has three phosphates. During RNA synthesis, a polymerase positions the incoming nucleoside triphosphate opposite a template or in another substrate-binding context. The 3′-hydroxyl at the end of the growing RNA attacks the alpha phosphate of the incoming triphosphate. The new phosphodiester linkage joins the old 3′ end to the 5′ phosphate of the added residue, and pyrophosphate is released. The result is extension of the RNA chain in the 5′ to 3′ direction. This chemistry is covered mechanistically for polymerases in Chapter 20 and Chapter 23, but the directional consequence is already essential here: RNA sequences are written from 5′ to 3′ because RNA chains have chemically distinct ends.
The phosphodiester linkage is the standard covalent connection between RNA residues. “Diester” means that phosphate is esterified to two sugar oxygens. Because phosphates remain negatively charged under most biological conditions, an RNA chain is a polyanion. The negative charge is not a minor detail. It makes RNA strongly hydrated and generally water-soluble. It causes RNA strands to repel each other and creates an energetic cost when a long RNA folds into a compact structure. It attracts counterions such as magnesium, potassium, sodium, and polyamines, and it gives RNA-binding proteins a recurring need to recognize or neutralize anionic surfaces.
Backbone charge explains why ions and proteins are inseparable from RNA structure in cells. A naked RNA helix in dilute solution is not the same object as the same RNA in a ribonucleoprotein complex, in a granule, in a viral particle, or in a lipid nanoparticle. Cations can shield repulsion without forming specific inner-sphere contacts. Some metal ions bind more specifically and shape active sites or tertiary folds. Proteins can bind the backbone, bases, or both. Later chapters separate electrostatic screening, metal-ion coordination, and RNA-protein recognition; this chapter emphasizes that all of them respond to the same anionic backbone.
Notation errors often become biological errors. A 5′-triphosphate RNA can activate innate immune sensors in contexts where a capped or monophosphorylated RNA may not. A 5′ cap is not simply “the first nucleotide”; it is a terminal chemical structure joined through unusual linkage chemistry and often accompanied by methylations. A 3′ poly(A) tail is not merely “many A bases”; it is a polymeric end structure with protein-binding, decay, and translation consequences. An exonuclease that acts 5′ to 3′ has different substrate requirements from one that acts 3′ to 5′. A guide RNA written in the wrong orientation can appear complementary when it is not the molecule that would pair in the cell.
Box 2.1. Notation Traps in RNA Chemistry
- Nucleoside vs. nucleotide. A nucleoside is base plus sugar, with no phosphate. A nucleotide is base plus sugar plus one or more phosphates. These terms are not synonyms; confusing them can misrepresent substrate chemistry, drug mechanism, or assay design.
- Base vs. residue. A residue is the unit remaining in a polymer after incorporation; it includes part of the ribose-phosphate framework. A base alone refers only to the heterocyclic ring system.
- Sequence polarity. RNA is written 5′ to 3′ because polymerases add nucleotides to the 3′-hydroxyl. Complementary strands run antiparallel. A guide RNA, antisense strand, or primer written in the wrong orientation appears complementary on paper but is not the active molecule in the cell.
- Modified base vs. modified sugar. Pseudouridine changes the glycosidic linkage and is a base modification. 2′-O-methylation changes the sugar hydroxyl and is a sugar modification. These are chemically distinct, detected by different methods, and have different effects on pairing, structure, and protein recognition.
- Cap structure vs. first transcribed nucleotide. The 5′ cap is joined through an unusual 5′-5′ pyrophosphate linkage and carries N7-methylguanosine plus often additional methylations. It is not simply the first incorporated nucleotide; its chemistry is distinct from the body of the transcript.
- RNA analog vs. natural RNA residue. A phosphorothioate residue, 2′-fluoro nucleotide, or morpholino unit behaves differently from canonical ribonucleotide residues. Chemical notation that obscures these substitutions can mislead interpretation of affinity, nuclease sensitivity, immune sensing, or enzyme compatibility.
Circular RNAs are an instructive boundary case. A circular RNA has no free 5′ or 3′ end, so end-specific enzymes and library methods behave differently. Nevertheless, a circular RNA derives from a transcript segment with a defined sequence orientation, and base pairing within or against the circle still depends on strand polarity. Another boundary case is a chemically synthesized oligonucleotide with nonstandard linkages. It may be written in 5′ to 3′ notation, but some residues, sugars, or backbones may not be ordinary ribonucleotide residues. The notation should not hide the chemistry.
RNA’s four canonical bases are adenine, cytosine, guanine, and uracil. Adenine and guanine are purines, with two fused rings. Cytosine and uracil are pyrimidines, with one ring. The bases contain nitrogen and oxygen atoms that can donate or accept hydrogen bonds, and the bases have aromatic surfaces that can stack against neighboring bases. The same base can therefore participate in multiple kinds of interactions: pairing through hydrogen bonds, stacking through aromatic surfaces, contacts with proteins, coordination through water or ions, and chemical modification at exposed atoms.

Figure 2.2. Base-Pairing and Stacking Vocabulary. RNA base interactions extend well beyond canonical Watson-Crick A:U and G:C pairs, and this figure provides a visual grammar for the main interaction types. Panels compare Watson-Crick pairing, a G:U wobble pair with shifted hydrogen-bond geometry, a Hoogsteen-edge contact using an alternative base face, and a sugar-edge interaction. A base-triple example shows how a third base contacts an existing pair through an unused edge, and an unpaired-but-stacked base illustrates that structural importance does not require hydrogen bonding. One panel shows how a chemical modification such as methylation can block a hydrogen-bond donor or acceptor, altering pairing or protein-recognition potential.
Watson-Crick pairing is the best-known use of base hydrogen bonding, but the chemical personality of a base is broader than its canonical partner. The edges of a base present different donor and acceptor patterns. Adenine can participate in A:U Watson-Crick pairing, but adenine can also be modified at N6, protonated under some conditions, or involved in noncanonical contacts. Guanine can pair with cytosine in a Watson-Crick geometry, pair with uracil in a wobble geometry, form G-quadruplexes in guanine-rich sequences, or become chemically oxidized. Uracil can pair with adenine, participate in wobble contacts, or be replaced by pseudouridine in many biological and engineered contexts. Cytosine can be methylated, deaminated, or protonated in specific structural settings.
Tautomerism is a reversible shift in proton position and bonding pattern. A base’s dominant tautomer usually determines its ordinary hydrogen-bonding behavior, but rare tautomers can expose different donor and acceptor patterns. Protonation is a related but distinct issue: adding a proton changes charge and hydrogen-bonding capacity. These effects are chemically plausible explanations for unusual pairing, pH-dependent folding, and transient mispairing. They require caution. Because rare tautomers and protonated states are difficult to observe directly in many biological systems, a functional claim about tautomerism should be supported by structural, spectroscopic, kinetic, mutational, or chemical evidence rather than by a drawing of a possible form.
Base stacking is at least as important as hydrogen bonding for many RNA structures. Stacking arises when flat aromatic bases align favorably with one another. In a helix, stacking between neighboring base pairs contributes substantially to stability. In loops and junctions, bases that are not paired may remain stacked, which can constrain local structure. In tRNA, stacked helices and noncanonical interactions help generate an L-shaped architecture. In riboswitches, stacking can position bases for ligand binding. In small interfering RNA duplexes, terminal pairing and stacking asymmetry can influence strand selection by Argonaute proteins, a topic developed in Chapter 87 and therapeutic chapters.
Chemical reactivity is often used experimentally to infer RNA structure or modification state. For example, a reagent may preferentially react with flexible ribose conformations, exposed adenines or cytosines, or particular modified bases. The resulting adducts may be read by reverse transcription, sequencing, mass spectrometry, antibody enrichment, or direct RNA sequencing. These methods are powerful because chemistry can report molecular state, but the signal rarely has a single interpretation. Reactivity can reflect solvent exposure, conformational flexibility, protein protection, reagent access, RNA modification, local pH, ion concentration, temperature, reverse-transcriptase behavior, and library construction bias.
Box 2.2. Chemical-Probing Interpretation Caveats
- Reactivity is not base pairing. A high probing signal indicates a solvent-exposed, conformationally flexible, or modified position, not automatically an unpaired base. A base in a breathing helix or a strained loop may show reactivity while still participating in pairing.
- Protein protection can mimic base pairing. A base shielded by a bound protein will show low reactivity, which can be misread as paired or structured in the absence of protein occupancy data.
- Modifications alter reagent behavior. Some chemical probes react preferentially or poorly with modified versus canonical bases, creating apparent structure signals that reflect modification chemistry rather than backbone conformation.
- Library construction introduces bias. Reverse-transcriptase stops, mutation rates, soft-clipping, and adapter-ligation preferences are sequence- and structure-dependent and can skew apparent reactivity profiles independently of true RNA state.
- High-value claims require orthogonal validation. Structural or modification conclusions from probing data should be supported by compensatory mutations, phylogenetic covariation, high-resolution structure, or probing with an orthogonal reagent chemistry.
One common error is to equate “reactive” with “unpaired.” A base in a loop may be protected by a protein. A base in a helix may become reactive because the helix breathes, the local structure is strained, or a modification changes reagent behavior. A low chemical-probing signal may indicate pairing, but it can also indicate protein protection, poor reagent access, or poor recovery of a damaged fragment. For this reason, high-value structural claims usually need orthogonal evidence: compensatory mutations, covariation across homologs, high-resolution structure, biochemical binding, or multiple probes with different chemistries. Chapter 131 treats high-throughput probing methods in depth, and Chapter 132 treats modification detection and mass spectrometry.
The same caution applies to biological claims about base modifications. A detected modification can be a stable enzymatic mark, a low-stoichiometry event, a damage product, an artifact introduced during handling, or a misassigned signal from a related chemistry. Base-resolution mapping has improved rapidly, but direct chemical identity and stoichiometry often still require carefully chosen standards and orthogonal assays. A reader should ask three questions: what chemical group is being detected, what molecule carries it, and what evidence links the signal to a biological function?
Base pairing allows RNA sequence to be interpreted by another nucleic acid strand or by another region of the same molecule. In a Watson-Crick A:U pair, adenine and uracil align their Watson-Crick edges to form compatible hydrogen bonds. In a Watson-Crick G:C pair, guanine and cytosine form a geometrically compatible pair with three hydrogen bonds. These pairs allow templated RNA synthesis, RNA-DNA hybrids, RNA-RNA duplexes, guide-RNA targeting, primer binding, many stems in secondary structures, and many sequence-alignment models.
The simplicity of Watson-Crick pairing is one reason RNA is programmable. A synthetic antisense oligonucleotide can be designed to bind a target RNA by complementarity. A CRISPR guide RNA can be designed to base-pair with a target nucleic acid in a protein complex. A primer can be designed to anneal to a specific template. A short hairpin can be drawn from a sequence by pairing complementary segments. All of these uses depend on the rule that A pairs with U and G pairs with C in ordinary RNA duplex geometry. The same rule can mislead when treated as complete. Real RNA structures contain mismatches, bulges, internal loops, pseudoknots, modified bases, and tertiary contacts.
G:U wobble pairing is the most familiar non-Watson-Crick interaction in RNA. In a G:U wobble pair, guanine and uracil form a pair with shifted geometry. G:U wobble pairs can be stable, recognized by proteins, conserved through evolution, and functionally important. In decoding, wobble describes flexible recognition between codons and anticodons, often shaped by modifications in the tRNA anticodon loop. In structural RNAs, G:U wobble pairs can create distinctive grooves and local shapes. The word wobble therefore has related but context-specific meanings: it can describe a pair geometry, a decoding principle, or a structural feature.
Hoogsteen pairing uses a different edge of the base than Watson-Crick pairing. Sugar-edge contacts use atoms on the side of the base near the glycosidic bond. RNA base-pair nomenclature commonly classifies interactions by the edge used by each base and by the relative orientation of the glycosidic bonds. That formal vocabulary appears more fully in Chapter 4 and Chapter 53, but the core principle belongs here: a base is not a one-purpose object. It has multiple edges, and each edge can support different contacts.
Noncanonical interactions enable RNA to do things that simple duplexes cannot do. A riboswitch aptamer can surround a metabolite with a pocket assembled from canonical pairs, noncanonical pairs, stacks, junctions, and direct ligand contacts. Ribosomal RNA uses many noncanonical interactions to stabilize the ribosome core and position functional groups. tRNA uses conserved tertiary contacts to fold a compact L-shaped molecule. Viral RNAs can use pseudoknots and unusual pairings to control frameshifting, replication, or packaging. Small regulatory RNAs and miRNAs may tolerate mismatches or noncanonical contacts that alter affinity, specificity, or regulatory outcome.
Base triples and higher-order networks further expand RNA structure. A base triple occurs when a third base contacts an existing pair, often through an unused edge. Such interactions can stabilize long-range contacts and ligand-binding pockets. Networks of noncanonical interactions can also create small-molecule binding preferences. This is why a structured RNA may bind a metabolite, antibiotic, or protein in a way that cannot be predicted from Watson-Crick complementarity alone.
Modified bases can preserve, weaken, or redirect pairing. Pseudouridine changes hydrogen-bonding possibilities and local hydration. Inosine pairs differently from adenosine and is important in decoding and RNA editing contexts. Methylations can block a hydrogen-bond donor or acceptor, alter stacking, or create a protein-recognition surface. Synthetic studies can isolate the pairing effect of a modification, but full biological interpretation requires the surrounding sequence, structure, protein context, modification stoichiometry, and cell type.
The practical lesson is to name the interaction that matters. If a claim depends only on complementarity, Watson-Crick notation may suffice. If a claim depends on structure, catalysis, ligand binding, decoding, or protein recognition, “base-paired” is usually too vague. A useful statement should say whether the interaction is Watson-Crick, G:U wobble, Hoogsteen-edge, sugar-edge, protonated, modified-base-dependent, transient, predicted, or experimentally observed.
The phosphodiester backbone must be stable enough to carry information, but reactive enough to be synthesized, processed, repaired, and degraded. RNA achieves this balance imperfectly and productively. Compared with DNA, RNA has a 2′-hydroxyl group adjacent to the backbone phosphate. Under suitable conditions, the 2′-hydroxyl oxygen can attack the neighboring phosphate. This intramolecular reaction can cleave the backbone and form products such as a 2′,3′-cyclic phosphate and a 5′-hydroxyl. The reaction is often faster under alkaline conditions, where deprotonation increases the nucleophilicity of the 2′-oxygen.

Figure 2.3. Phosphodiester Cleavage and Transesterification. The 2′-hydroxyl of RNA can act as an intramolecular nucleophile, attacking the adjacent phosphate to cleave the backbone and generate a 2′,3′-cyclic phosphate and a 5′-hydroxyl fragment. This figure traces the reaction pathway from substrate geometry through the transition state, highlighting the key requirements: 2′-oxygen positioning and activation, approach geometry at the phosphorus center, charge stabilization in the transition state, and leaving-group stabilization. Separate panels show how metal ions and general acid-base groups can assist catalysis in ribozymes and protein enzymes, and a product panel distinguishes the end chemistries produced by hydrolysis versus transesterification versus intramolecular cleavage.
A simplified intramolecular cleavage reaction proceeds in causal steps. First, the 2′-hydroxyl must be positioned near the adjacent phosphate. Second, the 2′-oxygen must become a better nucleophile, often by deprotonation or by active-site organization. Third, the attacking oxygen, phosphorus center, and leaving-group oxygen must approach a geometry compatible with reaction. Fourth, negative charge that develops in the transition state must be stabilized. Fifth, the leaving group must be protonated or otherwise stabilized. Ribozymes and protein enzymes accelerate reaction by solving some or all of these problems.
Hydrolysis and transesterification are related but not identical. Hydrolysis uses water as the nucleophile and breaks a bond by adding water across it. Transesterification uses an alcohol group as the nucleophile and transfers an ester linkage. The spliceosome illustrates transesterification on a large biological scale: intron removal proceeds through two transesterification reactions, and metal ions and active-site organization help position substrates for reaction. Recent mechanistic work continues to refine how ions promote spliceosomal chemistry. Self-cleaving ribozymes, group I and group II introns, RNase P RNA, and ribosomal RNA provide additional examples of RNA-centered or RNA-assisted catalysis, developed further in Chapter 8, Chapter 27, and Chapter 42.
Metal ions are central but should not be treated as magic catalysts. Magnesium ions and other cations can shield backbone charge, organize folds, bind specific sites, activate nucleophiles, stabilize transition states, or help leaving groups. Which role applies depends on the RNA and reaction. A metal-dependence curve alone does not prove a specific catalytic role. Stronger evidence comes from structural placement of the ion, metal rescue experiments, pH-rate profiles, substrate analogs, kinetic isotope effects, mutational effects, and product analysis.
Spontaneous RNA cleavage is biologically and technically important. RNA can degrade during extraction, storage, heating, alkaline treatment, metal contamination, or repeated freeze-thaw cycles. Environmental and origin-of-life studies also examine RNA hydrolysis at mineral-water interfaces, where surfaces can alter local concentration, orientation, and reactivity. These settings differ from regulated cellular RNA decay. In cells, ribonucleases and RNP assemblies control where, when, and how RNA is cleaved. Chapter 32 treats decay enzymes; Chapter 122 treats sample handling and RNA quality control.
Product identity is a key evidence standard. A shorter band on a gel can indicate cleavage, but it does not identify whether the product has a 5′-phosphate, 5′-hydroxyl, 2′,3′-cyclic phosphate, 3′-phosphate, or 3′-hydroxyl. Those ends determine whether ligases, exonucleases, phosphatases, repair enzymes, or sequencing adapters can act. Misreading end chemistry can mislead both biological mechanism and library interpretation.
RNA catalysis also teaches a broader lesson about function. RNA is not only a passive information carrier. RNA can position atoms, bind metals, create local electrostatic environments, and use functional groups from bases, ribose, and bound molecules to accelerate reactions. However, different ribozymes use different catalytic strategies, and a mechanism inferred for one ribozyme should not be copied to another without evidence. Mechanistic claims should specify substrate, active structure, cofactors, pH, products, and rate effects.
“Modified RNA” is too broad to interpret without classification. A natural RNA modification is a chemical alteration installed by an enzyme in a biological context. Examples include pseudouridine, 2′-O-methylation, inosine, N6-methyladenosine, 5-methylcytosine, N4-acetylcytidine, and complex tRNA modifications such as queuosine. Chemical damage is an unintended alteration produced by oxidation, hydrolysis, deamination, alkylation, or handling. A synthetic label is a chemical handle used for imaging, purification, crosslinking, sequencing, or biophysical measurement. A therapeutic analog is a designed chemical variant used to change nuclease resistance, affinity, distribution, immune sensing, potency, or toxicity.
Natural RNA modifications are chemically specific. Pseudouridine is not merely “modified uridine”; it changes the glycosidic connection and local interaction possibilities. Inosine is read differently from adenosine in many pairing contexts. 2′-O-methylation changes the sugar hydroxyl into a methylated ether and can affect nuclease sensitivity, immune recognition, and local structure. N6-methyladenosine changes a base surface and can create or disrupt protein recognition, alter local structure, or affect transcript fate depending on the transcript and RNP context. The effects of a modification in mRNA cannot be assumed from its effect in tRNA, and the effect in a short oligonucleotide cannot be assumed to match a full RNP.
Synthetic RNA chemistry uses the same atoms for different goals. Antisense oligonucleotides and siRNAs often incorporate sugar and backbone modifications to improve nuclease resistance, binding affinity, pharmacokinetics, and tissue exposure. Common sugar modifications include 2′-O-methyl, 2′-fluoro, and 2′-O-methoxyethyl-like chemistries. Backbone modifications include phosphorothioate linkages, in which a nonbridging oxygen is replaced by sulfur. Other platforms use locked nucleic acid-like sugars, morpholino backbones, peptide nucleic acid backbones, or other artificial nucleic acid architectures. These chemistries are covered in detail in Chapter 149 through Chapter 158; here the essential point is that each analog changes multiple properties at once.
The multiple-property rule is central to therapeutic design. A phosphorothioate backbone can increase nuclease resistance and protein binding, but it can also alter distribution and toxicity. A high-affinity sugar modification can improve target binding, but excessive affinity or poorly placed modifications may reduce specificity, disrupt RNase H recruitment, interfere with Argonaute loading, or create manufacturing challenges. A modification that reduces immune sensing in one platform may not solve immune activation in another because innate sensors respond to duplex character, length, end chemistry, contaminants, delivery vehicle, and cell type as well as nucleotide identity.

Figure 2.4. Chemical Substitution Sites and the RNA Properties They Change. Base, ribose, phosphate-backbone, and terminal substitutions act at different physical sites but often alter overlapping sets of RNA properties. A two-by-two chemical-site map highlights the substituted site on the same simplified RNA scaffold and assigns each site a bounded set of possible consequences: pairing, sugar pucker, nuclease resistance, protein binding, immune sensing, polymerase compatibility, charge and distribution, ligation and decay, enzyme access, translation, and stability. Repeated property badges show the many-to-many mapping without implying that every substitution has every effect. A conclusion band states that position and molecular context determine the outcome.
RNA labels and chemical handles require the same caution. Fluorophores can report localization, folding, or binding, but they add size, hydrophobicity, charge, or steric bulk. Photo-crosslinkers can identify contacts, but crosslink yield depends on distance, orientation, irradiation, and local chemistry. Clickable handles enable enrichment or imaging, but they can affect polymerase incorporation and RNA behavior. A labeled RNA is an experimental model of an RNA, not automatically the same molecule in all relevant properties.
Polymerase incorporation is a practical filter on modified-nucleotide use. A modified nucleoside triphosphate may be chemically attractive but poorly accepted by the polymerase used for in vitro transcription or by an engineered polymerase. Incorporation can be incomplete, biased by sequence context, or associated with pausing, misincorporation, abortive products, or altered RNA ends. Polymerase engineering can expand accepted substrates, but each polymerase-analog combination needs direct validation. Chemical usefulness does not guarantee efficient synthesis.
Post-synthetic chemistry offers another route. Instead of asking a polymerase to incorporate a modified nucleotide, chemists can synthesize an RNA with reactive handles or modify an RNA after synthesis. Post-synthetic approaches can install nucleobase modifications, labels, or crosslinking groups in defined positions. These methods are powerful for mechanistic experiments but can be limited by RNA length, yield, protecting-group compatibility, site selectivity, purification, and the need to verify final structure.
Synthetic mRNA provides a concrete example of modification tradeoffs. Modified nucleotides, cap structure, untranslated regions, coding sequence, poly(A) tail, purification, and formulation all influence translation, stability, innate immune sensing, and manufacturing quality. A change that improves protein output in one cell type or delivery vehicle may not generalize to another. The chemical design of mRNA therapeutics therefore belongs partly to this chapter, partly to innate immunity, and partly to manufacturing and delivery chapters.
RNA chemistry creates tradeoffs rather than one-directional rules. The phosphate backbone makes RNA soluble, separable, and recognizable, but the same negative charge makes compact folding and membrane crossing difficult. Complementary base pairing makes RNA programmable, but it also creates off-target pairing and competing structures. The 2′-hydroxyl supports A-form geometry, tertiary contacts, and catalysis, but it also enables cleavage chemistry. Modifications can stabilize one state while destabilizing another. Synthetic analogs can improve drug-like properties while introducing toxicity, altered specificity, or manufacturing problems.
RNA stability illustrates the danger of oversimplification. RNA is often described as unstable, but stability depends on the question being asked. Chemical backbone stability asks how quickly covalent bonds break under defined conditions. Structural stability asks whether an RNA adopts a particular fold. Cellular stability asks how long a molecule persists before decay. Functional stability asks whether the molecule retains its biological activity. A tRNA can be chemically vulnerable in principle but highly persistent in a cell because it is compact, modified, protein-associated, and protected from nucleases. An mRNA can be stabilized by cap chemistry, poly(A) tail length, untranslated regions, codon composition, RNA-binding proteins, and localization, or destabilized by deadenylation, decapping, endonucleolytic cleavage, and surveillance pathways.
Folding is also chemically constrained. RNA helices usually favor A-form geometry because of ribose and backbone constraints. A-form geometry differs from the familiar B-form DNA helix and influences groove accessibility, protein recognition, and tertiary packing. The 2′-hydroxyl can donate or accept hydrogen bonds and can help organize local hydration networks. Base stacking stabilizes helices and loops. Magnesium and other ions help overcome electrostatic repulsion. These chemical features are the physical basis for the thermodynamic and structural models introduced in Chapter 3 and Chapter 4.
Table 2.1. Chemical Features and Biological Consequences. Each chemical feature of RNA produces multiple consequences across biology and technology; this table maps the most important direct effects and flags common overgeneralizations.
| Feature | Direct chemical effect | Biological consequence | Technology implication | Common overgeneralization | Related chapters |
|---|---|---|---|---|---|
| 2′-hydroxyl | Adds nucleophilic hydroxyl at 2′ carbon; biases ribose toward C3′-endo pucker | Enables intramolecular cleavage; supports A-form geometry; contributes to tertiary contacts | Targeted by 2′-O-methyl and 2′-fluoro modifications to improve nuclease resistance | RNA is always unstable because of the 2′-hydroxyl | Chapter 3, Chapter 4, Chapter 8 |
| Phosphate charge | Backbone is negatively charged at physiological pH; attracts counterions | RNA is water-soluble; electrostatic repulsion costs energy in compact folding; cations screen charge | Phosphorothioate substitution alters charge, protein binding, and tissue distribution | Negative charge makes RNA impossible to deliver into cells | Chapter 3, Chapter 4, Chapter 149 |
| A-form geometry | Ribose pucker constrains helices into A-form with narrow deep major groove and wide shallow minor groove | Groove shape influences protein and small-molecule recognition differently from B-form DNA | Modified sugars can shift helix geometry and affect RNase H or Argonaute engagement | RNA and DNA duplexes have equivalent groove accessibility | Chapter 3, Chapter 4 |
| Base stacking | Aromatic base surfaces interact via pi-pi contacts, stabilizing adjacent residues | Stacking stabilizes helices; organized loops constrain flexible regions; unpaired bases can be structurally rigid | Terminal stacking asymmetry in siRNA duplexes influences strand selection by Argonaute | Unpaired bases are always flexible and solvent-exposed | Chapter 3, Chapter 87 |
| G:U wobble | Guanine and uracil form a shifted pair with two hydrogen bonds and altered groove geometry | Stabilizes helices; creates distinctive groove features for protein recognition; enables flexible codon-anticodon decoding | Wobble pairs in antisense or siRNA duplexes can reduce affinity or alter specificity | Only A:U and G:C pairs are stable in RNA helices | Chapter 4, Chapter 26 |
| Metal-ion coordination | Cations shield backbone charge; Mg²⁺ and others bind specific structural sites | Ion binding promotes compact folding, ribozyme catalysis, and ribosome architecture | Metal contamination can trigger RNA cleavage; Mg²⁺ concentration must be controlled in vitro | Mg²⁺ is a universal catalytic cofactor for all ribozymes | Chapter 4, Chapter 8 |
| Cap chemistry | 5′ end bears m7G joined via 5′-5′ pyrophosphate linkage, often with adjacent methylations | Protects mRNA from 5′-exonucleases; promotes translation initiation; modulates innate immune sensing | Cap analog identity and methylation state affect mRNA vaccine translation and immunogenicity | The 5′ cap is simply the first transcribed nucleotide | Chapter 20, Chapter 32 |
| Terminal end chemistry | 5′ end can carry triphosphate, monophosphate, or hydroxyl; 3′ end can be hydroxyl or cyclic phosphate | End state determines exonuclease sensitivity, ligation, decapping, sensor activation, and enzyme access | Library construction and adapter ligation depend on end chemistry; misreading ends misleads interpretation | All RNA ends are chemically equivalent for ligation and sequencing | Chapter 32, Chapter 122 |
Recognition depends on more than sequence complementarity. A short RNA oligonucleotide may bind a transcript by Watson-Crick pairing, but binding strength and biological outcome depend on length, mismatch tolerance, accessibility of the target site, RNA structure, competing proteins, chemical modifications, and cellular concentration. A folded RNA aptamer may bind a small molecule by shape complementarity, hydrogen bonding, stacking, and metal coordination. A protein may recognize a sequence motif only when presented in a particular structure. An innate immune receptor may respond to double-stranded character, end chemistry, cap methylation, nucleoside identity, or delivery-associated context. Thus the statement “RNA recognizes X” should be expanded into a chemical and biological description: which RNA form, which partner, which assay, which conditions, and which consequence?
Biotechnology deliberately exploits and manages these constraints. RNA nanotechnology uses predictable base pairing and tertiary assembly to build scaffolds and devices. RNA sensors and synthetic circuits use folding and ligand recognition to regulate translation or stability. RNA therapeutics use chemical modification, conjugation, purification, and formulation to make labile anionic molecules into medicines. RNA-directed small-molecule discovery uses structured pockets and noncanonical interactions as targets. Each technology depends on the same chemical rules but optimizes a different balance of stability, specificity, delivery, activity, and safety.
The boundary between chemistry and biology is especially visible in RNA measurement. A sequencing library is not a direct photograph of molecules in a cell. Extraction can damage RNA, enzymes can prefer some ends or modifications, reverse transcriptases can stop or misread at chemical lesions, antibodies can enrich related structures, and nanopore signals can be affected by neighboring sequence. The chemical state of the RNA shapes what the method can observe. Chapter 5 develops evidence standards, and Chapters 122 through 132 develop measurement workflows.
The strongest evidence for basic nucleotide chemistry comes from converging chemical, structural, enzymological, and biophysical observations. The composition and connectivity of nucleotides can be established by chemical analysis, enzymatic digestion, synthesis, spectroscopy, crystallography, cryo-electron microscopy, nuclear magnetic resonance, and mass spectrometry. For a chapter like this, many statements are long-established chemical principles, but mechanistic or context-specific claims still need evidence close to the claim.
For base-pairing claims, evidence can come from high-resolution structures, nuclear magnetic resonance, chemical probing, mutational rescue, comparative covariation, thermodynamic measurements, and biochemical binding experiments. Each evidence type has limits. A crystal structure may capture one conformational state. Cryo-electron microscopy may show well-ordered regions better than flexible regions. Chemical probing reports reactivity, not pairing directly. Covariation supports conserved pairing across evolution but can miss recent or lineage-specific interactions. Mutational rescue is powerful when compensatory changes restore function, but rescue can also occur through indirect structural effects.
For phosphodiester cleavage claims, a complete evidence package should identify substrates, products, reaction conditions, catalytic components, rate effects, pH dependence, ion requirements, and structural context when possible. Gel mobility is useful but insufficient by itself. Product-end mapping, mass spectrometry, nuclease sensitivity, ligation behavior, and chemically defined standards can distinguish hydrolysis from transesterification and distinguish one end chemistry from another.
For modification claims, the evidence must identify both the chemical group and its location. Antibody enrichment can suggest the presence of a modification class but can suffer from cross-reactivity and sequence or structure bias. Reverse-transcription signatures can report some modifications but may confuse modifications with damage, structure, or enzyme behavior. Mass spectrometry can identify modified nucleosides with strong chemical confidence but may lose positional information unless combined with sequence-specific workflows. Direct RNA sequencing can preserve native molecules and sometimes detect modification-associated signal shifts, but interpretation requires models, controls, and standards. Chapter 132 expands these points.
For synthetic analog claims, chemistry and biology must both be tested. A useful analog must be synthesized or incorporated reliably, purified, chemically characterized, and assayed in the relevant biological system. For therapeutic applications, potency is not enough. Specificity, nuclease resistance, protein binding, immune activation, distribution, metabolism, toxicity, and manufacturability all matter. A modification that performs well in a short cell-free assay may fail in serum, in a tissue, or during scale-up.
mRNA uses nucleotide chemistry to carry coding information and regulatory information in the same molecule. The coding sequence is read in codons by the ribosome, but the same mRNA also carries a 5′ end state, untranslated regions, possible internal modifications, structures, protein-binding sites, and a 3′ poly(A) tail. The chemistry of the cap, the accessibility of start codons, local structure near regulatory elements, codon composition, and modification state can all affect translation and decay.
Table 2.2. Modification and Analog Classes. RNA modifications and synthetic analogs span several chemical classes with distinct origins, uses, and caveats; each changes multiple molecular properties simultaneously.
| Class | Example | Chemical change | Common use or biological context | Major caveat |
|---|---|---|---|---|
| Base modifications (natural) | Pseudouridine (Ψ), N6-methyladenosine (m6A) | Changed glycosidic bond orientation (Ψ); N-methyl group on adenine exocyclic nitrogen (m6A) | tRNA and rRNA stability, ribosome biogenesis; mRNA abundance regulation and translation control | Effects differ by transcript type, RNP context, and stoichiometry; cannot be inferred across contexts |
| Sugar modifications | 2′-O-methylation, 2′-fluoro, 2′-O-methoxyethyl | Hydroxyl converted to methyl ether, fluorine, or bulkier methoxyethyl group at 2′ carbon | Nuclease resistance and improved binding affinity in antisense and siRNA therapeutics | Same sugar modification simultaneously affects nuclease resistance, immune sensing, helix geometry, and protein binding |
| Backbone analogs | Phosphorothioate (PS) linkage | One nonbridging oxygen of phosphate replaced by sulfur | Nuclease resistance and enhanced serum protein binding; widely used in antisense oligonucleotides | Alters charge distribution; increases off-target protein interactions and can elevate toxicity |
| Cap analogs | m7GpppG, anti-reverse cap analog (ARCA) | Modified 5′-5′ pyrophosphate linkage bearing N7-methylguanosine; ARCA blocks reverse incorporation | In vitro-transcribed mRNA for vaccines and protein-replacement therapeutics | Cap structure and flanking methylation state affect translation efficiency and innate immune activation |
| Fluorescent labels | Cy3, Cy5, fluorescein-nucleotide conjugates | Bulky aromatic fluorophore conjugated to base or sugar position | RNA localization imaging, FRET-based folding studies, single-molecule experiments | Fluorophore adds steric bulk and hydrophobicity that can perturb folding or protein binding |
| Clickable handles | 4-thiouridine (4SU), azide- or alkyne-modified nucleotides | Reactive thiol, azide, or alkyne installed at base or sugar | Pulse-chase metabolic RNA labeling, enrichment, and click-chemistry conjugation | Reactive group can reduce polymerase incorporation efficiency and alter RNA behavior in cells |
| Photo-crosslinkers | 4-thiouridine (4SU), 6-thioguanosine | Photo-reactive sulfur or halogenated group that covalently captures nearby molecules on UV irradiation | Identifying RNA-protein contacts (CLIP) and RNA-RNA contacts (CLASH) | Crosslink yield depends on irradiation geometry, local distance, and reagent access; rarely quantitative |
| Therapeutic oligonucleotide chemistries | Morpholino, locked nucleic acid (LNA)-like, peptide nucleic acid (PNA) | Morpholine ring replaces ribose (morpholino); bridged bicyclic sugar (LNA); amide backbone replacing phosphodiester (PNA) | Splice-switching, gene silencing, diagnostics; high-affinity hybridization to target RNA | Altered backbone changes polymerase compatibility, cellular uptake mechanism, and off-target binding profile |
tRNA is a compact example of chemical specialization. A tRNA is not simply an adaptor with an anticodon. It contains many modified nucleotides, conserved tertiary contacts, acceptor-stem identity elements, anticodon-loop chemistry, and structured surfaces for aminoacyl-tRNA synthetases, elongation factors, ribosomes, and quality-control enzymes. Wobble and modification chemistry in the anticodon loop directly affect decoding, while modifications elsewhere stabilize folding and recognition.
rRNA demonstrates how RNA chemistry can support a large molecular machine. Ribosomal RNA contains canonical helices, noncanonical contacts, modified residues, metal ions, and protein interfaces. Its chemistry helps scaffold ribosome architecture and position functional groups. The ribosome is not a simple protein enzyme with an RNA scaffold; it is a ribonucleoprotein machine whose active center depends deeply on RNA structure and chemistry.
Ribozymes and riboswitches show that RNA can fold and react. A self-cleaving ribozyme positions a phosphodiester linkage for reaction. A riboswitch aptamer uses bases, backbone, ions, and sometimes modifications to bind a metabolite and alter gene expression. These examples make clear why RNA chemistry cannot be reduced to sequence complementarity.
Viral RNAs and therapeutic RNAs expose the same chemistry to different selective pressures. Viral RNAs must be copied, translated, packaged, protected, and sensed or hidden from host defenses. Therapeutic RNAs must be manufactured, purified, delivered, and made active in a tissue without unacceptable immune activation or toxicity. In both cases, end chemistry, duplex character, modifications, structure, and protein interactions are central.
Computational RNA analysis often begins with sequence, but chemistry determines what sequence can mean. Secondary-structure prediction models usually treat base pairing and stacking through thermodynamic parameters, but they approximate ions, modifications, protein binding, co-transcriptional folding, and cellular crowding. A predicted helix is a hypothesis about chemistry, not a direct observation. Probing-constrained models improve realism but inherit the interpretation limits of the probes.
RNA design uses complementarity, folding, and modification rules to build molecules. Inverse-folding methods seek sequences that adopt desired structures. mRNA design balances coding sequence, codon usage, local structure, modification, untranslated regions, and manufacturing constraints. Small RNA design balances target affinity, specificity, strand bias, chemical stabilization, and delivery. These applications are not separate from basic chemistry; they are attempts to engineer within chemical constraints.
Clinical RNA technologies depend on chemistry at every step. Antisense oligonucleotides use backbone and sugar modifications to tune RNase H recruitment or steric blocking. siRNAs use duplex geometry, chemical stabilization, strand selection, and conjugation. mRNA vaccines and therapeutics use modified nucleotides, cap analogs, purification, and lipid nanoparticles. Antiviral nucleoside analogs exploit polymerase incorporation and chain termination or mutagenesis. RNA-directed small molecules rely on structured RNA pockets and selectivity over abundant cellular RNAs. Chapters 149 through 162 develop these technologies, but the chemical logic begins here.
Current consensus supports several broad conclusions. RNA’s ribose-phosphate backbone, nucleobases, and 2′-hydroxyl group jointly determine RNA behavior; no single feature explains RNA biology by itself. The phosphate backbone’s negative charge shapes folding, ribonucleoprotein assembly, purification, condensation, and delivery. The 2′-hydroxyl group contributes to RNA geometry, recognition, and cleavage chemistry, but RNA stability depends on context rather than on the 2′-hydroxyl alone. Canonical Watson-Crick pairs are fundamental, but noncanonical pairing and stacking are routine parts of RNA structure.
There is also consensus that chemical methods require careful interpretation. RNA chemical probing, modification mapping, and sequencing signatures can be highly informative, but the measured signal is mediated by reagent access, local conformation, enzymatic behavior, library construction, and analysis models. Orthogonal validation is especially important for claims about rare modifications, transient structures, or causal regulatory effects.
For engineered RNA, the consensus is pragmatic: modified nucleotides and synthetic analogs can improve stability, affinity, immune tolerance, delivery, or potency, but each chemical change creates a profile of effects rather than a universal benefit. Platform-specific testing remains necessary.
Open questions:
Common misconceptions:
Deprecated or weakened claims: