Chapter 53. RNA Structural Motifs, Long-Range Contacts, and Modular Architectures

Scope Note

This chapter explains recurring RNA secondary-structure motifs, tertiary contacts, long-range interactions, modular architectures, and noncanonical conformers as physical features of RNA molecules rather than as decorative diagram elements. It is the primary motif-by-motif owner for RNA G-quadruplexes, intramolecular RNA triple helices, and left-handed Z-RNA. It emphasizes how motifs are defined, how ions and sequence change their energetics, how cellular proteins remodel them, how each class is detected, and where motif language can overstate context-independent formation or function.

Executive Summary

RNA architecture is modular but not Lego-like. A stem, bulge, loop, junction, pseudoknot, kink-turn, A-minor contact, ribose zipper, or ligand-binding pocket can recur across unrelated RNAs, yet each motif works inside a sequence, ion environment, folding pathway, and RNA-protein complex. The useful abstraction is therefore “recurrent structural solution,” not independent plug-in part. Secondary-structure motifs define local Watson-Crick and wobble pairing patterns. Tertiary motifs define recurrent three-dimensional contacts among bases, ribose-phosphate backbones, ions, ligands, and proteins. Long-range contacts connect distant sequence segments and allow a one-dimensional transcript to become a compact functional object.

Canonical secondary structures are the first map of an RNA fold. Helices, hairpins, internal loops, bulges, and multibranch junctions can often be represented by dot-bracket notation or related graph encodings. These maps are valuable because they describe many stable base pairs and provide a scaffold for thermodynamic prediction. They are incomplete because they usually omit noncanonical base pairs, coaxial stacking, tertiary contacts, pseudoknots, ion binding, and protein stabilization. Reviews of RNA architecture emphasize that secondary structure is a necessary coordinate system, not a complete physical model.

Tertiary motifs make RNA architecture recurrent. A-minor interactions, tetraloop-receptor contacts, ribose zippers, kink-turns, sarcin-ricin-like loops, base triples, and other motifs use non-Watson-Crick edges and backbone geometry to pack helices and loops. Isostericity matrices and motif-centered annotation help explain why RNA sequences can vary while preserving three-dimensional geometry. Recent surveys of long-range tertiary interactions in noncoding RNA structures show that motif catalogs are now large enough to support systematic annotation, but they remain biased toward solved structured RNAs and high-resolution models.

Comparative evidence is the strongest way to separate conserved architecture from accidental pairing. Covariation, compensatory mutations, and preservation of noncanonical pair geometry can show that a structural feature is under evolutionary constraint even when the primary sequence changes. Rivas’s review of RNA sequence and structure conservation is a core source for this logic. Conservation is not automatic proof of one structure in every context; it is evidence that must be interpreted with phylogeny, alignment quality, expression context, and experimental validation.

Three noncanonical conformer families make the distinction between folding capacity and cellular occupancy especially important. RNA G-quadruplexes stack planar guanine tetrads around a monovalent-cation channel; RNA triple helices add a third strand to the major groove of a duplex; Z-RNA reverses the handedness and backbone path of an RNA duplex. Each can be stable in a favorable purified system. None should be declared constitutively folded in cells from sequence alone. Transcriptome-wide rG4 experiments, single-site structures, ligand-capture maps, helicase perturbations, and Z-form protein genetics instead support regulated ensembles whose occupancy depends on ion activity, neighboring structure, translation, proteins, and cellular state.

Concept Inventory

  • RNA structural motif: a recurrent local or long-range arrangement of RNA nucleotides, base pairs, backbone turns, ions, ligands, or protein contacts that is recognizable across one or more RNA structures. Motifs can be described at several levels: a secondary-structure motif such as a hairpin, a base-pairing motif such as a kink-turn, a tertiary packing motif such as a tetraloop-receptor interaction, or a functional motif such as a ligand-binding pocket.
  • Secondary structure: the pattern of intramolecular and intermolecular base pairs that can usually be drawn as helices and loops. The term is most often used for canonical Watson-Crick and G-U wobble pairs, but real secondary-structure annotations may include some noncanonical pairs. Secondary structure is a model of paired positions, not a complete three-dimensional structure.
  • Tertiary structure: the three-dimensional arrangement of the RNA chain after helices, loops, junctions, long-range contacts, ions, ligands, and proteins are considered. Tertiary structure includes base triples, stacking, backbone turns, noncanonical base pairs, metal-ion sites, ligand pockets, and packing between distant sequence segments.
  • Long-range contact: a physical interaction between RNA segments that are distant in primary sequence or belong to different RNA molecules. Long-range contacts include pseudoknots, kissing loops, A-minor loop-helix interactions, ribose zippers, long-distance base-pairing helices, and RNA-RNA regulatory interactions.
  • Pseudoknot: a base-pairing arrangement in which nucleotides from a loop pair with a complementary region outside the loop, producing crossing pair relationships that cannot be represented as a single nested secondary structure. Pseudoknots can stabilize catalytic centers, regulate translation, induce frameshifting, and shape viral RNA genomes.
  • Modular architecture means that an RNA fold can be understood as a hierarchy of reusable elements: helices, junctions, tertiary contacts, and stabilizing partners. Modularity is partial. A motif transplanted into a new context may fail if spacing, orientation, ion conditions, or protein partners are incompatible.
  • Comparative covariation: correlated sequence change across homologous RNAs that preserves a base pair or structural geometry. A classic example is a G-C pair changing to an A-U pair at aligned positions while retaining pairing. Covariation can also preserve noncanonical pair geometry through isosteric substitutions.
  • RNA G-quadruplex (rG4): a four-stranded arrangement in which guanines hydrogen-bond through their Hoogsteen edges to form planar G-quartets that stack around a central monovalent-ion channel. A guanine-rich sequence is an rG4-forming candidate, not proof that the RNA is folded into an rG4 in a particular cell.
  • RNA triple helix: an RNA architecture in which a third strand recognizes a duplex, commonly in its major groove, to create stacked base triples such as U•A-U. An RNA triple helix is distinct from a single isolated base triple and from an RNA-DNA triplex at chromatin.
  • Z-RNA: a left-handed RNA duplex whose zig-zag phosphate backbone, alternating syn and anti glycosidic conformations, and alternating sugar puckers distinguish it from the usual right-handed A-form duplex. Z-RNA is a higher-free-energy state whose formation is normally conditional and can be stabilized by Zα-domain proteins.
  • Folding capacity versus occupancy: the distinction between a sequence that can adopt a conformer under a specified in-vitro condition and the fraction of molecules that actually occupy that conformer in a defined cellular compartment and time window.

What to Know Before Reading This Chapter

The reader should know that RNA is a polymer with a 5′ to 3′ direction, a negatively charged phosphodiester backbone, ribose sugars with 2′ hydroxyl groups, and four common bases that can form canonical Watson-Crick pairs. The physical chain can fold back on itself, pair with another RNA, bind proteins, bind small molecules, and interact with metal ions. RNA structure is therefore not just a drawing of paired bases; it is a molecular ensemble.

Finally, the reader should read structural claims as evidence-scaled statements. A covariation-supported helix in a curated alignment, a cryo-EM-resolved tertiary contact in a ribosome, a chemical-probing-supported structure in living cells, and a computationally predicted motif are not equivalent. Each is useful, but each answers a different question. This chapter repeatedly separates observed structure, inferred conserved architecture, predicted motif, and functional mechanism.

The physical principles behind ion condensation, base stacking, conformational free-energy landscapes, and kinetic trapping are developed in Chapter 4. The present chapter applies those principles to named motifs. When a motif touches DNA topology, innate immunity, or RNA editing, this chapter owns the structural definition and evidence limits, while Chapter 97, Chapter 108, and Chapter 109 own the corresponding pathway mechanisms.

53.1. Canonical secondary-structure motifs

Canonical RNA secondary structure begins with paired stems. A stem is a run of base pairs, usually dominated by A-U, G-C, and G-U pairs, that forms an A-form helix. The A-form geometry matters because it positions the major and minor grooves differently from B-form DNA and makes the 2′ hydroxyl available for hydrogen bonding and backbone recognition. Stems provide the rigid rods of many RNA architectures. They can define domains, present loop motifs, organize ribozyme active sites, or create landing surfaces for RNA-binding proteins.

Figure 53.1. From Secondary Motifs to Tertiary Architecture

Figure 53.1. From Secondary Motifs to Tertiary Architecture. RNA structure is hierarchical. Secondary-structure motifs provide the paired scaffold, long-range contacts connect distant regions, and tertiary interactions, ions, ligands, and proteins stabilize the functional architecture.

A hairpin is a stem capped by a terminal loop. The loop is not merely unstructured slack. Loop sequence and size can determine whether the hairpin remains flexible, forms a stable tetraloop, docks into a receptor, binds a protein, or nucleates an alternative fold. GNRA tetraloops and UNCG tetraloops are classic examples of small loops with recurrent local geometry. In many RNAs, a hairpin is a display platform: the stem positions the loop, and the loop provides the recognition or docking surface. In miRNA precursors, by contrast, hairpin shape and processing context help recruit processing enzymes; the same generic motif name therefore covers different functional regimes.

Internal loops interrupt a helix with unpaired residues on both strands. Bulges interrupt a helix on one strand. These interruptions can bend the helix, widen or narrow grooves, expose bases, and create protein-binding or tertiary-docking sites. An internal loop may contain noncanonical base pairs that preserve a local helical stack while changing chemical recognition. A bulged nucleotide may flip out to contact a protein or ligand, or it may stack into the helix and subtly change curvature. Treating every unpaired residue as “single-stranded” is a common oversimplification; many loop and bulge residues participate in ordered interactions.

Multibranch junctions connect three or more helices. Junctions are architectural hubs because they determine helix orientation. A three-way junction can organize a riboswitch aptamer; a four-way junction can create a scaffold for ribozyme domains; a larger junction can anchor rRNA expansion segments or viral RNA domains. Junction geometry is influenced by sequence, coaxial stacking, divalent ions, ligand binding, and proteins. A junction drawn as a flat node in a secondary-structure diagram may correspond to a compact three-dimensional switch in the actual molecule.

The evidence for canonical secondary motifs comes from several layers. Comparative sequence analysis identifies conserved pairing patterns. Enzymatic and chemical probing report accessibility and local flexibility. Mutational studies test whether disrupting and restoring base pairs changes function. High-resolution structures provide atomic geometry. Each method has boundaries. Chemical reactivity can reflect protein binding or local dynamics rather than absence of pairing. Covariation depends on alignment quality and phylogenetic depth. A compensatory mutation may restore a pair but also change a protein-binding site. Good secondary-structure annotation integrates, rather than substitutes, these evidence classes.

Do not overgeneralize secondary-structure diagrams. A dot-bracket model is usually a representative model, not a complete ensemble. An RNA can sample alternative helices, co-transcriptional intermediates, ligand-bound and ligand-free structures, protein-stabilized states, and misfolded traps. Secondary motifs are therefore best read as a coordinate system for asking better structural questions: Which helices are conserved? Which loops are exposed? Which unpaired regions are structured by tertiary contacts? Which elements are stable alone, and which require proteins, ions, or ligand?

53.2. Long-range contacts and pseudoknots

Long-range contacts connect parts of an RNA that are far apart in sequence. They solve a central problem of RNA architecture: a linear polymer must bring selected segments together while leaving other segments accessible for translation, processing, catalysis, or recognition. Long-range contacts include canonical base-pairing helices, loop-loop kissing interactions, loop-helix docking, tertiary base triples, ribose-phosphate contacts, and protein-bridged interactions. They may form within one RNA molecule or between two RNA molecules.

Pseudoknots are the most familiar long-range secondary-structure exception. In a simple hairpin, a loop closes one stem. In a pseudoknot, loop nucleotides pair with a downstream or upstream complementary segment, creating interleaved base pairs. This crossing relationship cannot be represented by ordinary nested parentheses without extra notation. Pseudoknots are found in ribozymes, telomerase RNA, ribosomal RNA, viral untranslated regions, programmed ribosomal frameshifting elements, and many predicted structured RNAs. They can position catalytic residues, compact a fold, resist helicase movement, or create mechanical tension during translation.

Long-range RNA-RNA interactions also include intermolecular contacts. Bacterial small RNAs can base-pair with mRNAs to change translation or decay. Eukaryotic guide RNAs can direct editing or cleavage. Viral RNAs can circularize through contacts between terminal regions. Xue’s review of RNA-RNA interaction architecture emphasizes that RNA contacts can operate as structural scaffolds, regulatory switches, and assembly intermediates. The same base-pairing logic appears in many classes, but the cellular machinery and consequences differ.

Figure 53.2. Why a Pseudoknot Is Not a Nested Hairpin

Figure 53.2. Why a Pseudoknot Is Not a Nested Hairpin. A pseudoknot forms when loop residues pair with a distal complementary segment, creating interleaved base pairs that cannot be represented by ordinary nested notation. Functional interpretation requires evidence beyond a predicted crossing pattern, including covariation of both stems, mutational rescue, and ideally a high-resolution structure or functional assay.

The evidence for long-range contacts is strongest when independent methods converge. Comparative covariation can support distant pairing if compensatory changes preserve both sides of the interaction. Mutational rescue can show that the pairing, not the exact sequence, is important. Structure probing can reveal coordinated protection or reactivity changes. Crosslinking, proximity ligation, and high-throughput interaction mapping can nominate RNA-RNA contacts, but these methods can be biased by abundance, ligation efficiency, proximity without direct pairing, and cell lysis artifacts. High-resolution structures provide geometry but often use purified fragments or stabilized complexes.

Pseudoknots and long-range contacts challenge computational prediction. Standard dynamic programming algorithms for minimum free energy folding are efficient partly because they assume nested base pairs. General pseudoknot prediction is more complex and often requires restricted pseudoknot classes, heuristics, comparative information, or external constraints. Karan and Rivas’s all-at-once RNA folding framework with three-dimensional motif prediction illustrates the field’s move toward combining evolutionary information and motif-aware modeling rather than treating secondary and tertiary prediction as separate tasks. That approach is promising, but predictions remain evidence proposals until tested in the biological context.

The boundary case is alternative contact choice. A region predicted to participate in a pseudoknot may instead pair with an antisense RNA, bind a protein, form a local hairpin during transcription, or remain unfolded because a helicase or ribosome remodels it. In long transcripts, multiple weak contacts can compete. The functional structure may be a regulated ensemble, not a single global minimum. This is why long-range contact claims should specify whether the contact was observed in vitro, inferred from evolution, detected in cells, or predicted computationally.

53.3. Tertiary motifs and modular architecture

Tertiary motifs are recurrent three-dimensional solutions for packing RNA. They often use base edges that are not involved in Watson-Crick pairing, ribose 2′ hydroxyl groups, phosphate oxygens, base stacking, and hydrated metal ions. A tertiary motif may bring a loop into the minor groove of a helix, stack two helices across a junction, create a sharp backbone turn, or assemble a ligand-binding pocket. These motifs explain how RNAs with different sequences can build similar folds.

The A-minor interaction is one of the most important tertiary contacts. In an A-minor contact, an adenosine docks into the minor groove of a helix and makes shape- and hydrogen-bond-compatible contacts with a base pair. A-minor contacts are abundant in ribosomal RNA and other structured RNAs because they allow loops and single-stranded segments to recognize helical geometry. The motif demonstrates why the RNA minor groove is not chemically featureless. Its pattern of hydrogen-bond acceptors, donors, and backbone atoms can be read by RNA itself.

Figure 53.3. Gallery of Recurrent Tertiary Motifs

Figure 53.3. Gallery of Recurrent Tertiary Motifs. Recurrent tertiary motifs — including A-minor interactions, tetraloop-receptor contacts, ribose zippers, kink-turns, and base triples — use non-Watson-Crick base edges, ribose hydroxyl groups, backbone geometry, and helix packing to build compact RNA folds across unrelated RNA classes.

Tetraloop-receptor interactions are another modular architecture. A stable tetraloop, often a GNRA loop, docks into a receptor motif elsewhere in the RNA. This can fasten distant helices together with high specificity. The motif is modular enough to appear in natural RNAs and engineered scaffolds, but it is not context-free. The length and orientation of connecting helices, magnesium concentration, receptor integrity, and competing folds determine whether docking occurs. In ribozymes and riboswitches, tetraloop-receptor contacts often stabilize the active or ligand-bound architecture.

Kink-turns create sharp bends in RNA helices through a pattern of canonical and noncanonical pairs. They are common in ribosomal and small nucleolar RNPs and often serve as protein-binding motifs. A kink-turn may be partially preorganized by RNA sequence and then stabilized by a protein. This is a general theme for RNA motifs: the RNA can encode a tendency, while the final cellular fold may require an RNP partner. Motif annotation should therefore record whether a motif is RNA-intrinsic, protein-stabilized, or observed only in a complex.

Lescoute and colleagues emphasized recurrent RNA motifs, isostericity matrices, and sequence alignments as a unified way to connect sequence variation with three-dimensional preservation. Isosteric substitutions preserve the geometry of a base pair or motif even when the identities of the bases change. This explains why a noncanonical pair can be conserved structurally without being conserved as a literal sequence. The crucial lesson is that conservation of structure can look unlike conservation of letters.

Graph-based and motif-search methods now allow systematic annotation of tertiary base motifs and substructures. Emrizal and colleagues reviewed graph theoretical workflows for searching and annotating RNA tertiary motifs. Bohdan and colleagues surveyed long-range tertiary interactions and motifs in noncoding RNA structures, showing how recurrent contacts can be cataloged across solved structures. These resources help convert visual structural intuition into searchable data, but their coverage is shaped by which RNAs have solved structures and by the resolution and curation of those structures.

The common misconception is that finding a motif name explains function. It does not. A motif label explains a structural pattern. Function depends on whether the motif forms in the cellular context, whether it changes a molecular interaction, and whether that interaction affects the RNA life cycle. A kink-turn in a predicted structure may be a protein-binding hypothesis. A tetraloop may be a docking element or simply a stable cap. Annotation must be followed by mechanism.

Figure 53.4. Comparative Evidence for Structural Conservation

Figure 53.4. Comparative Evidence for Structural Conservation. RNA structure can be conserved at several levels. Exact sequence conservation, compensatory canonical base-pair changes, and isosteric noncanonical substitutions each support different structural inferences, and distinguishing among them requires well-curated alignments and phylogenetically informed models.

Figure 53.5. RNA Structural Annotation Workflow

Figure 53.5. RNA Structural Annotation Workflow. Structural annotation converts RNA folding evidence into searchable motif records by progressing from raw sequence, alignment, probing data, or atomic model through base-pair and contact annotation, motif search by graph or geometric methods, database cross-reference, and evidence grading, but each annotation must retain evidence type, uncertainty, and biological context.

Table 53.1. Major RNA Motif Classes and Interpretation Rules. Each major RNA motif class requires evidence appropriate to its geometry and proposed function before the motif label can support a mechanistic claim.

Motif class Minimal structural definition Typical evidence Functional examples Common overinterpretation Needed validation
Hairpin loop Stem closed by a terminal loop, commonly 4–6 nt Chemical probing, sequence conservation, thermal melting GNRA tetraloop, miRNA precursor processing, protein recruitment Any capped stem assumed to have a functional loop role Loop mutagenesis, protein-binding or docking assay
Internal loop or bulge Unpaired residues interrupting one or both strands of a helix Chemical probing, covariation, high-resolution structure Helix bending, groove exposure, protein or ligand contact Treated as unstructured rather than conformationally ordered Mutational rescue, binding assay, atomic-resolution structure
Multibranch junction Node connecting three or more helical arms Secondary-structure prediction, covariation, cryo-EM Riboswitch aptamer scaffold, ribozyme domain organization Flat junction assumed without considering three-dimensional geometry Tertiary-structure determination, ion-dependent folding assay
Pseudoknot Loop residues pair with a region outside the loop, creating crossing base pairs Covariation of both stems, mutational rescue, high-resolution structure Frameshifting elements, telomerase RNA, ribozyme active site Every crossing prediction assumed functional or frameshifting Mutagenesis of both stems, functional assay, structural verification
Kissing-loop interaction Watson-Crick pairing between loop residues of two separate hairpins Chemical probing, mutational rescue, crosslinking RNA dimerization, viral genome circularization, regulatory RNA pairing Loop complementarity alone taken as proof of contact Binding affinity measurement, covariation, in-cell detection
A-minor interaction Adenosine docking into the minor groove of an adjacent helix via shape-complementary hydrogen bonds High-resolution structure, conservation of adenosine identity Ribosomal RNA packing, ribozyme docking, helix-helix contacts Any conserved A near a helix assumed to form an A-minor contact Mutational disruption of adenosine, atomic-structure confirmation
Tetraloop-receptor interaction GNRA or related tetraloop docking into a specific receptor sequence on a distant helix High-resolution structure, mutagenesis, in vitro folding assay Group I intron folding, riboswitch tertiary docking, engineered scaffolds Tetraloop alone assumed sufficient; ignores receptor or helix geometry Mutational rescue of both tetraloop and receptor, thermodynamic measurement
Kink-turn Asymmetric internal loop flanked by canonical and G-A pairs creating a sharp backbone bend High-resolution structure, sequence motif conservation, protein-binding assay Ribosomal protein binding, snoRNP assembly, helix orientation Predicted kink-turn assumed functional without protein stabilization data Protein-binding assay, structural confirmation, protein-depletion experiment
Ribose zipper 2′-OH groups of one strand hydrogen-bonded to 2′-OH or base of another strand High-resolution crystal or cryo-EM structure, hydroxyl-specific mutagenesis Ribozyme tertiary packing, ribosomal RNA contacts, tetraloop-receptor interface Backbone contacts assumed nonspecific without structural evidence 2′-deoxy substitution, structural confirmation
Ligand-binding pocket RNA three-dimensional cavity with specific contacts to a small molecule, ion, or metabolite High-resolution structure with ligand, ITC or SPR binding assay Riboswitch aptamer, SAM-binding site, aminoglycoside-binding rRNA In vitro binder or predicted pocket assumed to be a validated drug target Cellular occupancy evidence, selectivity assay, functional consequence

Table 53.2. Evidence Ladder for RNA Structural Motifs. Evidence types answer different questions about RNA structure, and robust motif interpretation usually requires combining comparative, biochemical, structural, and functional evidence rather than relying on any single method.

Evidence type What it can support What it cannot prove alone Typical artifact or limitation Best companion evidence
Computational prediction Candidate base pairs, pseudoknot hypotheses, motif nominations In vivo structure, biological function Algorithm-dependent; pseudoknots often excluded or treated heuristically Comparative covariation, chemical probing
Thermodynamic folding Minimum free energy secondary structure, folding stability ranking Cellular fold, tertiary contacts, protein or ion effects Ionic conditions and temperature may not match cell; kinetics ignored Chemical probing, high-resolution structure
Comparative covariation Conserved base pairs, isosteric noncanonical pairs, conserved motif geometry Functional importance, three-dimensional coordinates Requires deep phylogeny; alignment errors and shared ancestry mislead inference Mutational rescue, high-resolution structure
Chemical probing Nucleotide accessibility, local flexibility, protection by protein or ligand Specific base pair identity, tertiary contacts, in vivo context Protein binding and local dynamics can mimic or mask pairing signals Comparative covariation, mutational rescue
Crosslinking or proximity ligation Candidate RNA-RNA or RNA-protein contacts, long-range interaction nominations Direct base pairing, biological role Ligation bias, proximity without pairing, cell-lysis artifacts Mutational rescue, high-resolution structure
Mutational rescue Functional dependence on a base pair or motif, pairing geometry Three-dimensional coordinates, molecular mechanism Compensatory mutation may alter protein binding or other structural elements High-resolution structure, chemical probing
High-resolution structure Atomic geometry of base pairs, motifs, ions, ligands, and backbone contacts Cellular fold under in vivo conditions, dynamic ensemble behavior Crystal packing, cryo-EM conformational selection, truncated or modified construct Chemical probing, functional assay
Functional assay Structural dependence on a biological activity or output Structural mechanism, which residues mediate the effect Indirect effects, reporter artifacts, off-target sequence changes High-resolution structure, mutational rescue

Box 53.1. Do Not Equate Motif Discovery With Function

  • A motif label describes a structural pattern, not a mechanism.
  • A conserved motif suggests evolutionary constraint, not a specific function.
  • A predicted motif is a structural hypothesis until tested in context.
  • A functional motif requires evidence that perturbing it changes an RNA fate or biological output.
  • The correct inference depends on evidence level: discovery, conservation, prediction, and demonstrated function are related but distinct claims.

Box 53.2. Minimum Metadata for an RNA Motif Record

  • Stable motif record ID.
  • RNA molecule name and organism.
  • RNA class (e.g., rRNA, riboswitch, viral genomic RNA).
  • Sequence interval and strand orientation.
  • Motif class (e.g., pseudoknot, A-minor interaction, kink-turn).
  • Structural evidence type and source (e.g., cryo-EM, chemical probing, covariation).
  • Functional evidence and phenotype tested.
  • Conservation evidence across phylogeny.
  • Bound ions, ligands, proteins, or RNA modifications affecting the motif.
  • Database cross-references (e.g., PDB, Rfam, RR3DD).
  • Source citations and explicit caveats on evidence strength.

53.4. RNA G-quadruplexes, RNA triplexes, and Z-RNA conformers

RNA G-quadruplexes, RNA triple helices, and Z-RNA are often grouped as “noncanonical” structures, but that label should not conceal their different molecular logic. An RNA G-quadruplex is built from stacked four-guanine planes. An RNA triple helix adds a third strand to a pre-existing duplex. Z-RNA is a left-handed state of a duplex rather than an additional strand or a four-stranded stack. All three are conformers: the same RNA sequence may partition among the named structure, ordinary A-form helices, hairpins, protein-bound states, and unfolded or partially folded states. A useful structural claim must therefore state sequence, ion composition, temperature, molecular concentration, binding partners, and whether the measurement reports folding capacity or occupancy.

RNA G-quadruplexes: quartets, topology, and ion dependence

An RNA G-quadruplex (rG4) begins with a G-quartet, also called a G-tetrad. Four guanine bases arrange in a near-planar ring and use Hoogsteen-edge hydrogen bonds rather than Watson-Crick pairing. Two or more quartets can stack, forming a central electronegative channel coordinated by monovalent cations. Potassium generally stabilizes many rG4s more strongly than sodium because the dehydrated ion fits favorably between adjacent quartet planes; lithium usually supports them poorly. This ion selectivity is not a universal ranking for every sequence, but it is a powerful diagnostic when a potassium-dependent spectral or reverse-transcription signature is lost after guanine substitutions.

Sequence determines what structures are accessible. Four runs of guanines connected by loops can form an intramolecular quadruplex; guanines supplied by two or four RNA molecules can instead create dimeric or tetrameric structures at sufficiently high concentration. The number of quartets, loop lengths, bulges within G-runs, flanking bases, and alternative Watson-Crick partners all change stability and kinetics. RNA’s 2′ hydroxyl and A-form preferences often favor compact, parallel-stranded rG4 topologies, but “parallel” is not a synonym for “identical”: loop geometries, bulges, quartet number, capping interactions, and multimeric state can differ. A short canonical sequence pattern such as G-rich tracts is therefore a search rule, not a structural measurement.

The central competition is frequently rG4 versus an ordinary stem-loop. If some guanines can pair with nearby cytidines, the same segment may choose between Hoogsteen quartets and Watson-Crick helices. RNA concentration and transcription history matter because an intermolecular quadruplex is disfavored at low strand concentration, whereas a local hairpin can form unimolecularly. Potassium can shift the rG4 side of the landscape, and a ligand can shift it further; a helicase, translating ribosome, or structure-disfavoring protein can push the population toward another state. A melting temperature measured for an isolated oligonucleotide is thus not the cellular occupancy of that motif.

The 5′ untranslated region of an NRAS transcript provided a classic single-site example. Circular dichroism, thermal measurements, mutagenesis, and cell-free reporters supported an rG4 capable of repressing translation. Small molecules could bind this folded construct and reduce reporter translation in vitro. Later transcript-resolved work supplied an essential boundary condition: across 14 tested cell lines, less than 1% of measured NRAS transcripts contained the proposed G4-bearing 5′ region, the predominant transcription start site lay downstream of that region, and a ligand that bound the isolated rG4 produced only moderate cellular effects. The lesson is not that the physical rG4 was unreal. The lesson is that isoform abundance, structure occupancy, compound selectivity, and cellular response are separate propositions requiring separate measurements.

Figure 53.6. RNA G-Quadruplex Potential, Occupancy, and Remodeling

Figure 53.6. RNA G-Quadruplex Potential, Occupancy, and Remodeling. A four-panel original schematic should show (1) guanines assembling into a Hoogsteen-bonded quartet, (2) stacked quartets coordinating potassium in the central channel, (3) competition between an rG4 and a Watson-Crick stem-loop, and (4) shifts in the ensemble caused by DHX36 or DHX9, a translating ribosome, an rG4-binding protein, or a stabilizing ligand. A side strip should distinguish rG4-seq folding potential, in-cell chemical probing, and fragment-scale G4RP-seq probe-dependent enrichment; the G4RP panel should label pre-lysis formaldehyde crosslinking, roughly 200–800-nucleotide fragments, and possible indirect crosslink carryover. The figure must not portray every guanine-rich RNA as folded in cells.

From rG4 folding potential to cellular occupancy

Single-site rG4 evidence normally combines orthogonal measurements. Circular dichroism can report a spectrum compatible with a parallel quadruplex but is not uniquely structural. UV melting near 295 nm, potassium-versus-lithium comparisons, native gels, nuclear magnetic resonance, X-ray crystallography, selective chemical modification, reverse-transcriptase pausing, and guanine-to-adenine or guanine-to-cytosine substitutions each constrain a different feature. A mutation is most informative when it disrupts quartets without simply creating a new stable hairpin or changing a protein motif. Ligand fluorescence and antibody staining can visualize candidate rG4s, but the probe may stabilize the structure it is intended to observe.

At transcriptome scale, rG4-seq exploits reverse-transcriptase stalling under rG4-favoring conditions, with comparison conditions and stabilizing ligands used to identify thousands of sites that can form rG4s in extracted human RNA. This is a map of sequence-dependent folding potential under the assay conditions. It is not an atlas of constitutively folded rG4s in living cells. Guo and Bartel compared in-vitro and in-cell chemical probing and found that thousands of mammalian regions with rG4 potential were overwhelmingly unfolded in eukaryotic cells; the same study supported a model of active eukaryotic unfolding and evolutionary depletion of strongly folding motifs in bacteria. That result corrected a common inference from motif counts while leaving room for transient, regulated, or condition-specific rG4 formation.

G4RP-seq asks a related but different question. It formaldehyde-crosslinks cells before lysis, fragments cellular material, captures RNAs with the G4-binding probe BioTASQ, and identifies enriched transcripts. The method detected transient candidate rG4s and changes after treatment with G4-stabilizing ligands. Capture depends on probe accessibility, affinity, crosslinking, transcript abundance, and enrichment thresholds. The protocol normally recovers fragments of roughly 200–800 nucleotides, so it does not localize a fold at nucleotide resolution, and a minor fraction of indirectly crosslinked RNA can accompany a captured target. Because BioTASQ recognizes and may stabilize the target conformation after crosslinking, enrichment means “capturable by this G4-selective workflow under this perturbation,” not a molecule-by-molecule occupancy fraction. Agreement among rG4-seq potential, in-cell probing, G4RP enrichment, site-specific mutation, and a functional assay is substantially stronger than any one layer.

Newer low-input implementations extend rG4-seq to scarce material and physiological perturbations. Ultra-low-input rG4-seq applied to mouse oocytes linked rG4 landscapes with acute or chronic DHX36 loss, while the divergent responses to the two perturbation regimes warned that adaptation can reverse a simple “less helicase means more folded substrate” prediction. Resolution, RNA abundance, developmental state, and knockout compensation remain important. The advance expands experimental reach; it does not erase the distinction between reverse-transcription signatures ex vivo and direct occupancy in every living molecule.

Protein remodeling and ligand stabilization of rG4 ensembles

Proteins interact with rG4s as binders, resolvases, folding partners, or competitors. DHX36 is a prominent ATP-dependent helicase with high affinity for many G4 substrates; DHX9, DDX3X, DDX5, DDX17, GRSF1, and other proteins have been linked to particular rG4 contexts. Affinity proteomics on the NRAS rG4 identified a large protein assembly and connected DDX3X binding to G4-containing 5′ untranslated regions. A separate study linked rG4s near upstream open reading frames to DHX36- and DHX9-dependent translation outcomes. These results do not define a universal direction of regulation. A helicase can remove a roadblock, remodel a structure needed for initiation, release a bound protein, or act on a competing helix; chronic depletion can also change expression and induce compensation.

Small molecules add another layer of selection. Many planar aromatic ligands stack on terminal quartets, while substituents contact grooves, loops, or flanking bases. A compound can increase lifetime without creating absolute specificity for one transcript. DNA G4s, other RNA G4s, duplex RNA, proteins, membranes, and cellular compartments all compete for ligand. An apparent decrease in target protein can reflect altered transcription, splicing, RNA decay, translation, stress responses, or toxicity. Good ligand evidence therefore combines biophysical affinity and selectivity, structural or footprint evidence for the binding site, target-site mutation, cellular exposure, transcript-isoform analysis, and a mechanism-linked output. The NRAS history illustrates why an elegant in-vitro target can remain a weak cellular target if the relevant isoform is scarce.

RNA triple helices: third-strand recognition of a duplex

An RNA triple helix forms when a third RNA segment binds along a duplex and makes a repeated series of base triples. In the major-groove triplexes best characterized in cellular and viral RNA, a pyrimidine-rich Hoogsteen strand recognizes the purine-rich strand of a Watson-Crick duplex. A U•A-U triple contains an A-U Watson-Crick pair plus a uridine that hydrogen-bonds to the adenosine’s Hoogsteen edge. C•G-C triples can depend on protonation and therefore show pH-sensitive stability. Stacking between consecutive triples, coaxial stacking with adjacent duplexes, terminal clamps, bulges, and A-minor contacts can matter as much as a simple count of U•A-U units.

The third strand may be continuous with the duplex-forming RNA or supplied by another RNA segment. This chapter uses “RNA triple helix” for RNA-only structures and distinguishes them from RNA-DNA triplexes proposed at chromatin. It also distinguishes an extended triple helix from a single base triple embedded in an ordinary tertiary fold. Triplex sequence searches are harder than canonical stem searches because the relevant segments can be separated in primary sequence, allowed triples are context-dependent, and peripheral helices can anchor the core.

The element for nuclear expression (ENE) in Kaposi sarcoma-associated herpesvirus polyadenylated nuclear RNA is a mechanistically concrete example. A U-rich internal loop clamps an A-rich segment of the poly(A) tail into stacked U•A-U triples. A 2.5-Å crystal structure, mutational analysis, binding assays, and nuclear-extract deadenylation experiments showed how this arrangement sequesters the tail from decay machinery. The structure supports a causal chain: the ENE provides two U-rich surfaces, the A-rich strand forms Watson-Crick and Hoogsteen contacts simultaneously, and the resulting clamp impedes deadenylase access or progression. The ENE can recognize an internal compatible register within a longer poly(A) tail and arrest shortening before the enzyme reaches the RNA body, so protection is not restricted to a triplex fixed exactly at the 3′ terminus. This mechanism does not imply that every U-rich hairpin and poly(A) segment form a stable triplex.

The 3′ end of the long noncoding RNA MALAT1 provides a related cellular architecture. RNase P processing creates a non-polyadenylated MALAT1 end containing an A-rich tract. Biochemical perturbation and reporter assays showed that the conserved U-rich element and A-rich tract form a protective triple helix. Systematic base-triple substitutions established that in-vitro thermodynamic stability was necessary but not sufficient for reporter stabilization in cells. METTL16 binding, detected with biochemical and in-cell association assays, supplied protein-sensitive evidence consistent with formation of the MALAT1 triplex in cells. Small-molecule studies further demonstrated that chemically similar triplexes can expose different pockets and respond differently to ligands. These experiments provide unusually convergent single-site evidence: atomic structure or structural modeling, mutational energetics, RNA accumulation, nuclease protection, protein recognition, and ligand response.

Figure 53.7. How ENE and MALAT1 RNA Triple Helices Protect RNA Ends

Figure 53.7. How ENE and MALAT1 RNA Triple Helices Protect RNA Ends. An original molecular schematic should compare the KSHV PAN ENE and MALAT1 3′ element. Each panel should identify the Watson-Crick duplex, U-rich Hoogsteen strand, A-rich strand, stacked U•A-U triples, flanking helices, and the RNA end protected from deadenylation or exonucleolysis. For the PAN ENE, show that the clamp can select an internal compatible register within a longer poly(A) tail and arrest shortening before a deadenylase reaches the RNA body; do not imply that the triplex must occupy the terminal adenosines. A causal inset should trace sequence and ion context to triplex assembly, reduced exonuclease access or progression, and increased RNA stability while noting that in-vitro melting stability is not sufficient for cellular stabilization.

Triplex stability remains conditional. Monovalent salt screens phosphate repulsion, magnesium can assist compaction, and pH changes protonation-sensitive triples. The length and sequence of the Hoogsteen tract, interruptions, terminal stacking, flanking duplexes, and the register of the third strand alter both equilibrium and kinetics. A short purified construct can overestimate accessibility if the same sequence is paired elsewhere in a full-length transcript. Conversely, a protein or adjacent helix can stabilize a triplex that appears marginal in isolation. Single-molecule FRET, NMR, crystallography, mutational rescue, thermal melting, nuclease-resistance assays, and RNA half-life measurements are complementary rather than interchangeable.

There is not yet a generally accepted transcriptome-wide occupancy assay for RNA-only triple helices comparable to rG4-seq. Covariance models and sequence searches can nominate ENE-like elements, and comparative work found many such candidates in transposable-element and viral RNAs, with reporter validation for representatives. Those results support family expansion, not automatic folding of every computational hit. For a new locus, the strongest workflow remains targeted: define the three strands, perturb Watson-Crick and Hoogsteen contacts independently, rescue compatible triples where possible, measure folding and stability under physiological ion conditions, and connect the structure to an RNA-fate phenotype.

Z-RNA: a conditional left-handed duplex

Ordinary double-stranded RNA adopts a right-handed A-form helix. Z-RNA is a left-handed duplex with a zig-zag phosphate backbone and a dinucleotide repeat in conformation. In the low-salt ADAR1 Zα-bound crystal structure, guanosines adopt syn glycosidic orientations and generally C3′-endo sugar puckers, whereas cytidines adopt anti orientations and C2′-endo puckers. Alternating purine-pyrimidine sequences, especially alternating C and G, are favorable model substrates, but sequence alone is not determinative. Hydration, cations, chemical modification, duplex length, junction cost, mechanical or topological context, and binding proteins determine the free-energy penalty for an A-to-Z transition.

Z-RNA is normally a higher-free-energy state than A-RNA. High salt and other nonphysiological conditions can drive model duplexes toward Z-like spectra in vitro, but a cellular mechanism cannot be inferred from that transition alone. Zα domains provide a biologically relevant route to stabilization. The Zα domain of ADAR1 binds and stabilizes Z-RNA in biochemical experiments, and a crystal structure of ADAR1 Zα bound to a short alternating-CG RNA revealed direct recognition of the left-handed backbone and its distinctive solvent organization. Binding can select a rare pre-existing Z-like state, lower the transition barrier, or stabilize the product; equilibrium binding alone does not distinguish those kinetic models.

ADAR1 p150 and Z-DNA-binding protein 1 (ZBP1) connect Z-RNA recognition to cell biology. Mouse genetics show that intact ZBP1 Zα domains are required for inflammatory pathology in several RIPK1-, FADD-, or caspase-8-perturbed settings, and that impairing the Zα function of ADAR1 permits ZBP1-dependent pathology; complementary studies link probe-detected endogenous Z-form RNA accumulation to ZBP1 activation and necroptosis. Endogenous retroelement-derived duplexes and interferon-stimulated transcripts are candidate sources, but these genetic and affinity-capture results do not by themselves assign nucleotide-resolved Z-RNA occupancy. This chapter owns the A-to-Z structural transition and evidence limits. The downstream sensing circuitry belongs to Chapter 108, while ADAR1 isoforms, editing, and self-RNA discrimination belong to Chapter 109.

Z-RNA detection is technically difficult because the measurement reagent can stabilize the conformation and because many reagents also recognize Z-DNA. Circular dichroism and Raman spectroscopy can follow an A-to-Z transition in purified duplexes; NMR and crystallography can define local geometry; Zα domains or antibodies such as Z22 can enrich Z-form nucleic acids from fixed or lysed material. Nuclease controls are essential to distinguish RNA from DNA, and input-normalized sequencing is needed to separate enrichment from abundance. Crosslinking, fixation, protein expression, and probe affinity can change the state being measured. A Z22 or Zα enrichment peak should therefore be annotated as probe-dependent evidence for a Z-form substrate, not as a direct quantitative occupancy map at nucleotide resolution.

Figure 53.8. The Conditional A-RNA-to-Z-RNA Transition

Figure 53.8. The Conditional A-RNA-to-Z-RNA Transition. An original structural comparison should place right-handed A-RNA beside left-handed Z-RNA, labeling handedness, the smooth versus zig-zag backbone, guanosine syn/C3′-endo and cytidine anti/C2′-endo geometry in the low-salt ADAR1 Zα-bound structure, and the energetic cost of the junction. Arrows should show conditions that can favor Z-RNA: alternating purine-pyrimidine sequence, hydration and cations, modifications, stress, and Zα-domain binding. A measurement inset should warn that Zα and Z22 capture can stabilize the state and can require nuclease controls to distinguish RNA from DNA.

One evidence ladder for three different conformers

For all three motif families, sequence prediction is the bottom of an evidence ladder. A guanine-rich pattern predicts rG4 potential; a U-rich internal loop beside an A-rich tract predicts triplex potential; an alternating purine-pyrimidine duplex predicts Z-RNA propensity. The next level is purified-RNA biophysics under explicitly stated ion, pH, temperature, and concentration conditions. Site-specific structural evidence, compensatory mutation, and kinetic measurements then test geometry and interconversion. In-cell probing or conformation-selective capture tests cellular accessibility, but must be interpreted with probe perturbation and abundance controls. Protein genetics, structure-sensitive rescue, and phenotype-linked assays connect conformer occupancy to mechanism.

The negative result is often the most informative result. The strong in-vitro rG4 potential of thousands of human RNA regions coexists with widespread unfolding in cells. A stable MALAT1 base triple in vitro can fail to stabilize a reporter in cells. A Z-prone repeat can remain A-form unless an ion environment, modification, stress, or Zα protein pays the transition cost. These are not contradictions. They show that cellular RNA structure is an actively maintained ensemble rather than the inevitable output of a motif regex.

The same separation prevents cross-chapter overreach. General free-energy, ion, and ensemble concepts are developed in Chapter 4. DNA G4 formation, R-loops, transcription-generated topology, and DNA-RNA crosstalk are treated in Chapter 97; the presence of an rG4 in a transcript should not be used as proof of a DNA G4 or R-loop at its gene. Z-RNA recognition by innate sensors is mechanistically developed in Chapter 108, and ADAR1-dependent editing and self-RNA discrimination in Chapter 109. Within this chapter, the primary question remains structural: what conformer is physically defined, what fraction of the RNA occupies it, which experiment supports that statement, and what alternative structures or assay artifacts remain plausible?

Table 53.3. Evidence Matrix for rG4, RNA Triplex, and Z-RNA Claims. The three conformer families share an evidence ladder but require different structural diagnostics and artifact controls.

Conformer Minimal physical definition Sequence-level nomination Strong single-site evidence Transcriptome-scale evidence Dominant artifact or overreach
RNA G-quadruplex At least two stacked G-quartets around a monovalent-ion channel Four G-rich tracts, including bulged or noncanonical variants Ion-dependent spectroscopy, NMR or atomic structure, structure-selective chemistry, guanine disruption, orthogonal functional rescue rG4-seq for folding potential; in-cell chemical probing for accessibility; G4RP-seq for fragment-scale probe-capturable enrichment Treating a motif count, potassium-folded extract, or ligand-captured RNA as constitutive in-cell occupancy
RNA triple helix Third RNA strand repeatedly recognizes a duplex, commonly by major-groove Hoogsteen triples U-rich tract and compatible A-rich/duplex segments; ENE covariance model Atomic structure or NMR, independent disruption of Watson-Crick and Hoogsteen contacts, triplex rescue, nuclease protection and RNA-fate assay Comparative searches can nominate ENE-like families; no universal RNA-only triplex occupancy assay Equating one base triple, a U-rich/A-rich juxtaposition, or an RNA-DNA triplex prediction with an extended RNA triplex
Z-RNA Left-handed duplex with zig-zag backbone and alternating syn/anti and sugar-pucker states Alternating purine-pyrimidine duplex and context-dependent Z propensity CD/Raman transition, NMR or crystal structure, Zα binding with structural controls, conformation-sensitive mutagenesis Z-form capture or RIP-seq under perturbation; endogenous site maps remain probe- and state-dependent Probe stabilization, RNA/DNA cross-reactivity, and inferring nucleotide-resolved occupancy from pathway genetics

Box 53.3. Five Questions Before Calling a Cellular Noncanonical Conformer

  • What is the exact physical structure being claimed: stacked G-quartets, a repeated third-strand triplex, or a left-handed duplex?
  • Does the experiment measure sequence potential, purified-RNA folding, probe-dependent enrichment, or occupancy in intact cells?
  • Which ion, pH, temperature, concentration, transcript isoform, compartment, and binding partners were present?
  • Could the reagent stabilize the conformer, cross-react with DNA or another RNA structure, or preferentially capture abundant transcripts?
  • Does an orthogonal mutation or perturbation connect the conformer to a specific RNA-fate or cellular mechanism without creating a new competing structure?

53.5. Metal-ion, ligand, and protein-stabilized folds

RNA is a polyanion. Every nucleotide contributes a negatively charged phosphate, so folding compact RNA structures requires electrostatic compensation. Monovalent ions contribute general charge screening, and divalent ions such as magnesium can stabilize compact folds. Some magnesium ions act diffusely; others occupy specific pockets and contact phosphate oxygens, bases, or water-mediated ligand networks. The distinction matters. A general magnesium requirement does not prove a site-specific metal-ion mechanism, while a resolved inner-sphere or outer-sphere metal site can indicate a precise structural role.

Ligands can convert an RNA ensemble into a functional architecture. In a riboswitch, the aptamer domain binds a metabolite or ion, and the ligand-bound architecture changes the expression platform controlling transcription, translation, splicing, or RNA stability. The ligand does not merely occupy a pocket; it shifts the probability of alternative RNA conformations. This explains why ligand affinity, folding kinetics, transcription speed, and competing structures all matter. RNA-ligand docking reviews emphasize both progress and persistent challenges in predicting small-molecule binding to RNA.

RNA-targeted ligand discovery is a useful boundary case. A small molecule may bind an RNA motif in vitro, but cellular function requires occupancy at achievable concentration, selectivity against other RNAs and proteins, access to the right compartment, and a measurable effect on RNA fate. Docking can prioritize hypotheses, but RNA flexibility, ion effects, hydration, and alternative conformations make RNA-ligand prediction difficult. Motif annotation can identify pockets, yet a pocket is not a validated drug target without binding, specificity, and functional evidence.

Proteins stabilize many cellular RNA folds. An RNA-binding protein can recognize a preformed loop, bind a single-stranded region and induce structure, bridge two distant RNA segments, stabilize a kink-turn, shield a labile helix, or remodel a fold with ATP-dependent helicase activity. Sasse and colleagues reviewed motif models for RNA-binding proteins, emphasizing that protein recognition often involves combinations of sequence, structure, and context rather than a single motif code. Chapter 56 treats this topic in depth.

The causal sequence for a protein-stabilized fold should be stated explicitly. Does RNA fold first and recruit the protein? Does the protein bind an unfolded segment and induce the motif? Does co-transcriptional protein binding prevent an alternative fold? Does the protein remain part of the final structure or act as an assembly factor? Different answers imply different experimental tests. Time-resolved probing, protein depletion, reconstitution, mutational rescue, and high-resolution RNP structures answer different parts of this question.

This subsection is intentionally conservative because the provided references for metal-ion and protein-stabilized folds are not sufficient for a full mechanistic literature base. The general principles above are established in RNA structural biology, and the local bibliography supports selected ligand, motif, and RNP examples. Final reference expansion should add more dedicated sources on magnesium-RNA interactions, ribozyme metal-ion catalysis, riboswitch ligand binding, RNP assembly, and helicase remodeling; until then, claims about specific ions or proteins should remain tied to cited local sources.

53.6. Motif conservation and comparative evidence

Comparative evidence asks whether evolution preserved a structure when sequences changed. If two aligned positions form a base pair, a mutation at one side may be compensated by a mutation on the other side. A G-C pair can become an A-U pair, preserving pairing. A G-U pair can be exchanged with another geometrically compatible pair. Multiple independent compensatory changes across a phylogeny provide strong evidence that the pair is functionally or structurally constrained.

Rivas’s review of evolutionary conservation of RNA sequence and structure is central for interpreting these signals. The main point is that RNA structure conservation is not identical to sequence conservation. Some RNAs preserve primary sequence because proteins or other RNAs read exact bases. Others preserve pairing while allowing sequence turnover. Still others preserve a higher-order architecture, such as a junction or tertiary motif, with only modest primary-sequence similarity. Comparative analysis can reveal all three regimes if alignments and models are good.

Covariation has limitations. It requires enough homologous sequences with enough variation. Closely related sequences may show conservation but not compensatory change. Distant sequences may be hard to align. Shared ancestry can make correlations appear stronger than they are if phylogeny is not modeled. Sequencing errors, paralogy, annotation errors, and selection on overlapping coding regions can mislead inference. In viral RNAs, a paired region may also encode protein sequence, immune evasion, or replication signals, complicating interpretation.

Isostericity extends comparative logic beyond canonical pairs. Lescoute and colleagues showed how recurrent structural motifs and isostericity matrices can guide sequence alignments. In a noncanonical base pair, a substitution may preserve the spatial arrangement even if it does not preserve Watson-Crick pairing. A model that recognizes only A-U, G-C, and G-U pairs may miss this conservation. This is especially important for tertiary motifs, where geometry can be conserved through base-edge substitutions.

Comparative evidence also supports long-range contacts. A distant interaction is more convincing when both sides covary together across homologs and when mutational disruption and rescue support the same pairing. For pseudoknots, covariation can be decisive because computational predictions alone often generate many alternatives. For tertiary motifs, recurring geometry across solved structures and isosteric sequence variation provide evidence of motif conservation.

Functional conservation is not guaranteed by structural conservation. A motif may be conserved because it maintains RNA stability, because it binds a protein, because it controls translation, or because it constrains an overlapping coding sequence. Conversely, a function can be conserved by different structural solutions. Comparative evidence should therefore be integrated with biochemical, genetic, and cellular experiments. The claim “this motif is conserved” is not the same as “this motif has the same function in every organism.”

53.7. Structural databases and annotation

Structural databases turn individual RNA structures into reusable knowledge. A database may store atomic coordinates, global fold classifications, motif annotations, aptamer models, base-pair classes, ligand contacts, sequence alignments, or metadata about experimental method and resolution. The value of a structural database is not only storage. It provides consistent identifiers, searchable annotations, and a way to compare motifs across RNA classes.

The Protein Data Bank is the general archive for experimentally determined macromolecular structures, including RNA-containing structures. RNA-focused resources build additional layers on top of coordinate archives. They may classify RNA global structures, annotate base-pairing networks, catalog tertiary motifs, or provide specialized aptamer models. Hong and colleagues introduced RR3DD as a global structure-based RNA three-dimensional classification database. Sato and colleagues described RNAapt3D as a database for RNA aptamer three-dimensional structural modeling. These sources illustrate two database scales: broad structural classification and specialized functional class modeling.

Motif annotation often uses graph representations. In a graph, nucleotides, base pairs, backbone links, and tertiary interactions can become nodes and edges. This representation supports motif search because a recurrent structural pattern can be found as a subgraph rather than by exact sequence. Emrizal and colleagues reviewed graph theoretical methods and workflows for RNA tertiary base motif search and annotation. Graph methods are useful because they transform visual structural patterns into explicit relationships that can be searched, compared, and audited across structures.

Annotation requires controlled vocabulary. A “hairpin,” “kissing loop,” “A-minor interaction,” “kink-turn,” “ribose zipper,” and “pseudoknot” should not be used as loose metaphors. Each should have a definition, boundary conditions, and evidence status. Some motifs have strict geometric criteria, while others are family names for related arrangements. A database may annotate a motif differently depending on cutoff distances, base-pair classification scheme, resolution, or whether proteins and ions are included.

Database bias is unavoidable. Solved structures overrepresent stable, abundant, crystallizable, cryo-EM-friendly, or heavily studied RNAs. Ribosomes, riboswitches, tRNAs, ribozymes, aptamers, viral elements, and engineered RNAs are much better represented than many long cellular RNAs. Low-resolution models may not support fine motif annotation. In vitro structures may lack cellular proteins, ligands, crowding, co-transcriptional folding, or modifications. Database users should treat absence of a motif as absence from current annotation, not absence from biology.

Computational annotation is improving. Karan and Rivas’s motif-aware folding framework points toward integrated prediction in which evolutionary information, secondary structure, and three-dimensional motif hypotheses are considered together. Such approaches may eventually fill gaps for RNAs without solved structures. The risk is overconfidence: predicted motifs need confidence scores, input evidence, and validation status. A predicted A-minor contact and a cryo-EM-observed A-minor contact should not be merged without provenance.

A reusable structural annotation should include a stable identifier, motif type, RNA class, organism, sequence interval, structural evidence, functional evidence, source references, database cross-references, and uncertainty. Such an annotation can say not only “this RNA has a pseudoknot” but “a pseudoknot is supported by covariation and mutational rescue in this viral untranslated region, with no high-resolution structure yet available.” That is the level of precision needed for scientific interpretation.

53.8. Functional interpretation across RNA classes

Structural motifs have different meanings in different RNA classes. In rRNA, motifs help build the ribosome’s scaffold, decoding center, peptidyl-transferase center, intersubunit bridges, and protein-binding surfaces. Many rRNA motifs are deeply conserved and embedded in a giant RNP. A local mutation can have global effects because the ribosome is a coupled machine. In tRNA, secondary and tertiary motifs define the L-shaped fold, anticodon presentation, aminoacylation identity, modification sites, and recognition by translation factors.

In riboswitches, motifs connect ligand recognition to gene control. The aptamer domain uses stems, junctions, tertiary contacts, metal ions, and ligand contacts to form a binding pocket. The expression platform converts the ligand-bound or ligand-free folding state into transcription termination, translation initiation, splicing, or RNA stability. The structural motif is therefore a switch component, not the whole switch. Ligand affinity, folding speed, transcription timing, and competing structures determine regulatory output.

In mRNAs, structural motifs can affect translation initiation, elongation, frameshifting, localization, decay, splicing, editing, and protein binding. A hairpin in a 5′ untranslated region may block scanning or recruit a protein. A pseudoknot can stimulate frameshifting. A 3′ untranslated-region structure can bind regulatory proteins or miRNAs. However, mRNAs are dynamic substrates for ribosomes, helicases, decay factors, and RNA-binding proteins. A motif predicted from naked RNA may not persist in the translating or protein-bound cell state.

In lncRNAs and other long noncoding RNAs, architecture is often harder to interpret. Some lncRNAs use local domains that bind proteins or nucleic acids, while the full transcript may be flexible, heterogeneous, or cell-state dependent. A structural motif in a lncRNA can be functionally important, but many lncRNA structure-function claims need careful controls for transcription, chromatin context, RNA abundance, localization, and protein binding. Putnam and colleagues’ discussion of RNA granules as functional compartments or incidental condensates is relevant here because RNA architecture can influence condensate participation, but condensate localization alone does not prove a specific structured mechanism.

Viral RNAs frequently use compact and multifunctional motifs. A viral genome may encode proteins, serve as template, recruit host and viral proteins, evade innate immunity, package into particles, and control translation. Flynn and colleagues mapped SARS-CoV-2 RNA-host protein interactions, showing how viral RNA function is inseparable from host RNP context. A structural motif in a viral RNA may simultaneously affect replication, translation, immune sensing, and packaging. Functional interpretation must separate these outputs rather than assigning one purpose too quickly.

Engineered RNAs use motifs as design components. Aptamers, guide RNAs, synthetic riboswitches, RNA scaffolds, and therapeutic RNAs can exploit stable stems, loop-receptor contacts, ligand pockets, and protein-binding motifs. Design makes the modularity assumption explicit: a motif is chosen because it is expected to behave similarly in a new context. Failures are instructive. A motif may misfold, become inaccessible, trigger innate immune sensing, bind unintended proteins, or fail under cellular ion conditions. Design therefore tests how transferable a structural module really is.

Recent Consensus

Recent consensus treats RNA structure as hierarchical, modular, and context-dependent. Local secondary-structure motifs provide an essential scaffold, but recurrent tertiary interactions and long-range contacts are required to explain compact functional folds. Motif catalogs, graph methods, and structure databases have matured enough to support systematic annotation rather than one-off visual description.

The field also agrees that evolutionary evidence is central. Conserved base pairs, compensatory mutations, isosteric substitutions, and conserved tertiary motif geometry can reveal structural constraint even when primary sequence conservation is weak. Comparative analysis is strongest when aligned with experimental probing, mutational rescue, and high-resolution structure.

A third consensus is that cellular RNA folds are conditional. Ions, ligands, proteins, transcriptional timing, RNA modifications, and molecular crowding influence which architecture forms. Motifs observed in purified RNA can be biologically meaningful, but they should be connected to cellular evidence before being treated as in vivo mechanisms. This is especially important for long mRNAs, lncRNAs, viral RNAs, and engineered RNAs.

For rG4s, triplexes, and Z-RNA, current consensus is best stated as an evidence rule rather than a single prevalence estimate. All three conformers are physically established, and selected cellular examples have strong mechanistic support. Their transcriptome-wide occupancy is not implied by motif frequency. rG4 studies directly demonstrate widespread in-vitro potential alongside active cellular unfolding, while triplex and Z-RNA evidence is strongest at defined loci or under perturbations that stabilize or reveal the conformer.

Open Questions, Controversies, Deprecated Models, and Common Misconceptions

Open questions:

  • How much of the transcriptome forms stable, specific tertiary architecture in cells? Many RNAs contain local structures and protein-bound regions, but stable global folds are easier to establish for rRNAs, tRNAs, riboswitches, ribozymes, snRNAs, viral elements, and engineered aptamers than for many long mRNAs and lncRNAs. The absence of a solved structure does not mean absence of structure, but the presence of predicted structure does not mean a stable cellular fold.
  • Which endogenous rG4s are transiently occupied without a stabilizing probe, which RNA-only triplex families extend beyond ENE-like elements, and which native duplex segments cross into Z-RNA at nucleotide resolution? Existing methods answer parts of these questions, but probe-induced stabilization, transcript abundance, isoform choice, and limited temporal resolution prevent a universal occupancy atlas.

Controversies:

  • A second controversy concerns functional assignment. Motif discovery often runs ahead of functional testing. A conserved hairpin, pseudoknot, or tertiary-like contact may matter, but it may affect RNA stability, processing, localization, translation, immune sensing, or protein binding rather than the first function imagined by the annotator. Functional claims should specify the tested phenotype and the evidence.
  • A third unsettled area is RNA-ligand targeting. RNA contains many pockets and dynamic conformations, and small molecules can bind RNA. The challenge is selectivity, cellular occupancy, function, and safety. Docking and motif annotation are useful for hypothesis generation, but they do not replace biochemical and cellular validation.
  • The cellular prevalence of rG4s remains context-dependent rather than reducible to “present” or “absent.” Chemical probing supports pervasive eukaryotic unfolding, whereas conformation-selective capture and perturbation studies support transient or regulated folding. These results measure different states and can coexist; disagreement arises when a folding-potential map, ligand-stabilized state, or enrichment signal is treated as unperturbed occupancy.
  • Z-RNA biology has strong protein-genetic support but incomplete site-level structural resolution. Zα-domain dependence links a left-handed-RNA recognition system to ADAR1 and ZBP1 phenotypes, yet capture reagents can stabilize Z-form nucleic acids and may recognize both RNA and DNA. Claims about a particular endogenous Z-RNA site therefore require stricter evidence than claims about pathway dependence.

Common misconceptions:

  • “Every hairpin is functional.” Hairpins can arise from ordinary RNA chemistry; function requires conservation, perturbation, binding, or output evidence.
  • “Every conserved sequence segment implies a conserved structure.” Sequence can be conserved for coding, protein-binding, processing, or regulatory reasons that do not require conserved base pairing.
  • “Every predicted pseudoknot causes frameshifting.” Frameshifting requires evidence for ribosome behavior and reading-frame output, not only a plausible RNA structure.
  • “Secondary structure is the whole RNA structure.” Secondary structure captures base-pairing patterns but not tertiary contacts, dynamics, ligand binding, protein assembly, or cellular context.
  • “A motif can be transplanted without context effects.” Motif behavior depends on surrounding sequence, folding path, expression context, proteins, ligands, and cellular state.
  • “Protein-bound RNA structures reveal the naked RNA fold.” Protein binding can stabilize, remodel, mask, or select RNA conformations, so the RNP state and naked RNA state must be interpreted separately.
  • “Every guanine-rich RNA folds into an rG4 in cells.” G-richness specifies potential; competing helices, helicases, ribosomes, proteins, and ionic context can keep the sequence unfolded or in another structure.
  • “A transcriptome-wide rG4 map measures one universal quantity.” Reverse-transcriptase stalling, in-cell chemical probing, and ligand-based capture measure different combinations of potential, accessibility, enrichment, and perturbation response.
  • “Every U-rich region beside an A-rich tract is an RNA triple helix.” A triplex requires a compatible duplex, repeated Hoogsteen contacts, correct register, and context-specific stabilization; targeted mutation and structural evidence are needed.
  • “Z-RNA is simply an unusual RNA sequence.” Z-RNA is a physical left-handed duplex state. Sequence affects propensity, but protein binding, hydration, ions, modifications, and transition costs determine occupancy.