Chapter 3. RNA Thermodynamics, Kinetics, Folding Landscapes, Energy Models, and Ensemble Concepts

Scope Note

RNA is a chemical polymer, but an RNA molecule is not usefully understood as a static string of letters or as one fixed drawing. RNA molecules sample conformations, exchange base pairs, bind ions and proteins, respond to temperature, and sometimes become trapped in structures that are not the most stable structures available to the full sequence. This chapter introduces the physical language needed to reason about those behaviors: free energy, enthalpy, entropy, temperature, salt, nearest-neighbor models, minimum-free-energy prediction, partition functions, base-pair probabilities, ensembles, folding landscapes, kinetic barriers, metastable states, and co-transcriptional folding. The chapter is conceptual and interpretive. Detailed algorithms for secondary-structure prediction, partition functions, probabilistic folding, probing-constrained modeling, tertiary modeling, and simulation are developed later in Chapters 60 through 65.

Executive Summary

RNA folding is best treated as a population and pathway problem. A single RNA sequence can populate multiple conformations, and the abundance of each conformation depends on the free-energy differences among states, the barriers between states, the ionic environment, temperature, pH, RNA concentration, sequence boundaries, chemical modifications, ligands, proteins, transcription history, translation, compartment, and active remodeling. The same nucleotide sequence can therefore behave differently in a purified buffer, in a bacterial transcription complex, inside a translating eukaryotic cytoplasm, in a viral particle, or in a lipid nanoparticle formulation. This ensemble view is a central principle of modern RNA biophysics.

Free energy is the thermodynamic quantity most often used to compare RNA states. A lower Gibbs free energy means a state is more favored at equilibrium under specified conditions. Free energy is not a single intrinsic property of a sequence. It is a conditional statement about a sequence, construct, buffer, temperature, salt, divalent ions, concentration, chemical composition, and model or measurement method. Enthalpic interactions such as base stacking, hydrogen bonding, ligand binding, and metal-ion coordination favor particular arrangements. Entropic effects oppose or favor folding depending on chain restriction, released ions, water reorganization, and binding-partner freedom. Reusable claims about RNA stability must therefore preserve the relevant conditions.

Nearest-neighbor models make secondary-structure prediction practical by estimating the free energy of an RNA structure from local sequence and loop features. The approach is powerful because it connects sequence to a physical hypothesis, but it rests on assumptions about additivity, parameter conditions, and the omission or simplification of pseudoknots, tertiary contacts, proteins, ligands, modifications, and cellular environments. The minimum-free-energy structure, or MFE structure, is the lowest-scoring secondary structure under a specified model. It is a useful hypothesis, not direct evidence that the RNA molecule adopts one exclusive structure.

Partition-function and ensemble methods reduce the false certainty created by single-structure drawings. A partition function sums modeled structures weighted by their Boltzmann factors; base-pair probabilities estimate how often each pair appears across the modeled ensemble. These calculations expose uncertain helices, competing alternatives, and regions in which the model expects a dominant fold. Model-derived probabilities are not direct experimental occupancies unless the model has been validated under matching conditions.

Kinetics explains why the most stable equilibrium state is not always the state that matters biologically. RNA folds while it is synthesized, processed, transported, translated, bound, and remodeled. Early structures can form before later sequence segments exist. Polymerase pausing can give a nascent RNA time to form a hairpin or ligand-binding fold. A metastable state can persist if a high barrier slows escape to another state. Co-transcriptional folding and kinetic traps are therefore not exceptions to RNA biology; they are ordinary consequences of polymer growth and rugged folding landscapes.

Cellular RNA folding is not merely dilute-solution folding with more parameters. RNA-binding proteins, ribosomes, helicases, metabolites, compartments, modifications, molecular crowding, and local ion composition can reshape both thermodynamic populations and kinetic transition rates. The appropriate scientific posture is neither to trust folding predictions blindly nor to dismiss them because cells are complex. Folding models are disciplined hypotheses. Experiments decide which predicted states are populated, which states are functional, and which missing cellular factors explain model disagreement.

Concept Inventory

  • Free energy: a thermodynamic quantity used to compare the relative stability of states under specified conditions. In RNA folding, the relevant free energy is usually Gibbs free energy, often written as Delta G. A negative folding Delta G means that the folded state is favored relative to a defined reference state, not that every molecule is folded.
  • Thermodynamic condition: the set of physical and chemical circumstances under which a stability claim is made. Temperature, monovalent salt, Mg2+, pH, RNA concentration, construct boundaries, modifications, proteins, ligands, and model assumptions can all change interpretation. “Stable RNA hairpin” is not a complete statement if the claim will be reused quantitatively.
  • Enthalpy: the heat-related component of free energy. Favorable enthalpic contributions in RNA include base stacking, hydrogen bonding, metal-ion coordination, ligand contacts, and protein contacts. Unfavorable enthalpic terms can arise from disrupted interactions or charge repulsion.
  • Entropy: the number and distribution of microscopic arrangements available to a system. RNA folding usually restricts chain motion, which is entropically costly, but folding can also release ions or water and can change the entropy of binding partners. Entropy is therefore not simply “disorder” in a casual sense; it is a quantitative contribution to population balance.
  • Nearest-neighbor model: estimates secondary-structure stability by adding energetic contributions from adjacent base pairs and loop features. The approach is often called an NN model. When common RNA thermodynamic parameter families are meant, the phrase Turner-style parameters is often used, although specific parameter sets and software implementations should be named when needed.
  • Minimum-free-energy structure: the lowest-energy structure predicted under a specified model and constraints. The abbreviation MFE should be read as “best under this model,” not as “observed molecular structure.”
  • Partition function: a sum over modeled states weighted by their Boltzmann factors. In RNA secondary-structure prediction, it supports calculation of ensemble free energy and base-pair probabilities. The partition function counts only the states allowed by the model and the chosen constraints.
  • Base-pair probability: a model-derived estimate that two positions are paired within the modeled ensemble. It is not direct evidence that the pair is formed in a cell unless supported by matching experimental evidence.
  • Ensemble: the collection of conformational states populated by an RNA under specified conditions. A folding landscape represents those states, their relative energies, and the barriers connecting them. Low-energy wells represent populated states; barriers influence how quickly molecules move among states.
  • Kinetic trap: a metastable conformation that persists because escape is slow, even if another state is thermodynamically more favorable. A metastable state can be functional, pathological, or experimentally misleading depending on context.
  • Co-transcriptional folding: folding that occurs while an RNA is being synthesized, before the entire sequence is available. Nascent RNA folding is a closely related phrase. The order and timing of sequence emergence can affect which base pairs form first.
  • Ion atmosphere: the diffuse and sometimes site-associated distribution of counterions around the negatively charged RNA backbone. Diffuse electrostatic screening and specific metal-ion binding are related but distinct physical phenomena.
  • Cellular folding environment: proteins, ribosomes, helicases, metabolites, modifications, compartments, crowding, ion composition, RNA decay machinery, and local concentration. These factors can alter RNA folding without changing the RNA sequence.

What to Know Before Reading This Chapter

The prerequisite chemical idea is simple: RNA is a single-stranded polymer with a negatively charged phosphate backbone and bases capable of stacking and pairing. Chapter 2 explains the chemical groups in detail. This chapter asks how those chemical features create populations of structures.

The most useful running example is a short RNA that can form a hairpin. A hairpin contains a stem, in which bases pair with a complementary region of the same strand, and a loop, in which unpaired nucleotides connect the two sides of the stem. At low temperature in favorable salt, many molecules may be folded into the hairpin. At higher temperature, more molecules may be unfolded. If a protein binds the loop, the folded state can become more populated. If a later transcribed downstream segment can pair with one side of the stem, the early hairpin may be replaced by an alternative helix or may persist as a kinetic trap.

The second running example is a riboswitch. A riboswitch is a structured RNA regulatory element that changes gene expression in response to a small molecule or ion. Many riboswitches contain an aptamer domain that binds the ligand and an expression platform that changes transcription, translation, splicing, or RNA stability. Riboswitch behavior is a natural place to see thermodynamics and kinetics together: ligand binding changes relative stability, while transcription timing and folding barriers influence whether the ligand has an opportunity to control the regulatory output.

The third running example is a viral frameshifting element. Some viral RNAs contain structured regions such as pseudoknots or stem-loops that influence ribosome movement and reading-frame choice. The biological effect depends not only on whether a structure is predicted, but also on how stable it is, how it responds to mechanical stress from the ribosome, how quickly it refolds, and whether alternative states are populated.

Several interpretive cautions recur throughout the chapter. “Stable” means stable under conditions. “Predicted” means produced by a model. “MFE” means lowest under a model, not necessarily present in cells. “Probability” often means model-derived probability, not measured occupancy. “Structure probing” reports chemical reactivity or accessibility, not a structure by itself. Keeping these distinctions clear prevents many common errors in RNA biology.

3.1. Thermodynamic variables: free energy, entropy, enthalpy, salt, and temperature

Figure 3.1. RNA Folding Landscape and Ensemble View

Figure 3.1. RNA Folding Landscape and Ensemble View. A single RNA sequence can adopt multiple conformations separated by energy barriers, with each conformation’s population determined by its free energy relative to temperature, ion conditions, and the presence of ligands or proteins. The folding landscape depicts low-energy wells as populated states and barriers as obstacles to interconversion, illustrating why population shifts rather than single rigid structures characterize RNA biology. A ligand or RNA-binding protein can stabilize one well and depopulate others, demonstrating that RNA function often operates through ensemble redistribution rather than all-or-none structural switching.

Thermodynamics describes how physical conditions determine the relative populations of molecular states at equilibrium. Equilibrium does not mean that molecules stop moving. It means that the rates of transitions among states balance so that the overall population distribution is stable over time. For RNA, the states may include an unfolded chain, a hairpin, a competing helix, a compact tertiary fold, a ligand-bound conformation, or an RNA-protein complex.

Table 3.1. Physical Variables and Interpretation Risks. A summary of physical and chemical variables that affect RNA stability and folding, with notes on model dependence and the overgeneralizations each variable invites.

Variable Effect on RNA Measurement or model dependence Common overgeneralization Related chapters
Temperature Increases thermal motion; destabilizes ordered structures at higher T Melting curves assume two-state behavior; Tm depends on buffer and concentration “Stable at 37 °C” omits ionic conditions, construct, and method Chapter 2, Chapter 59
Monovalent salt Screens phosphate charge; stabilizes helices Measured under defined [K⁺] or [Na⁺]; models often use 1 M NaCl convention “Salt stabilizes RNA” ignores ion species, concentration, and competing ions Chapter 59
Mg²⁺ Supports compact folds; can act diffusely or at specific sites Diffuse vs. site-specific effects require different models; concentration matters “Mg²⁺ stabilizes RNA” conflates charge screening with site-specific binding Chapter 59
pH Affects protonation of A and C residues; alters hydrogen bonding Measured at specific pH; parameters often assume pH 7.0 Ignoring pH misestimates stability of modified or catalytic RNAs Chapter 2, Chapter 4
RNA concentration Affects intermolecular duplex equilibrium and aggregation Concentration must be stated for intermolecular reactions Hairpin stability transferred to duplex context without concentration adjustment Chapter 2, Chapter 59
Crowding Excluded volume can favor compact states In-cell and in vitro conditions differ; direct modeling is difficult Assuming dilute-buffer predictions equal in-cell behavior Chapter 3, Chapter 6
Modification Alters base-pairing, stacking, hydration, and protein recognition Standard parameters assume unmodified RNA; modification effects are context-dependent Assuming modification effect is negligible without evidence Chapter 57
Protein binding Can stabilize, block, or remodel structures Requires binding assays; most folding models omit protein contributions Assuming naked-RNA fold equals the protein-bound fold Chapter 54, Chapter 56
Ligand binding Stabilizes one conformation; shifts ensemble populations Requires affinity and structural evidence; effect depends on ligand concentration Assuming the tightest-binding conformation is always the functional state Chapter 49
Transcription rate Affects cotranscriptional fold selection and kinetic trap formation Requires cotranscriptional assays or kinetic simulation; rate is polymerase- and context-dependent Assuming full-length equilibrium fold equals the nascent functional fold Chapter 24, Chapter 25

Gibbs free energy is the central thermodynamic quantity for comparing these states at constant temperature and pressure. If folding from an unfolded reference state to a hairpin has a negative Delta G, the hairpin is favored at equilibrium. If the Delta G is positive, the unfolded reference state is favored. The size of the difference matters. A very large favorable free-energy difference can make one state dominate the population. A small difference can leave substantial populations of multiple states. This is why an RNA with two nearly equal secondary structures can behave as a switch rather than as a single rigid object.

Table 3.2. RNA Folding Model Types and Evidence Requirements. An overview of major computational and experimental approaches used to model or measure RNA folding, with their outputs, strengths, key assumptions, and validation requirements.

Model or evidence type Output Strength Assumption Validation need Related chapters
MFE prediction Single lowest-energy secondary structure Interpretable; computationally fast; widely available Local additivity; naked unmodified RNA; equilibrium condition Mutagenesis, chemical probing, functional assay Chapter 60
Partition function Ensemble free energy; base-pair probabilities Exposes alternative states and structural uncertainty Same thermodynamic parameters as MFE; model-bounded state space Comparison with probing data; validation of ensemble features Chapter 61
Kinetic folding simulation Folding pathway; interconversion rates Captures barrier effects and metastable states Rate model; discrete state representation; parameter completeness Time-resolved or single-molecule experiments Chapter 64
Cotranscriptional model Structures at successive nascent transcript lengths Explains nascent fold commitment and kinetic routing Transcription speed; pause-site locations; sequence boundaries Cotranscriptional chemical probing; pause-site perturbations Chapter 24, Chapter 60
Probing-constrained model Structure model informed by reactivity data Integrates experimental signal with energy model Reactivity reflects base-pairing status; no confounding proteins or modifications Independent structural or mutational validation Chapter 63
Comparative (covariation) model Consensus secondary structure; covarying positions Phylogenetically supported base-pair evidence Sufficient sequence diversity; accurate alignment Functional assays; mutagenesis; structural experiments Chapter 65
Molecular dynamics Atomic-resolution conformational sampling Captures dynamics and tertiary interactions Force field accuracy; accessible simulation timescales Comparison with NMR, SAXS, or cryo-EM data Chapter 64
Machine-learning model Predicted RNA structure or property from sequence Can capture patterns beyond physics-based models Training data distribution; may encode assay biases Independent dataset; mechanistic perturbation testing Chapter 65, Chapter 139

Free energy combines enthalpy and entropy. Enthalpy captures favorable and unfavorable energetic interactions. In an RNA helix, base stacking is often a major favorable contribution; hydrogen bonds, ion interactions, and hydration also matter. Entropy captures how many microscopic arrangements remain available. Folding a flexible RNA into an ordered hairpin reduces the number of chain conformations, which is entropically unfavorable. However, folding can release ordered water or counterions, and binding can redistribute motions in the RNA and its partners. The observed stability is the sum of these competing contributions, not the consequence of one interaction type alone.

Temperature affects the enthalpy-entropy balance. Heating increases thermal motion and makes many ordered structures less populated. Thermal melting experiments take advantage of this principle: an RNA duplex or hairpin is heated while an observable such as ultraviolet absorbance, fluorescence, or calorimetric heat flow reports unfolding. The melting temperature is useful, but it is not a universal stability number. It depends on sequence, concentration for intermolecular complexes, salt, pH, buffer, and how the melting transition is analyzed. Biological temperature effects are also important. RNA thermometers regulate gene expression by changing structure over physiological temperature ranges, and organisms adapted to different temperatures face different constraints on RNA stability and turnover.

Figure 3.4. From Free-Energy Difference to Equilibrium Population and Melting Behavior

Figure 3.4. From Free-Energy Difference to Equilibrium Population and Melting Behavior. A horizontal quantitative bridge follows a simple intramolecular two-state hairpin from enthalpy, entropy, and temperature through the folding free-energy difference, equilibrium constant, folded population, and a condition-dependent melting curve. Free-energy sign is tied to folded-favored, equal-population, and unfolded-favored regimes. A separate boundary-case band adds a populated intermediate and overlapping transitions, showing why one melting temperature cannot summarize a multi-state RNA ensemble. The figure explicitly marks the two-state assumptions and keeps salt, pH, sequence, construct, and concentration as condition qualifiers rather than universal constants.

Salt and ions deserve special attention because RNA is a polyanion. Each phosphate carries negative charge, and folding often brings phosphate groups closer together. Monovalent ions such as K+ and Na+ reduce electrostatic repulsion by screening charge. Divalent ions such as Mg2+ can be especially important because they screen charge more effectively and can support compact tertiary folds. Some Mg2+ ions remain part of a diffuse ion atmosphere, while others occupy more specific sites that stabilize a local motif or participate in catalysis. These modes should not be collapsed into the vague phrase “magnesium stabilizes RNA.” The magnitude and mechanism of stabilization depend on the RNA, ion concentration, competing ions, pH, temperature, and whether the observed effect is diffuse screening, site-specific binding, or indirect change in hydration.

RNA concentration also matters. Intramolecular folding of one strand into a hairpin is less concentration-dependent than intermolecular association of two strands into a duplex, but concentration can still influence aggregation, multimerization, protein binding, and experimental interpretation. Construct boundaries matter because removing flanking sequences can remove competing base-pairing regions or protein-binding sites. Chemical modifications matter because methylation, pseudouridylation, editing, and synthetic nucleoside substitutions can alter base pairing, stacking, hydration, and protein recognition. A thermodynamic claim is reusable only when these relevant conditions are stated.

The evidence basis for thermodynamic statements includes optical melting, isothermal titration calorimetry, differential scanning calorimetry, equilibrium binding assays, native gels, fluorescence assays, and structural or probing data interpreted with explicit models. Isothermal titration calorimetry can be especially useful because it measures heat released or absorbed during binding or folding transitions and can separate enthalpic and entropic contributions through model fitting. The limitation is that every method imposes a model: melting curves may assume two-state behavior, calorimetric fits may assume a binding stoichiometry, and fluorescence labels may perturb the molecule being studied. A careful thermodynamic interpretation states both the measurement and the assumptions.

3.2. Nearest-neighbor parameters and model assumptions

The number of possible secondary structures for an RNA grows rapidly as the sequence gets longer. A 30-nucleotide hairpin can be drawn and reasoned about by hand; a 3,000-nucleotide viral genome or a long 3′ untranslated region cannot. Energy models make folding prediction manageable by assigning scores to structural features and then searching among possible structures.

The nearest-neighbor model is the most important baseline for RNA secondary-structure energetics. The term “nearest neighbor” reflects the fact that the stability of a base pair in a helix depends strongly on the adjacent base pair. A GC pair stacked next to a GC pair contributes differently from a GC pair stacked next to an AU pair. The model therefore treats stacked pairs, hairpin loops, bulges, internal loops, multibranch loops, dangling ends, and terminal mismatches as local features whose energetic contributions can be added to estimate the stability of a full secondary structure.

The model works because RNA secondary structure has a strong local component. Watson-Crick and wobble base pairs form helices; loop penalties and stacking energies capture much of the energetic logic of many small RNAs. This local-additivity approximation turns sequence into a physical hypothesis and enables practical computation. A predicted hairpin can be used to design compensatory mutations, to compare alleles, to identify candidate riboswitch elements, or to generate hypotheses for chemical probing.

The assumptions are as important as the output. First, nearest-neighbor models assume approximate additivity: the total free energy is estimated by summing feature contributions. Real RNAs can contain nonlocal couplings in which a tertiary contact, ligand, protein, or long-range interaction changes the stability of a local helix. Second, parameters come from experiments under defined conditions. A parameter measured in one salt and temperature setting is not automatically correct in another. Third, many standard secondary-structure models simplify or omit pseudoknots. A pseudoknot occurs when base pairs cross in the conventional secondary-structure diagram; pseudoknots are common in some functional RNAs and viral elements but make prediction more complex. Fourth, many models treat the RNA as naked, unmodified, and at equilibrium. Cells rarely satisfy all three assumptions.

The boundary cases are biologically important. A riboswitch aptamer can be weakly structured without ligand and stabilized upon ligand binding. A protein bound to an untranslated region can block a base-pairing interaction that the nearest-neighbor model favors. A methylated base can change stacking or pairing so that unmodified parameters misestimate stability. A nascent RNA can form a structure before the full-length sequence is available, making the pathway more important than the full-length equilibrium score. A long RNA can contain local domains whose predicted structures are useful even when the global fold is not well captured.

The correct interpretation is therefore conditional. A phrase such as “the nearest-neighbor model predicts a stable stem-loop at positions 40-68 under the stated parameter set” is scientifically stronger than “the RNA has a stem-loop.” A predicted structure becomes a biological claim only when supported by evidence such as chemical probing, covariation, mutational disruption and compensatory rescue, ligand binding, protein binding, or functional assays. Chapter 60 develops the dynamic-programming algorithms that use these energy models, while this chapter emphasizes the assumptions needed to interpret their output.

3.3. Minimum-free-energy folding as a physical approximation

Minimum-free-energy prediction asks a concise question: among the structures allowed by a chosen model, which structure has the lowest estimated free energy? The answer is the minimum-free-energy structure, abbreviated MFE structure. MFE prediction is attractive because it returns a single structure that can be drawn, compared, mutated, and discussed. It is often the first computational view of an RNA when no experimental model exists.

The physical approximation is useful but narrow. Equilibrium thermodynamics predicts that lower-free-energy states are more populated than higher-free-energy states, but the lowest state does not necessarily contain all molecules. If one structure is much lower in free energy than all alternatives, the MFE drawing may be a good representation of the dominant secondary structure. If several structures are close in free energy, the MFE drawing can hide substantial alternatives. The difference between these cases is not philosophical; it changes experimental design. A mutation intended to disrupt a predicted MFE hairpin may have little effect if an alternative helix already dominates in the relevant cell state.

A simple example illustrates the problem. Suppose nucleotides 10-20 can pair with nucleotides 40-50, while nucleotides 10-20 can alternatively pair with nucleotides 80-90. If the first helix is predicted to be slightly lower in energy, the MFE output shows only that helix. Yet the second helix may have a substantial population, especially if a protein binds it, if a ligand stabilizes it, or if transcription pausing gives it time to form. A single-structure output can therefore encourage a false sense of certainty.

MFE prediction can also fail when the relevant biology lies outside the model. Many viral recoding signals depend on pseudoknots or mechanically resistant structures. Many ribosomal RNA and transfer RNA functions depend on tertiary architecture and protein contacts. Many long transcripts are structured by RNA-binding proteins and translation. Many chemically modified RNAs contain residues whose energetic effects are not represented in standard parameter sets. In these contexts, MFE prediction remains a hypothesis generator, not a final structural model.

The language used to report MFE results should make the evidentiary status plain. “The model predicts an MFE hairpin” is appropriate. “The predicted MFE structure suggests a possible regulatory stem” is appropriate. “This RNA forms the MFE structure in cells” is not appropriate unless the statement is supported by orthogonal evidence. This distinction is the core of an MFE structure is a model prediction under specified conditions, not direct evidence that an RNA molecule adopts only that structure.

MFE prediction is most useful when it is connected to tests. If a predicted stem matters, disrupting one side of the stem should change structure or function, and compensatory changes on the other side should restore pairing and function. If a predicted riboswitch fold matters, ligand concentration should alter the relevant structure or regulatory output. If a predicted UTR structure controls translation, ribosome profiling, reporter assays, and probing can test the mechanism. Chapter 5 treats this style of causal evidence more fully.

3.4. Partition functions, base-pair probabilities, and ensembles

An ensemble is a population of states. For RNA, the ensemble may include different secondary structures, tertiary conformations, ligand-bound states, protein-bound states, and partially unfolded intermediates. An equilibrium ensemble assigns relative populations according to free energy. A kinetic ensemble may also depend on pathways and barriers. The ensemble idea prevents one drawing from being mistaken for the molecule.

The partition function is the mathematical tool that turns an energy model into an ensemble description. In simplified terms, each modeled structure is weighted by a Boltzmann factor that depends on its free energy and temperature. The partition function sums those weights over all structures allowed by the model. From that sum, one can calculate ensemble free energy and estimate the probability that particular base pairs occur across the modeled ensemble. Chapter 61 develops the algorithmic details; the key point here is interpretive.

Figure 3.2. MFE Versus Partition-Function Interpretation

Figure 3.2. MFE Versus Partition-Function Interpretation. The minimum-free-energy (MFE) structure presents the single lowest-scoring secondary structure under a specified energy model, but alternative suboptimal structures can have substantial populations when their free energies are close to the MFE. A base-pair probability heat map, derived from the partition function, reveals how frequently each nucleotide pair appears across the modeled ensemble, exposing competing helices and uncertain regions that are invisible in the MFE drawing alone. Together, the MFE structure and base-pair probability matrix provide a more complete and honest representation of structural uncertainty than any single diagram.

Base-pair probability answers a different question from MFE prediction. MFE prediction asks which single structure has the lowest score. Base-pair probability asks how often a pair appears among the modeled structures. A helix whose base pairs each have high probability is robust under the model. A helix whose base pairs have low or uneven probabilities may be a fragile feature of one structure. Competing helices can produce intermediate probabilities, revealing alternatives hidden by a single MFE drawing. This is why partition-function calculations and base-pair probability heat maps often provide a more honest first view of structural uncertainty.

Box 3.1. How Not to Overinterpret Predicted Structures

  • The MFE structure is a model prediction, not an observed molecular structure. It is the single lowest-scoring structure under a specified model and conditions, not evidence that an RNA adopts only that structure in a cell.
  • Base-pair probability is a model-derived estimate from a partition-function calculation. It reflects modeled ensemble frequency, not measured occupancy in a living cell or extract.
  • Chemical probing in cells reports nucleotide reactivity or accessibility, not structure directly. Protein occupancy, RNA modifications, local flexibility, RNA abundance, and library-construction biases can all alter probing signals independently of base-pairing status.
  • Disagreement between a model prediction and an experiment is biological information, not simply a software error. A missing protein, an unmodeled modification, a translation event, or a compartment-specific factor may account for the discrepancy.
  • Detailed folding prediction algorithms, partition-function methods, and probing-constrained modeling are developed in Chapters 6065.

Ensemble thinking also clarifies functional alternatives. Riboswitches, attenuators, splicing elements, viral frameshifting regions, internal ribosome entry sites, and 3′ untranslated regions can use alternative states as part of their mechanism. An alternative state is not automatically noise. It can be the ligand-bound state, the translation-permissive state, the frameshifting state, or the decay-promoting state. Recent work on RNA secondary-structure ensembles emphasizes that functional RNAs often use population shifts rather than permanent all-or-none structures.

Ensemble concepts are also essential in RNA design. A designed RNA should not only have the desired target structure as its MFE structure; the target should be sufficiently populated, and undesired alternatives should be sufficiently suppressed for the intended application. Metrics such as ensemble defect, target probability, and competing-structure penalties express this logic. For a synthetic riboswitch, the relevant design goal may be a large ligand-dependent shift between two ensembles rather than a single maximally stable fold.

Base-pair probabilities are not measurements by themselves. They depend on the structure space included in the calculation, the parameter set, temperature, salt assumptions, sequence boundaries, and any experimental constraints. If chemical probing data are added as constraints or pseudoenergies, the result is still a model that integrates reactivity data with assumptions about how reactivity maps to structure. Probing reactivity can reflect flexibility, accessibility, protein occupancy, chemical modification, or technical bias as well as base-pairing status. Chapter 63 treats this integration in detail.

The strongest ensemble claims combine evidence. A predicted competing helix is more convincing if chemical probing supports the alternative state, mutations shift the population as predicted, ligand or protein binding changes the distribution, and the structural shift explains a functional outcome. Without that evidence, ensemble prediction should be reported as a modeled hypothesis.

3.5. Kinetic folding, metastability, and cotranscriptional traps

Thermodynamics tells which states are favored at equilibrium. Kinetics tells how fast states are reached and how molecules move among them. These are not interchangeable questions. A road may lead to the lowest valley, but a high mountain pass can make the journey slow. In folding language, the valleys are populated conformations and the mountain passes are kinetic barriers.

A metastable state is a state that persists because escape is slow. It need not be the lowest-free-energy state available to the full sequence. A kinetic trap is a metastable state that affects the observed or biological behavior of the molecule. Some traps are nuisances in folding experiments. Others are part of normal regulation. In RNA, traps are common because base pairs can form locally and quickly, but escaping from a set of base pairs may require breaking many interactions before a better structure can form.

Co-transcriptional folding makes kinetics unavoidable. During transcription, RNA polymerase synthesizes RNA from 5′ to 3′. The 5′ region emerges first and can fold before downstream nucleotides exist. A hairpin formed at transcript length 60 may prevent a helix that would be favored at transcript length 120. Polymerase pausing can intensify this effect by giving the nascent RNA more time to fold at a specific length. Therefore, a full-length equilibrium prediction may miss the historical route that produced the functional state.

Figure 3.3. Cotranscriptional Folding Routes and Kinetic Traps

Figure 3.3. Cotranscriptional Folding Routes and Kinetic Traps. During transcription, RNA emerges 5′ to 3′ and can fold before downstream sequence is available, so an early hairpin formed at one transcript length may persist as a kinetic trap even after a more stable alternative becomes accessible at a longer transcript length. Transcriptional pausing extends the time a nascent RNA spends at a particular length, increasing the chance that a hairpin or ligand-binding conformation commits the molecule to one folding route. Protein or helicase binding can remodel these early intermediates, illustrating how the folding pathway and the biological outcome depend on transcription speed, pause sites, and cellular factors as well as RNA sequence.

Riboswitches provide a clear mechanism. During transcription of some bacterial riboswitches, the aptamer domain emerges and can bind a metabolite before the expression platform has fully formed. If ligand binds in the appropriate time window, the ligand-stabilized structure can favor a terminator or anti-terminator, or can expose or hide a ribosome-binding site. If ligand is absent or arrives too late, the RNA may commit to a different regulatory state. The regulatory decision therefore depends on ligand concentration, thermodynamic stabilization, transcription speed, pause sites, and folding barriers.

Viral frameshifting elements show another kinetic dimension. A ribosome moving along an RNA can encounter a structured element that resists unwinding. The effect on frameshifting depends on structure, stability, mechanical response, and refolding kinetics. A model that considers only the equilibrium secondary-structure score may miss the feature that matters to the translating ribosome. A recent kinetic and thermodynamic analysis of Chikungunya virus frameshifting-element mutants illustrates how kinetic traps can be targeted and interpreted in a sequence-specific viral context.

Kinetic folding is difficult to measure because the population changes with time. Bulk assays can average together different routes. Single-molecule methods can observe transitions but often require labels or surfaces. Chemical probing can be adapted to nascent RNAs, but reactivity reports accessibility and chemistry rather than a complete structure. Computational kinetic models can simulate pathways, but the number of possible transitions is enormous and the rates are model-dependent. Co-transcriptional chemical probing and co-transcriptional simulations are therefore most powerful when used together with perturbations that test a proposed route.

The boundary case is important: not every RNA requires kinetic modeling. Some short RNAs refold rapidly enough that equilibrium approximations are adequate for the question being asked. Some local helices are useful predictors even in long transcripts. The scientific task is to decide when timing, barriers, and transcription history are likely to change the biological conclusion. Kinetic claims are strongest when they include time-resolved evidence, rate estimates, or perturbations of transcription speed, pause sites, barriers, or remodeling factors.

3.6. Limits of energy models in cellular environments

Standard energy models are usually built from measurements on purified RNAs or short constructs under defined conditions. Cells are more complicated. Cellular RNAs fold while they are transcribed, processed, exported, localized, translated, modified, surveilled, degraded, and assembled into ribonucleoprotein particles. The relevant environment can include abundant RNA-binding proteins, ATP-dependent helicases, ribosomes, small molecules, polyamines, metabolites, membranes, compartments, and condensates. These factors can change both state populations and transition rates.

RNA-binding proteins can stabilize a structure by recognizing a loop, a single-stranded motif, a double-stranded region, or a larger RNA architecture. They can also prevent a structure by occupying a sequence that would otherwise pair. Some proteins act as RNA chaperones, increasing the chance that an RNA finds a functional fold. Helicases can use ATP to disrupt helices, remodel RNP complexes, or clear structures from translated regions. Ribosomes unwind coding-region structures as they translate. Decay and surveillance factors can selectively remove RNAs with particular structural or RNP features. A folding landscape inside a cell is therefore partly written by the RNA sequence and partly sculpted by the cell’s machinery.

Chemical modifications add another layer. A modified nucleoside can change base-pairing preference, base stacking, hydration, local flexibility, or protein recognition. N6-methyladenosine, pseudouridine, inosine generated by adenosine-to-inosine editing, 2′-O-methylated residues, and synthetic nucleoside analogs can all alter the meaning of an unmodified sequence. The energetic effect is context-dependent: a modification in a loop can influence protein binding, while the same modification in a helix can influence pairing or helix geometry. Chapter 57 covers RNA modifications in depth.

Compartments and biological assemblies matter because local concentration and binding-partner availability matter. A nuclear pre-mRNA, a cytoplasmic mRNA, a mitochondrial transcript, an RNA in a viral replication complex, and an RNA inside a therapeutic delivery particle can experience different ion composition, protein occupancy, crowding, and time constraints. A structure that appears unstable in a dilute buffer may be stabilized by a protein or ligand in one compartment. A structure predicted to be stable may be unwound by ribosomes or helicases in another.

In-cell probing and in-vivo-like parameter work are important because they bring models closer to biological conditions. For example, work on in-vivo-like nearest-neighbor parameters aims to improve prediction of fractional base pairing in cells. Such approaches are valuable, but they do not remove the need for interpretation. A probing signal can be altered by protein binding, RNA abundance, modification, local reactivity, reverse-transcription bias, or sequencing bias. A model that uses probing data can therefore fit the data while still being wrong about the mechanism if the data are interpreted too narrowly.

Predictive models of RNA cellular activity from sequence and conformational ensembles can be biologically useful even when they are incomplete. A predictive feature can identify a sequence design that works, a variant that changes expression, or a region worth mutating. Predictive success does not by itself prove the causal molecular mechanism. Mechanistic proof requires additional evidence, such as mutational rescue, direct binding, time-resolved folding, structural measurement, or functional perturbation of the proposed cellular factor.

The practical rule is to treat disagreement between model and experiment as information. If a predicted helix is absent in cells, the model may be wrong, the conditions may be mismatched, a protein may occupy the region, the RNA may be modified, translation may unwind it, or the probing method may be biased. If an unpredicted in-cell structure appears repeatedly and is functionally validated, the missing factor becomes part of the biology rather than a mere software error.

Experimental Foundations and Evidence

RNA thermodynamics and kinetics are inferred through multiple evidence types. No single method answers every question. The strongest interpretations align the question, the method, the model, and the perturbation.

Thermal melting assays follow folding or duplex formation as temperature changes. Ultraviolet absorbance can report base unstacking, fluorescence can report local environmental change, and calorimetry can report heat flow. These methods can estimate melting temperatures, enthalpy, entropy, and free energy under fitted models. Their main limitations are model dependence, construct dependence, concentration effects for intermolecular reactions, and the risk that a complex RNA unfolds through more than one transition.

Calorimetric and binding assays are useful when folding is coupled to ligand, metal ion, or protein binding. Isothermal titration calorimetry can measure heat changes during binding and support estimates of affinity, enthalpy, and stoichiometry. Fluorescence anisotropy, Forster resonance energy transfer, native gels, filter binding, and surface-based methods can measure binding or conformational changes, but each introduces possible artifacts such as labeling effects, surface effects, or separation of bound states during handling.

Chemical probing measures nucleotide reactivity or accessibility. Selective 2′-hydroxyl acylation analyzed by primer extension, dimethyl sulfate probing, hydroxyl radical probing, and related methods can identify flexible, accessible, or protected regions. These data are powerful when interpreted with controls and models, but probing does not directly return a structure. Reactivity can be influenced by base-pairing, tertiary contacts, protein occupancy, modification, local chemistry, RNA abundance, and library construction. Probing-constrained models are therefore integrated hypotheses, not assumption-free observations.

Comparative sequence analysis tests whether paired positions covary across evolution. If one side of a predicted pair changes and the opposite side changes in a compensatory way that preserves pairing, the evidence for the pair strengthens. Covariation is especially useful for conserved structured RNAs, but it is weaker for young, rapidly evolving, poorly aligned, or species- restricted RNAs. Covariation supports structural conservation; it does not by itself measure folding kinetics or cellular occupancy.

Mutational analysis tests causality. A disruption mutation should weaken the predicted structure and change the relevant function; a compensatory mutation should restore pairing and rescue function. This logic is powerful because it connects structure to mechanism. It can still be confounded if the mutation changes protein binding, codon usage, RNA modification, translation, localization, or decay independently of the intended structure.

Single-molecule methods, stopped-flow experiments, time-resolved probing, and co-transcriptional assays address kinetics more directly. They can reveal intermediates, rates, and molecule-to-molecule heterogeneity that bulk equilibrium assays hide. Their limitations include labeling, surfaces, temporal resolution, signal assignment, and the challenge of matching in vitro conditions to cellular conditions. Strong kinetic claims specify the system, time scale, observed intermediate, and connection to function.

Computational evidence is strongest when its assumptions are explicit. Minimum-free-energy prediction, partition-function analysis, kinetic simulation, molecular dynamics, machine-learning prediction, and probing-constrained modeling answer different questions. A computational model can be an excellent guide to experiment and design, but a model output should not be promoted to an observed mechanism without additional evidence.

Biological Contexts

Bacterial regulatory RNAs often expose the interplay of thermodynamics and kinetics because transcription and translation are closely coupled. A nascent bacterial transcript can fold before transcription is complete, and a ribosome can bind or translate the RNA while downstream sequence is still emerging. Riboswitches, transcription attenuators, small regulatory RNAs, and bacterial thermometers can therefore depend on timing as well as stability.

Eukaryotic pre-mRNAs fold while they are transcribed by RNA polymerase II and while they are coated with processing factors. Splicing, polyadenylation, export, RNA modification, and RNP assembly can all influence the structures available to a pre-mRNA. The larger size of many eukaryotic transcripts makes global equilibrium prediction less informative, but local structures and protein-stabilized regions can still be crucial.

Translation changes the folding problem for mRNAs. A coding sequence may be structured when naked, but ribosomes unwind local structures as they move. The 5′ untranslated region can influence initiation by exposing or hiding translation-initiation features. Coding-region structures can affect ribosome speed, frameshifting, mRNA decay, or cotranslational folding of the encoded protein. The relevant structure is therefore a time-dependent structure in a translating RNP, not just a full-length naked mRNA fold.

Ribosomal RNA and transfer RNA illustrate the boundary between secondary structure and full three-dimensional architecture. Their functional states depend on base-paired stems, tertiary contacts, metal ions, modifications, and protein interactions. A secondary-structure energy model can identify many local helices, but the mature ribosome or tRNA is an RNP assembly with a specific three-dimensional fold. Chapter 4 develops the structural principles that extend beyond this chapter’s thermodynamic baseline.

Viral RNAs frequently use structured elements for replication, translation, packaging, immune evasion, and recoding. Viral genomes can be long enough that alternative states, long-range contacts, and protein interactions matter, yet compact enough that conserved RNA structures can be experimentally and computationally tractable. Viral frameshifting elements are particularly useful examples because mechanical resistance, refolding, and ribosome interaction can matter as much as equilibrium stability.

Synthetic and therapeutic RNAs require the same caution. An mRNA vaccine or therapeutic mRNA is designed for translation, stability, manufacturability, and immune compatibility. Codon choice, untranslated-region design, chemical modifications, poly(A) tail length, formulation, and delivery environment can alter structure and activity. A structure prediction can help identify candidate features, but clinical or manufacturing performance cannot be deduced from MFE structure alone. Chapters on RNA therapeutics and delivery develop those applications in more detail.

Thermodynamic folding models support many practical technologies. Primer and probe design use estimates of duplex stability and off-target structures. Guide RNA design for CRISPR systems considers guide accessibility, alternative folds, target pairing, and protein-bound architecture. Antisense oligonucleotide and small interfering RNA design must consider target-site accessibility, local RNA structure, chemical modifications, and protein occupancy. RNA aptamer and riboswitch engineering uses ligand-dependent population shifts as a design principle.

Box 3.2. Required Qualifiers for Thermodynamic Claims

  • Sequence and construct boundaries, including flanking sequence, truncations, linkers, or scaffold additions.
  • Temperature, stated in °C or K.
  • Monovalent salt species and concentration (for example, 100 mM KCl or 1 M NaCl).
  • Mg²⁺ or other divalent ion concentration, or an explicit statement that none was present.
  • RNA concentration, especially for intermolecular or duplex reactions where concentration changes the equilibrium.
  • Modification status: which nucleosides are chemically modified, synthetic, or RNA-edited.
  • Protein or ligand presence, including identity and concentration when relevant.
  • Model and parameter set when the claim is computational, including software name and version if the comparison will be reused.

Partition-function and ensemble metrics are increasingly important in RNA engineering because function often depends on suppressing unwanted states, not only creating a desired state. A designed hairpin that is the MFE structure but has many near-degenerate competitors may fail in a cell. A riboswitch designed to change translation must produce a useful difference between ligand-free and ligand-bound ensembles. A therapeutic mRNA design may seek to balance translation efficiency, RNA stability, innate immune sensing, manufacturability, and avoidance of problematic structures.

Machine-learning models can use sequence, predicted structure, ensemble features, probing data, or experimental measurements to predict RNA behavior. These approaches can improve practical prediction when trained and validated well, but they inherit the evidence problem of all predictive models. A model may learn correlations involving sequence composition, expression, structure, or assay bias. A successful prediction should be distinguished from a causal explanation until perturbations test the proposed mechanism.

Clinical interpretation of RNA variants sometimes requires structure-aware reasoning. A variant in an untranslated region, splice element, viral genome, or therapeutic RNA can change a local structure, a protein-binding site, a modification context, or an ensemble distribution. Structure prediction can prioritize hypotheses, but clinical claims require genetic, biochemical, cellular, or patient-level evidence appropriate to the disease or intervention. Chapter 5 gives the general evidence framework.

Recent Consensus

Current consensus treats RNA folding as an ensemble and landscape problem rather than a single-structure drawing problem. A secondary-structure diagram is valuable, but it is a representation of selected interactions, not a full description of all populated states or all cellular influences.

Nearest-neighbor and MFE models remain essential baselines. They are interpretable, widely available, experimentally grounded, and often useful. Their limitations are also well established: parameter conditions, local-additivity assumptions, omitted pseudoknots or tertiary contacts, and missing cellular factors must be considered before drawing biological conclusions.

Partition-function calculations and base-pair probabilities are better suited than single MFE structures for communicating structural uncertainty. They do not solve every modeling problem, but they reveal alternative states and ambiguous regions that single-structure outputs hide.

Co-transcriptional folding is a major biological reality. Transcription speed, pause sites, nascent chain length, ligand timing, protein binding, and folding barriers can determine which states form and persist. Equilibrium prediction of a full-length RNA is therefore insufficient for mechanisms in which early folding choices affect later outcomes.

Cellular context must be part of interpretation. Proteins, helicases, ribosomes, modifications, ions, metabolites, compartments, and crowding can produce structures or activities that differ from naked-RNA predictions. The challenge is to identify which missing factor matters for a particular RNA, not to abandon physical modeling.

Open Questions, Controversies, Deprecated Models, and Common Misconceptions

Open questions:

  • How transferable can in-cell energy parameters be across organisms, cell types, compartments, growth states, and RNA classes? Parameters that improve average prediction in one setting may not capture local effects from a specific protein, modification, or compartment. More direct cellular measurements are needed before broad generalization is justified.
  • When is kinetic modeling necessary? Some RNAs behave well enough under equilibrium approximations for a given purpose. Other RNAs, especially nascent regulatory RNAs and recoding elements, cannot be understood without timing and barriers. The field needs clearer evidence standards for deciding when a kinetic explanation is required rather than merely plausible.
  • How can researchers combine probing data, thermodynamic models, evolutionary covariation, molecular simulation, and machine learning without circular validation? If probing data are used to train a model and then the same style of probing data are used to validate it, the apparent agreement may overstate biological accuracy. Independent perturbations remain important.

Common misconceptions:

  • “The MFE structure is the RNA structure.” The MFE structure is the lowest-scoring structure under a model. It may be a good hypothesis, but it is not direct evidence of cellular structure.
  • “High base-pair probability proves a base pair exists in cells.” Base-pair probability is model-derived unless validated under relevant conditions. It is stronger when supported by probing, covariation, mutational rescue, or functional evidence.
  • “Mg2+ simply stabilizes all RNA.” Mg2+ can screen charge, support compact folds, bind specific sites, and participate in catalysis, but its effects depend on concentration, competing ions, RNA architecture, and context.
  • “Chemical probing directly gives a structure.” Probing gives reactivity or accessibility data. Structure models inferred from probing require assumptions and can be confounded by protein occupancy, modifications, flexibility, and technical bias.
  • “Disagreement between prediction and experiment means the software is useless.” Disagreement may reveal missing biology, including proteins, modifications, kinetics, translation, ligand binding, or compartment-specific conditions.

Deprecated or weakened claims:

  • A deprecated simplification is the idea that each RNA has one native structure that can be read directly from the sequence. Some RNAs do have dominant functional structures under particular conditions, but many RNAs function through alternative states, binding-induced shifts, kinetic routes, or RNP assemblies. The “one sequence, one structure” habit is especially misleading for long transcripts, regulatory RNAs, and cellular RNPs.