Chapter 56. RNA-Protein Recognition, RBP Specificity, and RNP Assembly

Scope Note

This chapter explains how proteins recognize RNA, how RNA-binding proteins (RBPs) achieve specificity, and how RNA-protein complexes assemble into functional ribonucleoprotein particles (RNPs). The chapter treats recognition as a physical and biological process: an RBP does not simply read a sequence string. It encounters an RNA molecule with a backbone, bases, secondary and tertiary structure, chemical modifications, competing proteins, cellular localization, abundance, and kinetic history. RNP assembly therefore depends on molecular interfaces, timing, concentration, remodeling, and evidence from biochemical, structural, and transcriptome-wide assays.

Executive Summary

RNA-protein recognition is the set of molecular events by which a protein binds RNA with a defined affinity, residence time, and biological consequence. The simplest examples involve a folded RNA-binding domain contacting a short sequence motif. Many cellular examples are more complex. A protein may recognize the shape of an RNA helix, the pattern of single-stranded and double-stranded regions around a motif, a chemically modified nucleotide, a ribose-phosphate backbone geometry, or an RNP surface already occupied by other proteins. Recognition can occur by induced fit, in which binding reshapes the RNA or protein, or by conformational selection, in which the protein preferentially captures one state from a pre-existing RNA ensemble. These models are not mutually exclusive, and many RNPs use both principles.

RNA-binding domains define the vocabulary of recognition. RNA recognition motifs (RRMs), K-homology (KH) domains, zinc-finger domains, double-stranded RNA-binding domains (dsRBDs), S1 domains, cold-shock domains, Pumilio and fem-3-binding factor (PUF) repeats, La motifs, PAZ and PIWI domains, LSm rings, ribosomal proteins, and disordered arginine-rich regions solve different recognition problems. Some domains contact exposed bases and read sequence. Some contact the minor groove or backbone of A-form RNA helices and read structure more than sequence. Some bind with modest affinity alone but become selective when repeated in tandem or embedded in a larger protein. Domain names are therefore useful starting points, not complete predictions of cellular targets.

Specificity arises from layered filters. The first filter is the intrinsic interface: which bases, sugars, phosphates, or helical surfaces the protein can contact. The second filter is RNA accessibility: whether the motif is folded, protein-covered, modified, localized, or being translated or processed. The third filter is cellular competition: other RBPs, ribosomes, nucleases, small RNAs, and RNP assembly factors can occupy or remodel the same RNA. The fourth filter is stoichiometry: a high-affinity site may remain unoccupied if the protein is scarce or sequestered, whereas abundant proteins can occupy lower-affinity sites. A binding motif is therefore a conditional opportunity, not a guarantee of binding.

RNA modifications add another specificity layer. Modified nucleotides can change local base pairing, stacking, hydration, and protein recognition. Some proteins read modifications directly, whereas others bind indirectly because a modification changes RNA structure or prevents another protein from binding. The current literature supports modification-specific recognition as an important principle, but individual claims require careful evidence because detection methods, modification stoichiometry, and antibody or sequencing artifacts can distort interpretation.

RNP assembly is often ordered and kinetic. A box C/D small RNP, an mRNP, an influenza viral RNP, a ribosomal subunit precursor, a miRNA Argonaute complex, or an RNA granule does not appear fully formed in one equilibrium step. RNA folding, protein binding, modification, cleavage, transport, and remodeling occur in partially overlapping sequences. Assembly can follow alternative pathways, as shown for archaeal box C/D sRNPs, and large RNPs may require RNA helicases, chaperones, ATPases, and quality-control checkpoints to avoid off-pathway states.

RBP networks are competitive and collective. Many transcripts contain overlapping binding sites for stabilizing RBPs, decay factors, splicing regulators, translation regulators, miRNA-associated complexes, and structural assembly factors. The functional output depends on which factors bind first, which factors are present at limiting concentration, and which interactions are reinforced by multivalent contacts. Transcriptome-scale CLIP, single-cell interaction profiling, RNP-MaP, and RNA-dependent interactome methods have broadened the field from one-protein-one-target models to dynamic RNP assemblies, but these methods also require artifact controls.

Disease can arise from both sides of the interface. Mutations in an RBP can alter RNA affinity, specificity, localization, phase behavior, or assembly timing. Mutations in an RNA can create, remove, expose, hide, or chemically alter a binding site. Misregulation can also occur without mutation when stress, viral infection, inflammation, DNA damage, cancer signaling, or developmental state changes the concentration or modification state of an RBP or RNA. Because RBP networks are dense, disease mechanisms should not be inferred from binding alone; they require evidence that altered binding changes RNA fate, cellular phenotype, and, where relevant, organismal disease.

Concept Inventory

  • RNA-binding protein (RBP): a protein that binds RNA directly or as part of an RNA-containing complex. The term includes proteins with canonical RNA-binding domains and proteins whose RNA binding was discovered by interactome methods. Boundary case: a protein found near RNA by crosslinking or co-purification is not necessarily a direct sequence-specific RBP.
  • Ribonucleoprotein particle (RNP): an assembly containing RNA and protein. RNPs range from small complexes such as an RBP bound to one motif to large machines such as ribosomes, spliceosomes, telomerase, viral replication complexes, and RNA granules.
  • RNA recognition motif (RRM): a common RNA-binding domain of about 90 amino acids with a beta-alpha-beta-beta-alpha-beta fold. Canonical RRMs often present aromatic and charged residues on a beta-sheet surface to contact exposed RNA bases and backbone. Boundary case: not all RRMs bind RNA in the same orientation or with the same sequence length.
  • K-homology (KH) domain: an RNA-binding domain that uses a conserved loop, often described by a GXXG sequence pattern, to bind short single-stranded RNA segments. KH domains are modular and frequently occur in tandem.
  • Zinc-finger RBP: an RNA-binding protein that uses a zinc-coordinating fold, such as CCCH, CCHC, or related motifs, to position residues for RNA contact. Boundary case: some zinc fingers bind DNA, protein, or lipid-associated complexes rather than RNA, so zinc coordination alone does not prove RNA specificity.
  • Double-stranded RNA-binding domain (dsRBD): a domain that usually recognizes the shape and backbone geometry of A-form double-stranded RNA. dsRBDs often read structure more than primary sequence, although neighboring domains or local defects can add specificity.
  • Specificity: preferential binding of one RNA site, structure, modification state, or RNP context over alternatives. Specificity is not identical to high affinity; a protein can bind tightly but promiscuously, or weakly but selectively under cellular conditions.
  • Cooperativity: a situation in which one binding event changes the probability of another binding event. Cooperativity can arise from direct protein-protein contact, RNA conformational change, multivalent domain architecture, or competition for overlapping sites.
  • Multivalency: use of multiple binding modules or repeated interaction surfaces in the same molecule or assembly. Multivalency increases avidity, enables low-affinity motifs to contribute to stable complexes, and supports RNP granule organization.
  • RNP remodeling: active or passive rearrangement of RNA-RNA, RNA-protein, or protein-protein contacts within an RNP. Remodeling can involve RNA helicases, ATPases, chaperones, modification enzymes, nucleases, translation, or changes in local concentration.
  • CLIP: crosslinking and immunoprecipitation, a family of methods that use covalent RNA-protein crosslinks, immunoprecipitation, and sequencing to map RBP-associated RNA fragments. Boundary case: CLIP peaks are evidence of proximity and recovery under a protocol, not automatic proof that a site is functional.

What to Know Before Reading This Chapter

RNA is a charged polymer with bases that can pair, stack, loop out, form noncanonical contacts, and carry chemical modifications. A protein approaching RNA can therefore encounter several kinds of information at once: base identity, backbone shape, helix geometry, loop conformation, modification chemistry, and whether another molecule already occupies the site. RNA-protein recognition differs from DNA recognition because RNA is usually single-stranded or locally structured, folds into heterogeneous ensembles, and is continuously processed, translated, degraded, or assembled into RNPs.

Affinity is the strength of an interaction under a specified condition. Specificity is the preference for one partner or site over others. Occupancy is the fraction of sites bound in a cell or experiment. These three quantities should not be merged. A protein can have nanomolar affinity for a motif in vitro but low occupancy in cells if the motif is folded, the protein is absent from the compartment, or a competitor binds first. Conversely, a protein with modest intrinsic affinity can occupy many transcripts if it is abundant and multivalent.

Most RBPs contact RNA through several weak interactions rather than through one decisive bond. Hydrogen bonds, electrostatic contacts, cation-pi and pi-pi interactions, base stacking, shape complementarity, metal-ion coordination, and water-mediated contacts can all contribute. Because the RNA backbone is negatively charged, many RNA-binding surfaces are enriched in positively charged residues, but basic charge alone is not enough for specificity. Specificity usually requires geometry and context.

The central examples in this chapter are canonical RBP domains, modified RNA recognition, archaeal box C/D sRNP assembly, rRNA assembly during transcription, influenza viral RNP assembly, RNP granules, and transcriptome-wide mapping of RBP occupancy. Later chapters treat large RNP machines, phase separation, structural methods, CLIP families, and RNA-centric proteomics in more depth. This chapter focuses on the shared recognition and assembly logic that makes those later systems interpretable.

56.1. RRM, KH, zinc-finger, dsRBD, and other RBP domains

Table 56.1. Evidence Classes for RNA-Protein Recognition. This table compares eight evidence classes used to study RNA-protein recognition, summarizing what each method supports, what it cannot prove alone, critical controls, and the best complementary approach.

Evidence class What it supports What it cannot prove alone Critical controls Best paired method
Purified binding assay Affinity, kinetics, sequence dependence Cellular occupancy RNA folding and protein-quality controls CLIP or cellular reporter
Mutational binding assay Base or structure contribution to affinity Indirect effects in cells Compensatory or matched mutations Rescue experiment
CLIP or seCLIP Cellular proximity and recovered binding regions Direct functional regulation Antibody controls, size-matched input, biological replicates Perturbation or factor depletion
Single-cell interaction profiling Cell-state context for interactions Direct binding mechanism Expression and cell-quality controls Targeted biochemical validation
RNA-centric proteomics or RNP-MaP Assembly composition or network context Identity of direct RNA binders Non-target controls, RNase dependence Purified binding or structural data
Structural biology Interface geometry and contacts Pathway order or cellular occupancy Functional-state sample and construct controls Biochemical and cellular tests
Reconstitution Assembly sufficiency and component order Cellular competitors and context Component purity and stoichiometry controls Genetic perturbation
Functional rescue Causal output linking binding to phenotype Exact molecular contact identity Expression, localization, and off-target controls Binding assay and structural data

Box 56.1. Affinity, Specificity, Occupancy, and Function Are Different Claims

A purified protein may bind an RNA motif with high affinity, but the site may be unoccupied in cells if the motif is folded into a helix, modified, localized to a different compartment, or outcompeted by another factor. A CLIP peak may show cellular proximity, but proximity does not prove the binding regulates RNA fate. A mutation may change in vitro binding and still produce no phenotype if the network is buffered by redundant contacts. Strong mechanism claims connect all four levels: binding strength, preference over alternative sites, cellular occupancy, and functional output.

An RNA-binding domain is a folded or partially folded protein region that contributes directly to RNA contact. Domains matter because they reveal recurring physical solutions to the problem of recognizing a flexible nucleic acid. The same domain family can occur in many proteins and pathways, but each protein embeds the domain in a different architecture, expression pattern, and regulatory context. A domain name should therefore be read as a hypothesis about binding mode, not as a complete annotation of function.

The RNA recognition motif, abbreviated RRM, is one of the most common eukaryotic RNA-binding domains. A canonical RRM contains a four-stranded beta-sheet backed by alpha-helices. Conserved sequence features, often called RNP1 and RNP2 motifs, place aromatic and charged residues on the beta-sheet surface. In many RRMs, exposed bases from single-stranded RNA stack on aromatic side chains and make hydrogen bonds to side chains or backbone atoms. This explains why RRMs often recognize short single-stranded motifs in pre-mRNAs, mRNAs, snRNAs, and viral RNAs. It also explains why one RRM by itself may bind only a few nucleotides with modest affinity.

RRM specificity becomes sharper when domains work together. Tandem RRMs can bind adjacent RNA segments, bridge a loop and a single-stranded tail, impose a bend, or contact a structured RNA surface. Flexible linkers allow alternative registers, whereas fixed domain-domain orientations can require a precise spacing between motifs. In splicing regulators and heterogeneous nuclear ribonucleoproteins, the same RRM-containing protein may bind many RNAs, but functional specificity emerges from motif clusters, local RNA structure, neighboring RBPs, and recruitment to nascent transcripts. Structural reviews of RNA-protein recognition emphasize this induced-fit and modular-binding logic rather than a one-domain-one-code model.

Table 56.2. Domain, RNA Feature, and Specificity Source. This table summarizes major RNA-binding domain types, the RNA features they typically recognize, the main sources of their specificity, important boundary cases, and illustrative biological pathways.

Domain or architecture Typical RNA feature Common specificity source Boundary case Example pathway
RRM Single-stranded motifs or structured loops Base contacts on beta-sheet surface plus domain context Individual RRMs often have weak or broad specificity alone Splicing regulation and mRNA stability
KH domain Short single-stranded segments Cleft geometry and avidity from tandem spacing Specificity requires tandem arrangement or contextual partners Translation regulation, localization, splicing
Zinc finger Short motifs or structured RNA elements Metal-stabilized fold positions RNA-contacting residues Zinc fingers can bind DNA or protein, not only RNA RNA stability control and viral RNP contexts
dsRBD A-form double-stranded RNA Helix shape and phosphate-backbone contacts Limited sequence specificity when domain acts alone RNA editing, RNA interference, innate immune sensing
PUF repeat Extended single-stranded sequences Modular repeat-to-base recognition, one base per repeat Specificity depends on repeat number and spacing mRNA repression and localization
Disordered arginine-rich region Distributed contacts across RNA Charge complementarity, cation-pi interactions, multivalency Prone to nonspecific binding in vitro at high concentration RNA granule assembly and viral protein binding

KH domains solve a related but distinct problem. A KH domain usually binds a short, single-stranded RNA segment along a cleft formed by helices, sheets, and a conserved loop. The local geometry lets the domain read a small sequence patch while accommodating the ribose-phosphate backbone. KH domains often occur in repeated arrays, so one protein can contact several short motifs and build specificity from spacing and avidity. This architecture is useful for proteins that regulate splicing, translation, localization, or stability of many mRNAs. Final-reference expansion notes: the local Chapter 56 bibliography supports general RBP-domain principles but lacks a dedicated KH-domain structural review.

Figure 56.1. Domain Architectures for RNA Recognition

Figure 56.1. Domain Architectures for RNA Recognition. Major RNA-binding domains solve different recognition problems. A canonical RRM presents aromatic and charged residues on a beta-sheet surface to contact exposed bases and backbone of short single-stranded RNA. A KH domain binds a short single-stranded segment in a cleft formed by helices and a conserved loop. A zinc-finger domain uses metal coordination to position RNA-contacting residues, but zinc fingers can also bind DNA or protein, so direct biochemical evidence is needed to establish RNA specificity. A dsRBD reads the A-form geometry of double-stranded RNA, primarily contacting the minor groove and backbone, making it more sensitive to RNA shape than to primary sequence.

Zinc-finger RBPs use metal coordination to stabilize compact folds that position RNA-contacting residues. CCCH zinc fingers are common in post-transcriptional regulators that bind U-rich or AU-rich elements, CCHC zinc knuckles occur in viral and cellular RNA-binding contexts, and some C2H2-like domains can bind RNA as well as DNA or protein. Zinc fingers illustrate a key boundary case: the presence of a zinc finger does not by itself define the nucleic acid substrate. Direct biochemical or structural evidence is needed to show whether a given zinc finger recognizes RNA, DNA, a protein partner, or a composite RNP surface.

Double-stranded RNA-binding domains, or dsRBDs, typically recognize the A-form geometry of double-stranded RNA. A-form RNA has a deep, narrow major groove and a wide, shallow minor groove compared with B-form DNA. Many dsRBDs contact the minor groove, phosphate backbone, and helix shape rather than reading exposed base edges in a sequence-specific way. This allows proteins in RNA editing, RNA interference, antiviral sensing, and processing pathways to bind duplex RNA, hairpins, or structured precursors. Specificity often comes from neighboring domains, local mismatches, loops, terminal structures, or cellular localization rather than from the dsRBD alone.

Other RBP architectures expand the recognition repertoire. PUF-repeat proteins use repeated modules to recognize bases in an extended single-stranded RNA, often one base per repeat. S1 and cold-shock domains bind single-stranded RNA surfaces in ribosomal, bacterial, and stress-related proteins. PAZ and PIWI domains contribute to small-RNA guide recognition in Argonaute-family proteins. LSm and Sm rings encircle or bind RNA ends and form stable RNP cores. Ribosomal proteins use globular domains, extensions, and basic tails to stabilize rRNA folds. Arginine-rich and glycine-rich low-complexity regions often bind RNA through distributed charge, cation-pi interactions, or multivalent contacts rather than through a single rigid pocket. Final-reference expansion notes: these domain-specific examples need a later targeted bibliography pass for PUF, Sm/LSm, PAZ/PIWI, and ribosomal protein domain references.

The common lesson is that RBP domains are interpretable only in context. A canonical domain can bind noncanonical substrates, a disordered region can supply much of the cellular affinity, and a protein without a classic RBP domain can still crosslink reproducibly to RNA. Recent discussions of “rethinking” RBPs emphasize that riboregulation includes enzymes, metabolic proteins, structural proteins, and stress-responsive assemblies whose RNA interactions may be conditional rather than obvious from domain annotation alone.

56.2. Sequence, structure, and modification specificity

Sequence specificity means that a protein preferentially binds one nucleotide pattern over alternatives. The pattern may be a short word such as a U-rich element, a longer motif with degenerate positions, or a set of spaced elements. Sequence recognition is easiest to see in purified binding assays: mutate a base, measure lower affinity, and restore binding by restoring the motif. Cellular sequence specificity is more conditional because the same sequence may be hidden in a helix, occupied by another protein, edited, methylated, translated by a ribosome, or absent from the relevant isoform.

Structure specificity means that a protein recognizes RNA shape or conformational state. An RBP may bind a hairpin loop, internal loop, kinked helix, bulged nucleotide, pseudoknot, G-rich structure, double-stranded region, or tertiary junction. Structure-specific recognition can be sequence-dependent if the structure requires particular bases, but the recognized object is not simply the linear sequence. For example, a dsRBD that contacts A-form geometry may tolerate many base sequences, whereas a loop-binding RRM may need both the loop sequence and the stem that presents the loop. RNA 3D-structure prediction reviews emphasize that RNA structure is an ensemble problem, so a “binding structure” is often one populated state among alternatives (WangRNA3D2023).

Induced fit and conformational selection are two useful models for structure-specific recognition. In induced fit, the protein binds an RNA and then reshapes it, perhaps flipping a base, stabilizing a loop, or bending a helix. In conformational selection, the RNA transiently samples a structure, and the protein preferentially binds that pre-existing state. Real systems often combine the two: a protein first captures a rare exposed motif and then stabilizes a more ordered RNP interface. This is why measured affinity can depend on folding conditions, temperature, ion concentration, transcript length, and prior protein binding.

Modification specificity adds chemical information. RNA modifications such as N6-methyladenosine (m6A), N6,2′-O-dimethyladenosine (m6Am), pseudouridine, 2′-O-methylation, 5-methylcytidine, inosine, acetylcytidine, queuosine, and many tRNA modifications can alter protein binding by several mechanisms. A modification can create a direct recognition surface for a reader protein. It can destabilize a local helix and expose a motif. It can stabilize a structure and hide a motif. It can change reverse-transcription behavior in sequencing assays, creating apparent binding or modification patterns that require validation. Reviews of modification recognition and modification detection emphasize both biological importance and method caution.

The distinction between direct and indirect modification recognition is essential. Direct recognition means the protein contacts the modified chemical group in a way that contributes to binding specificity. Indirect recognition means the modification changes RNA structure, localization, decay, translation, or competition, and the protein responds to that altered state. Both mechanisms are biologically real, but they require different evidence. Structural data or targeted chemical substitution can support direct recognition. Structure probing, binding competition, modification stoichiometry, and rescue experiments are needed to support indirect mechanisms.

Modification stoichiometry is a recurring boundary case. A transcriptome site may be modified in only a fraction of molecules. If an RBP binds only the modified fraction, bulk RNA measurements may dilute the effect. If a modification is inferred from a sequencing signature with false positives or context bias, a claimed reader interaction may rest on an uncertain substrate. Direct RNA sequencing and other modification-detection methods are improving, but systematic reviews still emphasize the need to benchmark signals against orthogonal chemistry, genetic perturbation of writer enzymes, and known standards.

Sequence, structure, and modification can conflict or cooperate. A protein may prefer a sequence motif only when it is single-stranded. A modification may weaken a helix and thereby increase motif exposure. A structure may place two weak motifs close enough for tandem domains to bind. Conversely, an RBP can protect a modification site from enzymes, remodel a structure so a motif disappears, or compete with another RBP that reads the same nucleotide patch. Specificity is therefore best described as a layered recognition grammar rather than as a single motif logo.

Box 56.2. Cautions for RNA Modification Reader Claims

A modified nucleotide can recruit a reader protein directly, alter local RNA structure to expose or hide a binding motif, block a competing RBP, change RNA decay rate, or bias sequencing detection chemistry. These mechanisms are distinct and require distinct evidence. A direct reader claim should ideally include measurement of modification stoichiometry, loss and rescue experiments using writer-enzyme perturbation or synthetic modified RNA, direct binding comparison of modified and unmodified RNA, and structural or chemical evidence that the protein contacts the modified chemical group rather than an indirectly altered conformation.

Concrete viral examples show this logic. Viral RNAs often contain structured untranslated regions, coding-region structures, long-range contacts, and host- or virus-encoded RBP binding sites. Next-generation sequencing approaches have broadened the study of viral RNA-protein interactions, but viral RNA abundance, replication intermediates, protein expression, and infection-stage mixtures complicate interpretation. A recent structural and mechanistic study of hnRNPA1 binding to hepatitis C virus RNA illustrates how a host RBP can interact with a defined viral RNA feature, but the biological meaning depends on viral life-cycle stage and competing host factors.

Figure 56.2. Layered Specificity of an RBP Binding Site

Figure 56.2. Layered Specificity of an RBP Binding Site. RBP specificity is filtered by several layers beyond intrinsic sequence affinity. The same sequence motif can be unavailable when buried in an RNA helix, accessible when the helix is remodeled, altered by a chemical modification, occluded by a competing protein or ribosome, or stabilized by neighboring motifs that allow tandem-domain binding. Cellular target selection is therefore a conditional property of motif sequence, RNA structure, modification state, localization, competition, and stoichiometry, not a simple readout of the sequence alone.

56.3. Cooperative and multivalent RNA-protein binding

Cooperative binding occurs when one binding event changes the likelihood of another. In RNA-protein recognition, cooperativity can be positive or negative. Positive cooperativity means that the first protein or domain makes later binding easier, perhaps by exposing a second motif, stabilizing an RNA conformation, or recruiting a partner protein. Negative cooperativity means that the first event blocks later binding, perhaps by covering an overlapping site, changing RNA structure, or using up a limiting scaffold. Both forms shape RNP assembly.

Multivalency is a common source of cooperativity. Many RBPs contain multiple folded RNA-binding domains plus intrinsically disordered regions. A single RRM may bind weakly, but two RRMs connected by a linker can bind adjacent motifs with much higher avidity. An arginine-rich region may make many transient contacts that increase residence time without specifying exact bases. A protein oligomer can present several RNA-binding surfaces. An RNA molecule with repeated motifs can recruit several proteins and create a local assembly platform. The result is often a sharp cellular response from individually weak interactions.

Multivalent recognition helps explain why short motifs are predictive but insufficient. A U-rich element, a stem-loop, or a modification site may be common across the transcriptome. The RBP still needs to select biologically relevant targets from many possible sites. Clustering of motifs, local structure, transcript abundance, RNA length, subcellular localization, and nearby protein partners can raise the effective affinity of one site over another. This is why in vitro motif discovery, CLIP peaks, and functional regulation often overlap imperfectly.

Cooperative RNP assembly can occur through RNA-induced protein proximity. An RNA containing two sites can bring proteins together and increase their local concentration. Those proteins may then interact directly, recruit an enzyme, block a nuclease, or remodel the RNA. Conversely, proteins can create RNA proximity by bridging distant segments, promoting RNA-RNA contacts, or concentrating RNAs in a granule. RNP granule studies show that competing protein-RNA interaction networks can organize multiphase intracellular compartments, but the material state of a granule does not by itself prove that every RNA inside is specifically regulated.

Stress-responsive assemblies illustrate positive and negative effects. G3BP-driven RNP granules can promote inhibitory RNA-RNA interactions, while the DDX3X helicase can resolve some of those interactions to regulate mRNA translatability. This example links multivalent RNP assembly to RNA remodeling: assembly can create a repressed state, and an ATP-dependent remodeler can return selected mRNAs to a more translatable state. The mechanism is not merely “condensation”; it is a sequence of RNA recruitment, multivalent contact formation, inhibitory RNA-RNA interaction, and helicase-mediated resolution.

Multivalency also creates hazards for interpretation. High local concentration can make weak nonspecific interactions visible in crosslinking or pull-down experiments. Disordered regions can bind many RNAs in vitro at salt conditions that do not match cells. Phase-separation assays can exaggerate assembly if proteins are overexpressed or purified at high concentration. A rigorous claim should specify whether multivalency changes affinity, specificity, kinetics, localization, material state, or biological output. Reviews of high-throughput RNA-protein interaction methods emphasize that cellular validation and orthogonal assays are needed before treating an interaction map as a regulatory network.

Cooperative binding is also central to immune and metabolic responses. The RBP RRP1 has been reported to brake macrophage one-carbon metabolism and suppress autoinflammation, illustrating how an RBP can connect RNA binding to cellular metabolism and inflammatory output. The broader lesson is that RBP function often emerges from networks of RNA targets and partner proteins, not from one binding event. However, disease or physiology claims require perturbation evidence because an RBP can bind many RNAs without regulating all of them.

56.4. RNP assembly pathways and remodeling

RNP assembly is the process by which RNA and proteins form a functional complex. The simplest assembly pathway is a protein binding a preformed RNA site. Many biological pathways are more ordered. A nascent RNA folds while being synthesized, one protein binds and stabilizes a local structure, another protein modifies or cleaves the RNA, a remodeling factor removes a temporary component, and the mature RNP acquires a different set of proteins. Assembly is therefore a kinetic pathway, not only an equilibrium endpoint.

Figure 56.3. Ordered RNP Assembly and Off-Pathway Branches

Figure 56.3. Ordered RNP Assembly and Off-Pathway Branches. Functional RNPs often assemble through ordered or alternative pathways rather than forming in a single equilibrium step. A nascent RNA folds, an early protein stabilizes a local structure, a modification or processing factor changes the substrate, an ATP-dependent remodeler removes a temporary contact, and a mature RNP forms. A side branch illustrates an off-pathway trap or damaged RNA-protein crosslink that must be remodeled or routed to surveillance and repair machinery.

The archaeal box C/D small RNP provides a compact example. Box C/D guide RNAs assemble with proteins to guide RNA modification. A reconstitution study showed that archaeal box C/D sRNP assembly can occur through alternative pathways and may require temperature-facilitated sRNA remodeling. The important principle is not only the specific archaeal temperature condition; it is that a small RNP can have more than one assembly route and that RNA remodeling can determine which route succeeds.

Ribosomal RNA assembly during transcription shows the large-scale version of the same principle. rRNA does not wait as a naked full-length transcript for ribosomal proteins to attach. rRNA folds co-transcriptionally, binds ribosomal proteins and assembly factors, receives chemical modifications, and undergoes cleavage and quality control. A roadmap for rRNA folding and assembly during transcription emphasizes that the order of transcription, local folding, protein recruitment, and processing decisions shape the final ribosome. This chapter treats those events as recognition and assembly logic; Chapter 42 and Chapter 43 treat ribosome biogenesis in greater detail.

Viral RNPs demonstrate another assembly logic. Influenza virus packages genomic RNA segments as ribonucleoprotein complexes with nucleoprotein and a viral polymerase. Structural and mechanistic studies of influenza RNP assembly and processive RNA synthesis show that viral RNA, nucleoprotein organization, and polymerase engagement must be coordinated for replication and transcription. Viral RNPs are especially useful examples because they must assemble in host cells, compete with host RBPs, and maintain specificity for viral genome segments under high RNA abundance and immune pressure.

Small-RNA RNPs add guide-strand logic. In miRNA and siRNA pathways, a small RNA is loaded into an Argonaute-family protein, one strand is retained as a guide, and target recognition occurs through base pairing rather than through a conventional protein domain reading a motif. The protein supplies structure, stability, slicing or repression capacity, and partner recruitment. The RNA supplies sequence specificity. Disease-focused reviews of miRNAs and signaling pathways illustrate the therapeutic interest in these assemblies, although a pathway-level review is not a substitute for direct assembly biochemistry.

RNP granules and germ granules show that assembly can be spatial as well as molecular. Rbm24a has been reported to dictate mRNA recruitment for germ granule assembly in zebrafish, connecting an RBP’s target selection to developmental RNP organization (ZhangRbm24a2025). In stress granules, G3BP-centered assemblies can alter RNA-RNA contacts and translation state. These systems are not simply bags of RNA and protein. They contain selective recruitment rules, multivalent contacts, remodelers, and exchange with the surrounding cytoplasm.

Remodeling is the process that changes an RNP after initial assembly. Remodeling can expose a cleavage site, remove an assembly factor, release a guide RNA, switch a transcript from repression to translation, or mark a damaged crosslink for repair. Formaldehyde-induced RNA-protein crosslinks can be marked by K6-linked ubiquitylation for resolution, illustrating that cells treat some RNA-protein covalent adducts as damage rather than as normal RNP states. This is a useful boundary case: not every stable RNA-protein association is a productive RNP.

The evidence standard for assembly pathways is higher than for pairwise binding. To claim an ordered pathway, one needs time-resolved intermediates, perturbation of proposed order, reconstitution, structural snapshots, or genetic epistasis. A static cryo-EM structure can show an assembled state, but it cannot by itself prove the route. CLIP can show that a protein contacts an RNA, but it cannot by itself prove when the contact forms. Assembly is a mechanism claim and should be supported by mechanism-level evidence.

Box 56.3. Minimal Evidence for an Ordered RNP Assembly Pathway

A pathway claim requires more than a mature complex structure. Stronger evidence includes time-resolved intermediates captured during assembly, in vitro reconstitution showing the pathway can proceed with defined components, ordered addition or sequential depletion of factors to establish dependency, remodeler perturbation that traps or bypasses a step, structural snapshots of intermediates, and functional rescue that restores the mature complex when a missing factor is restored. Static binding measurements and final-state structures establish components and interfaces but do not define the order of assembly or the existence of alternative routes.

56.5. RBP networks, competition, and stoichiometry

In cells, an RNA molecule is rarely alone with one RBP. A pre-mRNA can be contacted by spliceosomal factors, hnRNPs, SR proteins, cap-binding proteins, polyadenylation factors, export adaptors, modification enzymes, surveillance factors, and chromatin-associated proteins. A mature mRNA can be contacted by translation factors, ribosomes, decay enzymes, miRNA-associated complexes, localization factors, and stress granule proteins. These factors form an RBP network: a set of direct and indirect interactions that determine RNA fate.

Figure 56.4. From Pairwise Binding to RBP Networks

Figure 56.4. From Pairwise Binding to RBP Networks. A single transcript can carry mutually exclusive, cooperative, and transient RBP contacts simultaneously. Network-scale methods extend the view from one-protein-one-target pairs to assemblies, RNA-dependent protein networks, and single-cell interaction states. Functional interpretation of these networks still requires quantitative occupancy measurements and perturbation experiments to distinguish binding from regulation.

Competition is unavoidable because many RBPs prefer similar RNA features. U-rich regions, AU-rich elements, structured loops, poly(A) tails, cap-proximal regions, splice-site neighborhoods, and coding-region motifs can attract multiple factors. Competition can be steric, when two proteins cannot occupy overlapping nucleotides. It can be allosteric, when one protein changes RNA structure and removes another site. It can be kinetic, when the first protein to bind creates a long-lived state. It can be stoichiometric, when a limiting protein is titrated by abundant decoy RNAs.

Stoichiometry matters because binding depends on concentration as well as affinity. If a transcript has many weak sites and an RBP is abundant, cumulative occupancy can be high. If a high-affinity site is present on a rare transcript but the RBP is sequestered in a granule or nucleus, occupancy can be low. If a repeat expansion or viral RNA produces many binding sites, it can titrate an RBP away from normal targets. If stress changes RBP phosphorylation, localization, or modification, the effective concentration of active RBP can shift without any change in RNA sequence.

Network-scale methods try to measure these layers. RNA-dependent interactome analysis can assign RBP function by identifying protein networks that depend on RNA, not just individual RNA targets. RNP-MaP uses mutational profiling logic to study RNA-protein interaction networks and can support hypotheses about coordinated occupancy. Single-cell methods can co-profile in situ RNA-protein interactions and transcriptomes, adding cell-state context that bulk assays average away. SPIDR enables multiplexed mapping of RNA-protein interactions and has been used to uncover selective translational suppression under cell stress. irCLIP-RNP and Re-CLIP reveal dynamic protein assemblies on RNA rather than only individual RBP sites.

These methods make an important conceptual change. A transcript is not merely a list of independent binding peaks. It is an RNP state that may contain mutually exclusive proteins, cooperative clusters, transient processing factors, stable structural proteins, ribosomes, and decay factors. The relevant biological question is often not “does this protein ever bind this RNA?” but “under this condition, what fraction of RNA molecules carries this RNP state, how long does the state persist, and what output follows?”

Network claims are vulnerable to indirectness. A protein may appear in an RNA-dependent interactome because it binds another RBP that binds RNA. A CLIP peak may reflect crosslinkability rather than high occupancy. A stress-induced interaction may reflect global RNA abundance changes rather than specific recruitment. A network edge may be condition-specific and absent in another cell type. For this reason, network-level studies should be interpreted alongside perturbation, rescue, quantitative stoichiometry, and orthogonal assays.

Cancer and DNA-damage contexts illustrate why network thinking matters. RBPs participate in DNA damage responses, radiation responses, splicing changes, RNA stability changes, and stress programs in cancer cells. A cancer phenotype may not result from a single lost binding site. It may result from altered expression of an RBP, changed phosphorylation, changed RNA modification, alternative polyadenylation, splicing shifts, and competition among many RNAs. Functional interpretation must therefore separate binding, regulation, and phenotypic causality.

56.6. Disease mutations and misregulation

RBP-related disease mechanisms fall into several classes. A mutation in an RBP can weaken binding to normal RNA targets. Another mutation can create a new preference and redirect the protein to abnormal targets. A mutation in a low-complexity region can alter localization, granule dynamics, aggregation, or partner recruitment. A mutation in an RNA can create or destroy an RBP motif. A change in RNA modification can change reader binding. A change in RBP abundance can shift network competition. These mechanisms can coexist in the same disease.

Mutations in RNA regulatory regions are especially difficult to interpret because one nucleotide can affect several layers. A noncoding mutation can alter a splice motif, an RBP binding motif, RNA secondary structure, miRNA targeting, polyadenylation, transcription-factor binding at the DNA level, or RNA modification. Reviews of noncoding RNA mutations in cancer emphasize that regulatory mutation claims require careful separation of these mechanisms. A mutation-induced RNA structure change can be modeled and visualized computationally, as in MutaRNA, but computation should be treated as hypothesis generation unless paired with functional evidence.

RBP misregulation also occurs without mutation. Stress can relocalize RBPs into granules, infection can redirect RBPs to viral RNAs, inflammation can change RNA metabolism, and cancer signaling can alter RBP expression. Editorial and review literature on post-transcriptional regulation and misregulation frames these events as translationally important because altered RNA processing, stability, localization, and translation can drive disease phenotypes. However, clinical relevance requires evidence beyond an association between an RBP and a disease state.

Neurodegeneration is a major RBP disease area, with TDP-43, FUS, hnRNPA1, repeat-containing RNAs, and stress granule biology often discussed together. The local Chapter 56 bibliography does not yet contain dedicated neurodegeneration references, so this draft does not use those examples as claim-level provenance. Final-reference expansion notes: add targeted reviews and primary papers for TDP-43, FUS, hnRNPA1, FMRP, repeat-expansion RNAs, and ALS/frontotemporal dementia before final disease expansion.

Viral infection creates a different disease context. RNA viruses mutate rapidly, generate structured replication intermediates, and recruit or antagonize host RBPs. Classic work on RNA virus mutation rates provides a background reason why viral RNA-protein interfaces can change under selection. Studies of viral RNA-protein interactions and chikungunya virus protein-targeting drugs illustrate how viral RNA and viral proteins become both biological mechanisms and therapeutic targets. A specific antiviral claim still needs virus-specific evidence because RNP architecture differs among positive-strand RNA viruses, negative-strand RNA viruses, retroviruses, and segmented viruses.

Therapeutic correction can target RNA outcomes rather than the original DNA mutation. CRISPR-free RNA base editing has been reported to mediate premature-termination-codon readthrough and restore hearing in mice with an Otof nonsense mutation. This example belongs at the edge of this chapter because the therapeutic strategy changes RNA sequence interpretation and protein production, but it is not primarily a natural RBP specificity mechanism. The broader point is that RNA-level interventions must account for RBP binding, RNA structure, modification, innate sensing, and decay.

The main caution is that binding is not disease mechanism by itself. A disease-associated RBP may bind thousands of RNAs, but only a subset may drive pathology. A mutation may change a CLIP peak without changing RNA abundance or protein output. An RBP aggregate may be a cause, a consequence, or a buffering response. Mechanistic disease claims should specify the altered molecule, the changed interaction, the RNA fate consequence, the cell-type context, the phenotype, and the evidence connecting these steps.

56.7. CLIP, biochemical, and structural evidence

RNA-protein recognition is supported by several evidence classes, each with a distinct strength. Biochemical assays such as electrophoretic mobility-shift assays, filter binding, fluorescence anisotropy, isothermal titration calorimetry, surface plasmon resonance, microscale thermophoresis, nuclease footprinting, and competition binding can measure affinity, stoichiometry, kinetics, and sequence dependence under controlled conditions. Their weakness is context: purified proteins and RNAs may lack modifications, cofactors, competitors, folding history, and cellular compartmentalization. This chapter owns the biological interpretation of recognition and assembly; Chapter 124 owns the reusable measurement models, concentration regimes, active-fraction controls, fitting assumptions, and artifacts of the quantitative assays themselves.

Structural methods show interfaces. X-ray crystallography, nuclear magnetic resonance spectroscopy, cryo-electron microscopy, crosslinking-mass spectrometry, chemical probing, and integrative modeling can identify contacts between bases, sugars, phosphates, protein side chains, metal ions, and partner proteins. Structural evidence is powerful when the sample represents the functional state. It is weaker when a structure uses truncated constructs, engineered RNAs, high-affinity variants, nonphysiological salts, or static snapshots of a dynamic pathway. Reviews of RNA-protein structural modeling emphasize that flexible RNA, transient complexes, and incomplete restraints remain difficult.

CLIP-based methods map RNA regions associated with a protein in cells. In a typical CLIP workflow, cells or tissues are crosslinked, an RBP is immunoprecipitated, RNA fragments are recovered, libraries are sequenced, and peaks or crosslink-induced signatures are identified. seCLIP-seq provides a standardized protocol framework for transcriptome-wide binding-site identification. Recent reviews of CLIP methodologies emphasize that crosslink chemistry, nuclease digestion, antibody quality, library complexity, biological replicates, controls, and computational peak calling all influence the result.

CLIP evidence should be interpreted as protocol-conditioned evidence of proximity and recovery. A peak near a splice site can support RBP association with that region, but it does not automatically prove direct base recognition, regulatory function, or occupancy of every transcript molecule. UV crosslinking favors certain amino acid and nucleotide chemistries. Formaldehyde can capture indirect contacts and can create damage-like crosslinks that cells may resolve through repair pathways. Immunoprecipitation can enrich abundant RNAs or sticky complexes. RNase digestion can move apparent peak boundaries. These limitations do not make CLIP unreliable; they define the controls needed for strong inference.

Newer methods extend beyond one RBP at a time. Single-cell RBP target discovery can connect interactions to cell-state heterogeneity. Co-profiling in situ RNA-protein interactions and transcriptomes can preserve tissue and single-cell context. SPIDR enables multiplexed mapping and can connect interaction patterns to translational suppression under stress. irCLIP-RNP and Re-CLIP reveal patterns of dynamic protein assemblies on RNA, a step toward measuring RNP states rather than isolated binding events. RNA-centric and network-centric methods, including RNP-MaP and RNA-dependent interactome analysis, help identify coordinated assemblies.

The best-supported recognition claims combine methods. A strong motif-specific claim might include in vitro binding with base substitutions, CLIP showing cellular occupancy, structural data showing contacts, reporter assays showing regulation, endogenous editing or mutation showing phenotype, and rescue by restoring the interaction. A strong assembly-pathway claim might include time-resolved intermediates, reconstitution, structural snapshots, remodeler perturbation, and functional output. A strong disease claim might add patient or model-system genetics, cell-type specificity, and correction of the altered RNA fate.

Artifact control is not a separate afterthought. It is part of the biological claim. If a protein binds RNA only after cell lysis, the result may reveal an intrinsic affinity but not a cellular interaction. If a crosslink peak disappears after protein depletion, that supports specificity, but depletion can indirectly change RNA abundance. If a structure is solved with a short RNA motif, the structure may miss flanking elements that determine cellular selectivity. If a modification-reader interaction is detected with an antibody-enriched RNA sample, modification validation is required. Evidence should be reported with the exact assay, condition, and inference level.

Recent Consensus

The current consensus is that RNA-protein recognition is combinatorial. Canonical domains such as RRMs, KH domains, zinc fingers, and dsRBDs provide recurring binding modes, but cellular specificity arises from domain combinations, RNA structure, chemical modification, timing, localization, concentration, and competition.

The field has moved away from a strict motif-only model of RBP targeting. Sequence motifs remain useful, but many RBPs bind conditional sites whose accessibility depends on folding, RNP state, modification, ribosome movement, and neighboring proteins. Modification-specific recognition is accepted as a major principle, but individual sites and readers require careful validation because modification-detection technologies remain method-dependent.

RNP assembly is now treated as a pathway with intermediates, not a static final complex. This is established for large systems such as ribosome biogenesis and viral RNPs, and it is increasingly measurable for smaller and more dynamic assemblies through reconstitution, structural snapshots, and interaction-mapping methods.

Transcriptome-wide and single-cell methods have expanded the evidence base from isolated interactions to RBP networks. The consensus is cautious: maps are powerful discovery tools, but functional claims require perturbation, quantitative controls, and orthogonal evidence.

Open Questions, Controversies, Deprecated Models, and Common Misconceptions

Open questions:

  • How much of cellular RBP specificity is encoded by intrinsic domain-RNA affinity, and how much is imposed by transcript age, localization, stoichiometry, and competition? The answer probably differs among splicing factors, translation regulators, viral proteins, structural RNP proteins, and granule-associated RBPs.
  • Which RNA modifications are directly read by proteins, and which mainly act by changing RNA structure, stability, or competition? This distinction remains difficult when modification stoichiometry is low, detection methods disagree, or modification enzymes have indirect effects.
  • Can high-throughput interaction maps be converted into quantitative occupancy models? A peak or interaction edge is not the same as molecules-bound-per-cell, residence time, or regulatory strength.
  • How often do low-complexity RBP regions contribute specific recognition rather than nonspecific affinity, localization, or granule partitioning? Multivalent and disordered interactions are biologically important but can be overinterpreted in simplified in vitro systems.

Controversies:

  • Controversy: Some RBP network diagrams imply causal regulation from co-occurrence or crosslinking. Such diagrams are useful hypotheses but can overstate mechanism unless supported by perturbation, rescue, and quantitative RNA fate measurements.

Common misconceptions:

  • “An RBP motif predicts binding.” A motif predicts a possible binding opportunity. Cellular binding also depends on structure, modification, localization, expression, competition, and protein state.
  • “CLIP peaks are functional binding sites.” CLIP peaks are evidence of recovered RNA-protein proximity under a protocol. Functional binding requires additional evidence.
  • “Specificity means high affinity.” Specificity is preference among alternatives. A high-affinity interaction can be nonspecific, and a lower-affinity interaction can be specific in the right cellular context.
  • “A granule-localized RNA is necessarily regulated by every granule protein.” Granule localization can reflect recruitment, retention, passive partitioning, or stress-induced crowding. Regulation must be shown for the specific RNA-protein pair and output.