Chapter 138. RNA-Centric Proteomics, RNA-Capture Chemistry, and RNP Composition Analysis

Scope Note

RNA-centric proteomics asks which proteins are associated with a defined RNA molecule, RNA class, RNA state, or RNA-containing assembly under a specified biological condition. The question is deliberately broader than direct RNA binding. A recovered protein may contact the RNA surface, bind a partner protein in the same ribonucleoprotein (RNP), share a condensate or chromatin neighborhood with the RNA, recognize a chemical modification, interact with a capture reagent, or appear because of abundance and nonspecific adhesion. This chapter explains how RNA pulldown, endogenous antisense capture, aptamer tagging, proximity labeling, crosslinking chemistry, peptide- and protein-centric mass spectrometry, contamination control, and orthogonal validation are combined to move from a proteomic hit list toward a defensible model of RNP composition and function. It owns mass-spectrometric analysis of the protein material recovered by RNA capture. Mass spectrometry of nucleosides, oligonucleotides, and intact RNA belongs to Chapter 132, whereas native mass spectrometry of intact RNP composition and stoichiometry belongs to Chapter 59.

Executive Summary

RNA-centric proteomics reverses the usual direction of protein-centric RNA biology. Instead of starting with a protein and asking which RNAs it binds, the experiment starts with an RNA target or RNA-defined state and asks which proteins are enriched with that material. This direction can discover unexpected proteins associated with lncRNAs, viral genomes, localized mRNAs, repeat RNAs, guide RNAs, and synthetic therapeutic RNAs. It is also vulnerable to overinterpretation because enrichment is not equivalent to direct binding, specificity, stoichiometry, or function.

The chapter’s central distinction is between four experimental meanings. In vitro RNA pulldown tests association with a prepared RNA bait. Endogenous antisense capture enriches cellular RNA material through complementary oligonucleotides. Aptamer tagging adds a recoverable or recruitable RNA element that can support purification, imaging, or proximity labeling but can perturb RNA metabolism. Proximity labeling records proteins near an RNA-recruited labeler during a defined window; it is a neighborhood assay, not a direct-contact assay.

Crosslinking chemistry determines what can survive capture and therefore what the experiment can claim. UV irradiation favors direct nucleic acid-protein contacts but is inefficient and biased by nucleotide, amino acid, and geometry. Formaldehyde and related chemical crosslinkers preserve wider molecular neighborhoods, including protein-protein and chromatin-associated connections, but they blur directness. Photoactivatable nucleotides, diazirines, aryl azides, benzophenones, psoralen-like reagents, and clickable handles add specialized selectivity at the cost of perturbation and chemistry-specific blind spots. Native workflows preserve labile complexes but admit post-lysis reassortment and indirect co-purification; denaturing workflows require covalent stabilization and reduce many indirect contaminants but can lose native stoichiometry.

Quantitative proteomics is not optional. RNA captures almost always recover background proteins, including abundant RBPs, ribosomal proteins, cytoskeletal proteins, metabolic enzymes, chaperones, bead binders, probe binders, stress-response proteins, and compartmental bystanders. Replicates, effect sizes, missing-value handling, peptide evidence, isotope or isobaric labeling, label-free quantification, spike-ins, RNase tests, mutant RNAs, localization-matched controls, reagent-only controls, and recurrent-contaminant filtering determine whether a hit list becomes a credible candidate list. Condensate proteomics is especially difficult because scaffolds, clients, regulators, passengers, and purification-induced aggregates can all be enriched from the same sample.

The usual analytical endpoint is bottom-up proteomics. Proteins released from the captured RNA material are denatured, reduced, alkylated, digested into peptides, separated by liquid chromatography, ionized, and analyzed by tandem mass spectrometry. Peptide spectra are matched against a specified sequence database, and proteins are inferred from unique and shared peptides. This workflow is sensitive and scalable, but digestion breaks the physical link between peptides from the same intact proteoform and destroys any record of which proteins occupied the same individual RNA molecule. Label-free, metabolic-label, and isobaric-label strategies then estimate relative enrichment, each with distinct mixing points, missingness patterns, and interference risks. Absolute RNP stoichiometry requires calibration and an RNA-molecule denominator that ordinary discovery capture rarely supplies.

A robust RNP model integrates RNA-centric proteomics with orthogonal evidence. CLIP or related protein-centric methods can localize binding sites. Imaging can test spatial overlap and state specificity. Mutating the RNA, the protein RNA-binding domain, the localization signal, or the condensate scaffold can test dependence. Purified-component binding assays can test direct interaction and affinity. Functional perturbation and rescue can test biological consequence. The most defensible language is calibrated to the evidence: “protein P is enriched with captured RNA R after formaldehyde crosslinking” is different from “protein P directly binds RNA R at motif M and controls phenotype X.”

Concept Inventory

  • RNA-centric proteomics: a method family in which an RNA target, RNA class, RNA state, or RNA-proximal label is used to enrich proteins for mass spectrometry or targeted validation.
  • Ribonucleoprotein (RNP): an assembly containing RNA and protein. RNPs range from stable machines such as ribosomes to dynamic messenger RNPs, viral replication complexes, chromatin-associated lncRNA assemblies, and RNA-rich condensates.
  • In vitro RNA pulldown: an affinity experiment in which a prepared RNA bait is immobilized or captured and incubated with purified proteins, extracts, or lysates.
  • Endogenous antisense capture: enrichment of cellular RNA material through hybridization between the target RNA and complementary probes, often biotinylated oligonucleotides recovered on streptavidin supports.
  • Aptamer tagging: insertion or attachment of an RNA sequence that binds a known ligand or protein, such as MS2, PP7, BoxB, streptavidin-binding aptamers, fluorogenic aptamers, or small-molecule aptamers.
  • Proximity labeling: enzymatic or chemical labeling of proteins near a recruited enzyme or reactive handle, followed by enrichment of the label, commonly biotin.
  • Crosslinking: covalent stabilization of nearby molecules before or during capture. Crosslinking can favor direct contacts or broader neighborhoods depending on chemistry.
  • Direct RNA binder: a protein that physically contacts RNA through a defined or nonspecific RNA-binding surface. Direct binding is not automatically functional regulation.
  • Indirect RNP component: a protein recovered because it binds another protein, DNA, membrane, condensate scaffold, or RNA-associated compartment rather than the RNA surface itself.
  • Quantitative enrichment: reproducible abundance increase in target capture relative to matched controls, interpreted with effect size, replicate structure, peptide evidence, and background behavior.
  • Bottom-up proteomics: analysis in which recovered proteins are digested into peptides before liquid chromatography-tandem mass spectrometry; protein identities and abundances are inferred from peptide evidence rather than observed as intact molecules.
  • Peptide-spectrum match: assignment of an observed tandem mass spectrum to a candidate peptide sequence under a defined search space, scoring model, and false-discovery procedure.
  • Protein inference: the process of mapping identified peptides to proteins or protein groups, including explicit treatment of peptides shared among paralogs, isoforms, and proteoforms.
  • Proteoform: a molecular form of a protein produced by sequence variation, RNA processing, proteolysis, or post-translational modification. Bottom-up evidence often cannot reconstruct the complete proteoform carried by an RNP.
  • RNP stoichiometry: the number of copies of each component per intact assembly or per RNA molecule. Relative peptide enrichment does not by itself provide this quantity.
  • Contaminant: a recovered protein whose apparent enrichment is better explained by reagent binding, abundance, sample handling, post-lysis reassortment, stress, or compartmental co-purification than by target-specific RNA association.
  • Claim calibration: matching the wording of a biological claim to the highest level of evidence actually obtained.

What to Know Before Reading This Chapter

The reader should start from the idea that cellular RNA is rarely naked. Nascent transcripts are coated by processing factors. mRNAs cycle among export, translation, localization, storage, surveillance, and decay states. Small RNAs function inside Argonaute, PIWI, CRISPR-Cas, spliceosomal, or ribonuclease complexes. Viral RNAs assemble replication and packaging machines. Nuclear lncRNAs can remain near chromatin, nucleoli, speckles, paraspeckles, or other nuclear domains. A protein can therefore be associated with an RNA because it binds the RNA directly, because it binds another RNP component, because it is in the same cellular body, or because a purification protocol created a new association after lysis.

Affinity purification is the technical background. A bait is captured, co-purifying material is washed, and recovered proteins are identified by mass spectrometry or targeted assays. The proteomic output is not a simple membership list. Peptide detection depends on protein abundance, digestion, ionization, instrument duty cycle, database search settings, sequence uniqueness, and missing values. Biological interpretation depends on controls. A protein detected in a target capture but absent from a control is a candidate; a protein reproducibly enriched across independent conditions, lost after RNA-domain mutation, supported by binding-site evidence, and required for an RNA-dependent phenotype is a much stronger mechanistic actor.

Readers should also distinguish the analyte from the bait. In this chapter the RNA defines what is captured, but proteins and their peptides are the mass-spectrometric analytes. A nucleoside or oligonucleotide spectrum asks what the RNA is chemically made of and is treated in Chapter 132. A native mass spectrum of an intact RNP asks which assembly masses, ligand states, or component stoichiometries survive transfer into the gas phase and is treated in Chapter 59. A peptide spectrum after RNA capture instead supports a sequence assignment and, after aggregation and comparison with controls, a protein-enrichment estimate.

A running example throughout the chapter is a nuclear lncRNA suspected to organize a chromatin-associated RNP. An antisense capture experiment may recover the lncRNA, chromatin regulators, hnRNP proteins, nucleolar proteins, ribosomal proteins, DNA-binding factors, and stress-response proteins. Some may be true lncRNA partners. Others may be abundant nuclear background, chromatin hitchhikers, probe binders, proteins crosslinked indirectly through DNA, or factors associated with only a small subpopulation of the lncRNA. The experiment becomes informative when RNA recovery is measured, off-target RNAs are monitored, matched controls are quantified, RNase and mutation tests are used, and candidate proteins are followed by independent binding, imaging, and perturbation evidence.

138.1. RNA pulldown and antisense capture

RNA pulldown and antisense capture are the two classic entry points for RNA-centric proteomics. They share a capture logic: base pairing or affinity chemistry enriches an RNA and the proteins that remain associated with it. They differ in the starting state. In vitro RNA pulldown begins with a prepared RNA bait outside its native cellular context. Endogenous antisense capture begins with the RNA inside cells or lysate and uses complementary probes to retrieve it. The difference seems simple, but it controls almost every interpretation that follows.

In vitro RNA pulldown is often the most direct way to ask whether a sequence element, RNA structure, modification, or mutation can recruit proteins. The RNA bait may be chemically synthesized, transcribed in vitro, folded under defined salt and temperature conditions, modified enzymatically or chemically, biotinylated at an end, internally labeled, or fused to an affinity handle. The bait is incubated with purified proteins, cell lysate, nuclear extract, cytoplasmic extract, ribosomal fractions, or other biochemical preparations. Proteins retained after washing are eluted for immunoblotting, silver staining, mass spectrometry, or targeted proteomics.

The strength of this workflow is control. A researcher can compare wild-type and mutant RNA, full-length and domain-level RNA, folded and heat-denatured RNA, modified and unmodified RNA, sense and antisense RNA, or disease-associated and reference alleles. A riboswitch aptamer, a viral untranslated-region element, a repeat-expansion RNA, a splicing enhancer, or a synthetic mRNA untranslated region can be tested as a discrete bait. If a purified protein binds the wild-type RNA but not a motif-mutant RNA under physiological salt, the result supports direct biochemical specificity more strongly than a discovery capture from crude lysate.

The weakness of in vitro pulldown is context loss. An RNA made by in vitro transcription may lack native base modifications, co-transcriptional folding history, ribonucleoprotein maturation, compartment-specific cofactors, and normal competition from the transcriptome. A sequence that is exposed in vitro may be buried in a cellular RNP. Conversely, a sequence that is hidden in vitro may be exposed when the RNA folds co-transcriptionally or when a helicase remodels it. Proteins in lysate can reassort onto naked RNA after cell disruption, especially abundant basic proteins, helicases, ribosomal proteins, and low-complexity RBPs. Therefore an in vitro pulldown can demonstrate binding potential, not automatically cellular occupancy.

Endogenous antisense capture tries to keep the RNA closer to its biological state. Biotinylated DNA, locked nucleic acid, or other modified oligonucleotide probes hybridize to accessible segments of the target RNA. Streptavidin beads then recover the probe-bound RNA and associated molecules. If crosslinking occurs in cells before lysis, endogenous capture can preserve low-abundance or fragile associations that might otherwise dissociate. Hybridization capture has been adapted from RNA localization and RNA-chromatin methods to proteomics by replacing or adding a mass-spectrometry readout. For nuclear lncRNAs, viral RNAs, abundant structured RNAs, and chromatin-associated transcripts, this strategy often asks a better biological question than a naked in vitro bait.

Endogenous capture is not automatically physiological. Probe binding depends on target accessibility, which is shaped by RNA secondary structure, protein occupancy, transcript fragmentation, chemical modifications, and isoform choice. A probe set tiled across a transcript can improve recovery, but every probe adds possible off-target hybridization and probe-specific protein background. Repetitive RNAs, homologous gene families, antisense transcripts, pseudogenes, and transcript isoforms are especially vulnerable to ambiguous capture. A target with low copy number may require large input, strong crosslinking, or aggressive amplification of the proteomic signal; those choices can increase background.

Good capture design starts with the RNA model. The investigator should define which isoforms are meant to be captured, whether the RNA is mostly nuclear or cytoplasmic, whether it is translated, whether it is chromatin-associated, whether repeats or paralogs complicate probe specificity, and whether a subregion or the full transcript is the biological target. Probe sequences should avoid highly repetitive or homologous regions when specificity is the priority, but they should also cover regions likely to be accessible. Capture efficiency should be measured by quantitative reverse transcription PCR, targeted sequencing, or another RNA recovery assay. Specificity should be checked against related RNAs, abundant RNAs, compartment-matched RNAs, and negative-control regions.

The most important interpretive rule is that protein data should not outrun RNA data. A long protein table is weak if the target RNA was poorly enriched, if off-target RNAs were co-captured, or if capture efficiency varied strongly across replicates. The recovered target fraction also matters. If only a small fraction of the target RNA is recovered, the proteome may describe an accessible or crosslinked subpopulation rather than the whole RNA. Fragment size matters as well: short fragments can provide local neighborhood information, whereas long fragments can recover larger assemblies but blur positional resolution.

Figure 138.1. Four entry points into RNA-centric proteomics

Figure 138.1. Four entry points into RNA-centric proteomics. In vitro RNA pulldown tests association with a prepared RNA bait; endogenous antisense capture enriches native RNA material with complementary probes; aptamer tagging adds a recoverable or recruitable module; proximity labeling marks proteins near an RNA-recruited enzyme or chemical handle. The visual should emphasize that each route supports a different claim and creates a different background.

Several controls are essential. A scrambled or sense-probe control estimates probe and bead background. An unrelated RNA with similar abundance and localization estimates compartment background. A mutant RNA or deleted domain tests sequence or structural specificity when genetic manipulation is possible. RNase treatment can test RNA dependence, but it should be interpreted carefully because RNase can disrupt an RNP indirectly or expose new sticky surfaces. Competition with excess unlabeled probe or RNA can help distinguish specific hybridization from nonspecific capture. Spike-in RNAs or exogenous standards can monitor technical recovery. None of these controls is sufficient alone; together they convert a discovery experiment into a controlled measurement.

Box 138.1. Control Checklist for lncRNA Antisense Capture

  • Box body (render-ready Markdown):

Before interpreting a lncRNA capture proteome, check five measurements. First, confirm that the intended lncRNA isoform or region was recovered reproducibly, not merely detected. Second, measure plausible off-target RNAs, including antisense transcripts, paralogs, repeats, abundant nuclear RNAs, and neighboring chromatin-associated transcripts. Third, compare the target capture with reagent controls such as beads-only, probe-only, scrambled-probe, or sense-probe samples. Fourth, include a localization-matched RNA control when the target is nuclear, chromatin-associated, nucleolar, or condensate-enriched. Fifth, test RNA dependence with RNase, domain deletion, or sequence mutation while verifying that the perturbation actually changed the RNA material being captured. A protein that survives this checklist is still a candidate association, not automatically a direct binder, but the candidate is now worth mechanistic follow-up.

For the running lncRNA example, an antisense capture that recovers chromatin modifiers is promising but not conclusive. The next questions are whether the lncRNA itself is enriched, whether neighboring genomic DNA or nascent chromatin fragments are co-purified, whether an unrelated nuclear lncRNA produces the same proteome, whether RNase or domain deletion removes candidate proteins, and whether candidate proteins show independent evidence of binding or colocalization. The strongest conclusion available from the capture alone is an association under the capture conditions. Direct binding and function require further evidence.

Table 138.1. What each RNA capture strategy can claim. Protein-centric, RNA-centric, proximity, crosslinking, and affinity capture strategies support different association or contact claims; bait behavior, abundance, localization, and background determine the required controls.

Strategy Primary bait or label Best-supported claim Main artifact risk Essential controls
In vitro RNA pulldown Prepared RNA bait, often biotinylated or affinity-tagged A protein associates with the bait under defined assay conditions Nonnative folding, missing modifications, and post-lysis reassortment onto naked RNA Mutant RNA, unrelated RNA, beads-only control, folded-state control, purified reconstitution
Endogenous antisense capture Hybridization probes against the target RNA A protein is enriched with captured endogenous RNA material Off-target RNA capture, inaccessible probes, and probe-specific protein background Scrambled or sense probes, target RNA recovery, off-target RNA checks, localization-matched RNA
Aptamer capture Engineered RNA tag such as MS2, PP7, BoxB, or affinity aptamer A protein associates with the tagged RNA state Tag-induced changes in RNA abundance, folding, localization, translation, or decay Untagged RNA, tag-only control, endogenous expression or titration, functional rescue
Proximity labeling RNA-recruited enzyme or reactive handle A protein lies near the RNA during the labeling window Compartmental bystanders and broad labeling radius Enzyme-only control, localization-matched RNA, substrate or pulse controls, time course
Crosslinked denaturing capture Covalently stabilized RNA-protein material or chemical handle A protein survives stringent capture because of stabilized contact or proximity Crosslinking bias, low efficiency, and loss of native stoichiometry No-crosslink control, chemistry controls, crosslink reversal checks, RNA recovery metrics

138.2. Aptamer tagging and proximity approaches

Aptamer tagging changes the capture problem by adding a known interaction module to the RNA. The tag may bind a protein, a small molecule, a dye, streptavidin, or another affinity reagent. Bacteriophage-derived hairpin systems such as MS2 and PP7 recruit cognate coat proteins that can be fused to fluorescent proteins, affinity handles, enzymes, or other domains. BoxB-lambdaN systems use a different hairpin-protein pair. Other aptamers bind small molecules or engineered proteins. The tag can make an RNA visible, recoverable, recruitable, or chemically labelable.

The attraction is obvious. Many endogenous RNAs are hard to capture because they are scarce, structured, repetitive, poorly accessible, or transiently localized. A tag can provide a strong, standardized handle. Repeated hairpins can increase signal. A coat protein fused to GFP can support imaging; a coat protein fused to an affinity tag can support purification; a coat protein fused to APEX, BioID, TurboID, or another enzyme can support proximity labeling. A tagged reporter mRNA can test how a 3′ UTR recruits localization factors, decay machinery, translational repressors, or granule components. A tagged lncRNA can test which proteins travel with a defined transcript domain.

The tag is also part of the RNA. It adds sequence, structure, mass, and binding sites. Repeated hairpins can change transcription elongation, splicing, export, localization, translation, decay, and folding. A tag in a 5′ untranslated region may interfere with scanning; a tag in a coding sequence may change the peptide; a tag in an intron may alter splicing; a tag in a 3′ untranslated region may change localization or miRNA regulation; a terminal tag may affect exonucleolytic decay. A small RNA or guide RNA can be dominated by the tag. Even a long lncRNA can be altered if the tag changes folding, protein occupancy, or nuclear retention.

The correct control question is not simply whether the tag works. The question is whether the tagged RNA behaves like the untagged RNA for the biological state being studied. Expression level should approximate the endogenous level when endogenous biology is being claimed. Processing, isoform structure, localization, half-life, translation status, and function should be measured when relevant. Genomic insertion at the endogenous locus is usually more physiological than plasmid overexpression, but it is not a guarantee. A genomic tag can still disrupt regulatory elements, change RNA folding, or create new protein-binding surfaces. An overexpressed tagged RNA can saturate RBPs or nucleate nonphysiological condensates.

Proximity labeling extends the tag concept from capture to neighborhood measurement. A recruited enzyme converts nearby proteins into enrichable products, commonly by biotinylation. APEX-type peroxidases generate short-lived radicals under supplied substrate conditions and can label proteins over short time windows. BioID and TurboID-type biotin ligases label proximal proteins through activated biotin intermediates over longer or enzyme-dependent windows. Other proximity strategies use chemically reactive probes, guide-directed recruitment, or split systems that assemble activity only when components coincide.

Proximity labeling is valuable because many RNA-associated states are weak, transient, insoluble, membrane-adjacent, or disrupted by lysis. Proteins near a localized mRNA at mitochondria, a viral replication compartment, a stress granule, a nuclear speckle, a paraspeckle, or a translation site may be difficult to recover by conventional pulldown. Labeling them in living cells can preserve a record of spatial exposure before extraction. The output, however, is a proximity proteome. A labeled protein may bind the RNA, bind a protein on the RNA, pass through the same compartment, or merely be reactive and nearby during the labeling pulse.

Box 138.2. Proximity Labeling Is a Neighborhood Assay

  • Box body (render-ready Markdown):

Interpret a proximity hit by asking what the labeler could reach. The answer depends on enzyme type, tag position, expression level, localization, substrate concentration, pulse length, and the local environment. A rapidly labeled protein near a tagged mRNA at a membrane may be a direct RBP, a translation factor, a membrane protein, a trafficking factor, or a bystander that entered the same nanoscale neighborhood during the pulse. Longer labeling windows increase sensitivity but also integrate movement through a compartment. The strongest proximity experiments therefore include enzyme-only and tag-only controls, substrate-free or inactive-enzyme controls when available, localization-matched RNAs, and time courses. Proximity enrichment is excellent evidence for spatial exposure under the tested condition. Direct binding, binding site, stoichiometry, and functional requirement need CLIP, reconstitution, imaging, mutational, or perturbation evidence.

This distinction should shape experimental design. A proximity experiment needs an enzyme-only control, a tag-only control, a catalytically inactive or substrate-free control when applicable, and a localization-matched control. If an RNA is targeted to mitochondria, the negative control should test the same organelle background. If a tag is overexpressed, expression-matched controls are needed because enzyme abundance and localization affect labeling radius and background. Time-course and substrate-dose experiments can distinguish rapid local labeling from slow accumulation of broad compartment signal.

Figure 138.2. Aptamer and proximity-labeling architectures

Figure 138.2. Aptamer and proximity-labeling architectures. A tagged RNA can recruit a coat protein, affinity handle, fluorescent protein, or labeling enzyme. Proximity labeling records proteins exposed to the labeler during a time window, whereas affinity capture recovers material that remains associated through extraction. The figure should show tag-only, enzyme-only, substrate-free, and localization-matched controls.

The aptamer and proximity strategies are especially useful for RNA state transitions. A reporter mRNA can be tagged and followed during stress, translational repression, localization, or decay. Candidate proteins can be identified at early and late time points. A lncRNA domain can be tagged to ask whether a local neighborhood changes after differentiation. A viral RNA segment can be tagged to map host proteins near replication intermediates. In each case, the strongest claim is conditional: proteins were enriched near the tagged RNA under a defined expression, localization, and labeling condition. Direct binding must be tested separately.

The boundary cases are important. Some RNAs cannot tolerate tags. Some tags recruit coat proteins that oligomerize, bind endogenous RNAs, or alter condensate formation. Some proximity enzymes perturb redox state, biotin metabolism, organelle function, or local protein behavior. Some cell types poorly take up substrates. Some labeling windows are too slow for fast RNA dynamics. Negative results can reflect failed recruitment, inaccessible tag folding, low enzyme activity, or insufficient mass-spectrometry depth rather than absence of neighbors. Positive results can reflect compartment background rather than target specificity. A serious study reports tag performance as part of the result, not as an assumed technical detail.

Table 138.2. Aptamer and proximity-labeling control logic. Aptamer tags and proximity-labeling systems can alter folding, localization, expression, or labeling radius; matched tag, enzyme, localization, and expression controls determine whether enrichment is interpretable.

Design feature What can go wrong Measurement or control Interpretation if control fails
Repeated RNA hairpins Tag arrays can alter RNA stability, folding, processing, or localization Compare tagged and untagged RNA abundance, isoforms, half-life, and localization Avoid endogenous-function claims or redesign the tag placement and copy number
Coat-protein fusion Fusion proteins can oligomerize, bind off-target RNAs, or reshape condensates Include fusion-only, tag-only, and expression-matched controls Down-weight proteins shared with controls and interpret remaining hits as tag-state specific
Proximity enzyme Labeling can reflect a whole compartment rather than the target RNA neighborhood Use localization-matched controls and labeling time courses Interpret enrichment as broad neighborhood exposure, not target-specific contact
Substrate or pulse condition Substrate dose, pulse length, or chemistry can induce stress or change labeling radius Include substrate-free, short-pulse, dose-response, and stress-readout controls Separate method-induced labeling from RNA-dependent biology
Overexpression Excess tagged RNA can saturate RBPs or nucleate ectopic assemblies Prefer endogenous tagging or expression titration with localization and function checks Limit conclusions to the reporter or overexpression system

138.3. Crosslinking chemistry for RNA-centric proteomics

Crosslinking is used because RNPs can rearrange during lysis. Cell disruption dilutes compartments, releases nucleases, changes salt and metabolites, breaks membranes, exposes naked RNA surfaces, and allows proteins that never met in vivo to reassort. Crosslinking attempts to preserve proximity or contact before this disruption. The chemistry determines the molecular distances and functional groups that can be stabilized, so crosslinking is not a generic “fixation” step. It is a selective observation filter.

Ultraviolet crosslinking is the best-known route for direct RNA-protein contacts. UV irradiation can create covalent bonds between nucleic acids and amino acid side chains that are close enough and positioned appropriately. The advantage is interpretive: compared with formaldehyde, UV crosslinking is more biased toward proteins in direct contact with RNA or DNA. The limitations are equally important. Crosslinking efficiency is low. Reactivity depends on nucleotide identity, amino acid chemistry, base exposure, local geometry, and irradiation conditions. Some genuine RNA-protein interactions crosslink poorly or not at all. Some proteins crosslink at only a subset of their binding sites. Therefore absence from a UV-based capture is not strong evidence of absence from the RNP.

Photoactivatable ribonucleoside strategies can increase recovery by incorporating analogs into RNA before irradiation. These methods are powerful when labeling is efficient and tolerated, but metabolic incorporation is uneven across organisms, cell types, RNA classes, growth conditions, and nucleotide pools. The analog can perturb RNA metabolism. Labeled RNA populations may not represent all RNA states. In viral infection or therapeutic-RNA contexts, analog incorporation may be impractical or biologically confounding. The experiment must state which RNA population was labelable and what unlabeled population was invisible.

Formaldehyde and related reversible crosslinkers preserve broader neighborhoods. Formaldehyde can stabilize protein-protein, protein-nucleic acid, and chromatin-associated interactions over short distances. This breadth is valuable for nuclear lncRNAs, chromatin-associated RNAs, condensates, and large RNP assemblies that would fall apart under native lysis. It also broadens the claim. A protein recovered after formaldehyde capture may be linked through another protein, through DNA, through chromatin, or through a fixed compartment. Formaldehyde capture is therefore strong evidence for association under fixation and weaker evidence for direct RNA binding unless combined with direct-contact assays.

Other chemistries target specialized problems. Psoralen-like reagents stabilize base-paired nucleic acids and are most relevant when RNA-centric proteomics is integrated with RNA-RNA or RNA-DNA interaction mapping. Diazirines, aryl azides, benzophenones, and related photoactivatable groups can be installed in nucleotides, probes, ligands, or small molecules to capture local contacts after light activation. Clickable handles can permit selective enrichment of labeled material. Crosslinkers with defined spacer lengths can estimate distance classes, although cellular accessibility and reaction kinetics complicate simple distance interpretation. Each chemistry offers selectivity, and each creates blind spots.

Figure 138.3. Crosslinking chemistry defines observable contact classes

Figure 138.3. Crosslinking chemistry defines observable contact classes. UV irradiation preferentially captures direct nucleic acid-protein contacts but has low efficiency and sequence bias. Formaldehyde preserves broader neighborhoods including protein-protein and chromatin-mediated associations. Photoactivatable nucleotides and reactive handles add specialized selectivity. Denaturing capture then tests which stabilized contacts survive stringent extraction.

Crosslinking conditions interact with fragmentation and capture. In endogenous antisense capture, cells may be crosslinked, lysed, fragmented by sonication or nuclease, hybridized to probes, washed, and reversed. Short fragments can sharpen local association but may lose long-range RNP context. Long fragments can preserve assemblies but reduce positional precision and increase indirect recovery. Harsh reversal can damage proteins or reduce peptide recovery. In denaturing capture, strong detergents, chaotropes, high salt, or heat can remove many noncovalent contaminants, but only covalently stabilized material survives. The resulting proteome reflects both biology and crosslink chemistry.

Crosslinking has quantitative artifacts. More crosslinking is not always better. Over-fixation can reduce probe accessibility, impair solubilization, trap abundant compartments, and lower peptide identification. Under-fixation can allow dissociation or reassortment. Crosslinking can create differential recovery unrelated to biology if two samples differ in cell density, light penetration, chromatin compaction, stress, or RNA abundance. Crosslink reversal may be incomplete. Crosslinked peptides can digest poorly and be difficult to identify. A good study therefore treats crosslinking dose, timing, sample thickness, quenching, reversal, and extraction as variables that require optimization and reporting.

The most careful language separates contact from capture survival. “UV-crosslinked to RNA” suggests closer contact than “formaldehyde-preserved with RNA-containing chromatin.” “Recovered after denaturing hybridization capture” suggests covalent stabilization or extremely stable association. “Recovered in native pulldown” suggests association under extraction conditions. These differences should be preserved in claim wording. A proteomics table should not flatten all crosslinking strategies into a single category called “RNA binders.”

138.4. Quantitative proteomics of RNPs and condensates

Mass spectrometry transforms recovered material into peptide and protein measurements. The central challenge is that RNP captures are rarely clean. Proteins bind beads, streptavidin, antibodies, oligonucleotides, resin surfaces, affinity tags, nucleic acids, detergents, and plastics. Abundant proteins can appear because they are abundant. Many RBPs bind RNA broadly or electrostatically. Ribosomal proteins, hnRNP proteins, helicases, chaperones, cytoskeletal proteins, metabolic enzymes, translation factors, splicing factors, and stress-response proteins can appear across unrelated captures. Quantitative proteomics is the discipline that decides which of these observations are specific enough to follow.

From captured protein material to peptides

The protein workflow starts only after the RNA-defined material has been recovered and its RNA yield and specificity have been measured. Proteins may be released by reversing formaldehyde crosslinks, digesting the RNA, displacing an affinity interaction, heating in detergent, or processing bead-bound material directly. The release method is not neutral. Incomplete crosslink reversal lowers recovery; prolonged heat or extreme pH can modify proteins; residual detergent, salt, nucleic acid, polyethylene glycol, or probe material can suppress digestion and ionization. Low-input captures are especially vulnerable because adsorption to tubes and cleanup media can remove a large fraction of the sample. A method report should therefore connect RNA recovery, protein elution, cleanup losses, and the amount injected into the mass spectrometer rather than presenting the proteomic run as an independent black box.

In a typical bottom-up workflow, proteins are first denatured so that folded domains and complexes no longer shield cleavage sites. Disulfide bonds are reduced, commonly with dithiothreitol or tris(2-carboxyethyl)phosphine, and liberated cysteines are alkylated to prevent disulfide re-formation. Trypsin is the usual protease because cleavage after lysine and arginine often generates peptides well suited to positive-ion tandem mass spectrometry, but lysine-rich RBPs, very basic ribosomal proteins, small proteins, low-complexity domains, and crosslinked regions can yield peptides that are too short, too long, too modified, or poorly retained. Lys-C, Glu-C, chymotrypsin, or sequential digestion can recover complementary sequence space. The protease, enzyme-to-protein ratio, digestion time, allowed missed cleavages, and cleanup method therefore shape which RNP proteins are observable.

Digestion is analytically powerful and conceptually destructive. Once a captured mixture is converted to peptides, the instrument does not know which peptides came from the same intact protein molecule, which protein molecules occupied the same RNA molecule, or which RNA-state subpopulation carried them. A peptide unique to protein P can support the presence of P in the recovered material; it does not reveal whether P was full length, truncated, modified, or accompanied by protein Q on the same RNP. Bottom-up analysis is consequently well suited to broad discovery and relative enrichment, but it cannot by itself reconstruct intact proteoforms or single-particle RNP composition.

Liquid chromatography-tandem mass spectrometry and sequence assignment

Liquid chromatography separates the peptide mixture over time before electrospray ionization introduces peptide ions into the mass spectrometer. An MS1 survey records precursor mass-to-charge values and signal intensities. In data-dependent acquisition, the instrument selects a subset of precursors, usually favoring intense ions, for fragmentation; low-intensity peptides can be skipped stochastically. Data-independent acquisition fragments broader precursor windows and can reduce selection stochasticity, but the resulting multiplexed fragment signals require library-based or predicted-spectrum deconvolution. In both cases, MS/MS fragmentation produces sequence-informative ions from which candidate peptide sequences are inferred.

A peptide-spectrum match is conditional on the search space. A conventional search specifies a protein FASTA database, the protease and permitted missed cleavages, precursor and fragment tolerances, fixed modifications such as cysteine alkylation, variable modifications such as methionine oxidation, and a scoring and target-decoy procedure. Adding every possible splice isoform, sequence variant, post-translational modification, contaminant, viral protein, and synthetic construct increases biological coverage but also enlarges the search space and identification burden. An RNP experiment must include sequences that could actually be present: host and pathogen proteomes for infected cells, engineered coat-protein or labeler fusions for tagged workflows, and common laboratory contaminants. Conversely, a search that omits a viral protein, fusion junction, noncanonical isoform, or sample-specific variant cannot identify it regardless of spectral quality.

False-discovery rate should be controlled and reported at the relevant levels. A threshold applied to peptide-spectrum matches does not automatically guarantee the same error rate for distinct peptides, proteins, or protein groups. Crosslinked or chemically modified peptides may fall outside a standard search and disappear from the identifiable pool. Search-engine version, database release, decoy strategy, peptide length, charge-state rules, modification list, and protein-inference settings are therefore scientific assumptions, not clerical metadata.

Protein inference begins after peptide assignment. A peptide sequence unique to one database protein is discriminating evidence under that database. A shared peptide may map to several paralogs, splice isoforms, protein families, or proteoforms. Parsimony-based algorithms often report a protein group that explains the observed peptides with a minimal set of entries, but that group is not proof that every listed member was present. Isoform-specific claims require isoform-discriminating peptides or independent evidence. Modification-site claims require localized modified-peptide evidence. A single-peptide identification can be real, especially for a small or low-abundance RNP protein, but its spectrum, uniqueness, reproducibility, and behavior in controls deserve direct inspection before mechanistic promotion.

A minimal discovery experiment compares target capture to matched controls across biological replicates. Label-free quantification estimates protein abundance from peptide intensities or spectral evidence across runs. Stable isotope labeling can mix samples before processing and reduce some run-to-run variation. Isobaric tags can multiplex multiple conditions in one mass-spectrometry experiment, although ratio compression and missing values require care. Targeted proteomics can quantify candidate peptides with higher precision after discovery. Spike-in standards can monitor recovery and instrument variation. The exact platform matters less than the design principle: enrichment must be measured against the right background.

The mixing point determines which technical variation a design can control. Label-free quantification keeps samples separate through chromatography and acquisition, enabling large cohorts but exposing comparisons to run order, retention-time alignment, ion sampling, and between-run missingness. Stable isotope labeling by amino acids in cell culture (SILAC) can mix target and control cells or lysates early, so later preparation losses are shared, but complete incorporation, label swapping, arginine-to-proline conversion, cell-type compatibility, and mixing accuracy must be checked. Tandem mass tag (TMT) and related isobaric labels multiplex peptide digests after separate upstream processing; multiplexing improves throughput and within-set comparison, yet co-isolated peptides can distort reporter-ion ratios unless acquisition and interference controls are appropriate. A label does not repair an unmatched biological control or a failed RNA capture.

Quantification can occur at the peptide, protein-group, or targeted-peptide level. Feature-intensity label-free methods compare chromatographic ion areas; spectral counting uses how often peptides were sampled and is generally less precise for modest changes. Protein-level summaries combine peptide measurements through explicit rules, so discordant peptides should not be hidden automatically: they can indicate interference, a shared peptide, an isoform difference, a modification, proteolysis, or a poor match. For targeted follow-up, selected- or parallel-reaction monitoring can repeatedly measure proteotypic peptides, and heavy synthetic peptides can support calibrated peptide amounts. Immunoblotting or immunoassays provide a different measurement chain, although antibody specificity and dynamic range remain limiting.

The right background depends on the method. In vitro pulldown needs beads-only, affinity-handle-only, unrelated RNA, mutant RNA, and sometimes folded-state controls. Antisense capture needs scrambled or sense probes, unrelated but localization-matched RNA, off-target RNA measurements, and probe-set comparisons. Aptamer capture needs tag-only and untagged controls plus expression and localization checks. Proximity labeling needs enzyme-only, substrate-free, localization-matched, and time-course controls. Crosslinked capture needs no-crosslink, crosslink-dose, and reversal controls. RNase treatment is useful but not decisive because it can remove indirect RNA-dependent assemblies as well as direct contacts.

Figure 138.4. Quantitative enrichment separates candidates from background

Figure 138.4. Quantitative enrichment separates candidates from background. Target capture should be compared with matched controls such as mutant RNA, unrelated RNA, RNase treatment, reagent-only capture, and localization-matched proximity labeling. Proteins that show reproducible effect-size enrichment and adequate peptide evidence become candidates; proteins that track with beads, probes, tags, enzyme expression, stress, or general compartment background are down-weighted.

Protein-level evidence has grades. A protein identified by one peptide in one replicate is a weak lead. Multiple unique peptides across replicates are stronger. Reproducible enrichment over controls with a substantial effect size is stronger still. Consistent behavior across independent capture designs, loss after RNA mutation, and persistence in orthogonal assays are much stronger. Missing values require caution: a protein absent from controls and present in target may be highly specific, but it may also be near the detection limit. Statistical analysis should include replicate structure, multiple-testing awareness, effect size, peptide evidence, and missingness behavior rather than relying only on presence or absence.

Missingness has more than one cause. A peptide can be absent because its protein truly was not recovered, because its abundance fell below the detection limit, because data-dependent selection chose other ions, because chromatography or ionization failed, or because the identification algorithm did not accept the spectrum. These mechanisms are not interchangeable. Replacing every missing control value with a small number can manufacture enormous target-to-control ratios, whereas deleting all incomplete proteins can remove specific low-abundance candidates. A defensible analysis reports detection patterns, distinguishes censored low-intensity observations from apparently random acquisition gaps when possible, tests whether conclusions depend on the imputation rule, and gives effect sizes alongside adjusted statistical evidence. Targeted reacquisition is often more informative than increasingly elaborate imputation for a short list of decisive candidates.

Concrete RNA-capture studies illustrate why the proteomics layer changes interpretation. Xist chromatin isolation by RNA purification followed by mass spectrometry (ChIRP-MS) recovered a broad Xist-associated protein set using crosslinked material and quantitative comparisons, whereas RNA antisense purification followed by mass spectrometry (RAP-MS) combined stringent antisense capture, UV-biased contact preservation, SILAC, and functional follow-up to prioritize a smaller direct-contact-biased set that included SHARP/SPEN. The different list sizes should not be read as one experiment being simply right and the other wrong. Crosslink chemistry, probe design, washing, labeling, search depth, inference thresholds, and the biological state define different observable sets. Both studies became mechanistically informative because candidate enrichment was linked to Xist recovery and followed by perturbation or other orthogonal evidence.

Condensates and membraneless bodies add a conceptual layer. RNA-rich condensates such as stress granules, processing bodies, nucleoli, paraspeckles, germ granules, and viral inclusions can concentrate many proteins through multivalent interactions, crowding, phase separation, active transport, or stress-induced remodeling. Capturing an RNA from a condensate may recover scaffolds that help build the body, clients that partition into it, enzymes that act there, regulators that control assembly, passengers that are transiently concentrated, and contaminants that become insoluble during purification. These categories have different biological meanings.

For condensates, quantitative proteomics should ask state-specific questions. Is the protein enriched only when the RNA-containing body forms? Does enrichment depend on the RNA scaffold, a low-complexity protein domain, translation inhibition, stress condition, or phase-separation trigger? Does the protein exchange rapidly or remain stably associated? Does removal of the protein change condensate formation, RNA localization, or RNA function? Does the protein remain enriched after stringent washing, or is it lost because it is a weak but real client? A scaffold/client/passenger vocabulary is more accurate than a binary member/nonmember vocabulary.

Quantification also needs stoichiometric humility. Mass spectrometry signal is not the same as copy number without appropriate calibration. A protein with high ionization efficiency may appear abundant. A low-abundance regulator may be biologically important but difficult to detect. Crosslinked proteins can produce fewer peptides. Membrane proteins, very basic proteins, low-complexity proteins, and small proteins may be underrepresented. RNP composition is often heterogeneous: different molecules of the same RNA can carry different protein sets. Bulk proteomics averages these states. The measured proteome is therefore an enriched ensemble, not necessarily a single defined complex.

Relative enrichment is not RNP stoichiometry. Estimating copies of protein P per RNA R requires, at minimum, calibrated recovery of P, a defensible amount or molecule count for the captured RNA population, recovery corrections for both analytes, and evidence that the recovered material represents a defined assembly rather than a mixture of RNP states. Heavy peptide standards can calibrate peptide amount, but they do not correct automatically for incomplete protein extraction, digestion, crosslink reversal, or RNA recovery. Equal protein signal for two candidates also does not imply equal copy number because their peptides differ in yield and response. Ordinary discovery capture should therefore report relative enrichment or calibrated recovered amount, not subunit stoichiometry.

Intact or top-down protein analysis can be useful when an RNP question depends on a proteoform, cleavage product, or combination of post-translational modifications that peptide aggregation cannot reconstruct. These methods analyze whole proteins before or during fragmentation and thereby preserve more proteoform connectivity, but complex captured mixtures, low input, dynamic range, intact-protein separation, ionization, deconvolution, and false-discovery control constrain routine use. Intact-protein evidence remains protein-centric. It does not become native-RNP evidence unless an intact RNA-protein assembly is maintained and measured as an assembly.

Table 138.3. Proteomic evidence levels for RNP candidates. RNP candidates progress from reproducible enrichment to direct contact, endogenous association, perturbation response, and mechanism; claim language should reflect the strongest completed evidence level.

Evidence level Typical observation Main risk Appropriate claim language
Detection only Protein appears in one target capture sample Stochastic detection, carryover, or common contaminant Candidate detected in the RNA capture
Replicate detection Protein appears across biological replicates Recurrent background protein with no target specificity Reproducibly detected candidate
Quantitative enrichment Protein is enriched over matched controls with peptide support Incorrect background model or missing-value artifact RNA-associated candidate under the specified conditions
Orthogonal capture support Enrichment persists with another probe set, tag, or workflow Shared artifact or shared compartment background Robust RNA-associated candidate requiring directness tests
Direct-contact support CLIP signal, purified binding, or motif-dependent recovery supports contact Condition mismatch between assays Direct binder under the tested conditions
Functional support Perturbation and rescue connect the association to an RNA-dependent output Indirect pathway effect or broad toxicity Functional regulator of the RNA-dependent process

The practical output of RNA-centric proteomics should be a ranked candidate list with annotations, not a final membership roster. Useful annotations include enrichment effect size, replicate reproducibility, peptide count, RNA recovery for the sample, sensitivity to RNase, sensitivity to mutation, appearance in reagent controls, known RNA-binding domains, known localization, prior CLIP evidence, known protein-protein partners, and functional perturbation data. Proteins can then be prioritized for validation based on the biological question. For example, a viral RNA study might prioritize host factors that are enriched only during replication and whose knockdown reduces viral RNA accumulation. A lncRNA study might prioritize chromatin regulators that depend on an RNA domain and colocalize with the RNA locus.

The chapter boundary can now be stated operationally. If digestion and LC-MS/MS identify peptides from proteins that followed an RNA capture, Chapter 138 owns the measurement and inference. If the measured ions are ribonucleosides, RNase-generated oligonucleotides, or intact RNA products, Chapter 132 owns the RNA analyte, even when proteins were used during purification. If the measured object is an intact RNP or RNA-protein assembly and the output concerns assembly mass, oligomeric state, bound ligands, collision behavior, ion mobility, or subunit stoichiometry, Chapter 59 owns the native-complex interpretation. These boundaries prevent the shared phrase “RNP mass spectrometry” from collapsing three different analytes and evidence models.

138.5. Native versus denaturing workflows and contamination control

Native and denaturing workflows answer different questions. Native workflows use gentle lysis and washing to preserve noncovalent assemblies. They can recover labile RNPs, protein-protein bridges, enzymatic complexes, ribosomes, localization particles, and condensate clients. They are useful when the biological state depends on weak or reversible interactions. Their weakness is that they preserve indirect associations and allow new associations after lysis. If naked RNA is exposed in extract, RBPs can bind it. If compartments rupture, proteins can mix across spaces that never contacted each other in the cell. If salt and ATP conditions change, remodeling enzymes can rearrange complexes.

Denaturing workflows use strong disruption to reduce noncovalent carryover. High salt, detergents, chaotropes, heat, or organic conditions can strip many protein-protein and RNA-protein interactions. Denaturing capture is most meaningful when crosslinking, chemical handles, or extremely stable hybrids preserve the target material. This can improve specificity for covalently stabilized contacts and allow stringent washing, but it does not make the result universally more biological. Denaturation can destroy real weak interactions, bias recovery toward crosslinkable residues, lower peptide detection, and obscure stoichiometry. It asks which proteins survive the chosen chemical and extraction filter, not which proteins were present in every native RNP state.

Contamination control should be viewed as a hierarchy. The first layer is reagent background: beads, streptavidin, antibodies, oligonucleotides, tags, enzymes, dyes, and plastics bind proteins. The second layer is sample background: abundant proteins and broadly nucleic-acid-binding proteins recur across many captures. The third layer is post-lysis background: proteins can reassort onto RNA or exposed surfaces after extraction. The fourth layer is compartment background: chromatin, membranes, ribosomes, granules, spliceosomes, and viral factories can co-purify. The fifth layer is interpretation background: prior expectations can make a familiar protein seem meaningful and an unfamiliar protein seem like contamination without evidence.

Table 138.4. Contamination classes and control logic. Abundant proteins, sticky surfaces, beads, tags, neighboring compartments, and sample handling create distinct contamination classes; persistent background narrows the supported claim even when enrichment is reproducible.

Background class Typical source Warning sign Control or mitigation Interpretation if persistent
Reagent background Beads, streptavidin, probes, affinity tags, enzymes, dyes, or plastics Same proteins recur in reagent-only or tag-only samples Reagent-only, tag-only, beads-only, and probe-only controls Remove or strongly down-weight unless target-dependent enrichment is clear
Abundance background Ribosomal proteins, hnRNPs, chaperones, cytoskeletal proteins, and metabolic enzymes Protein appears across unrelated RNA captures Unrelated RNA controls, contaminant repositories, and effect-size thresholds Treat as candidate only when the pattern follows the target RNA or state
Post-lysis reassortment Exposed RNA surfaces and mobile RBPs after extraction Signal increases with slow, gentle, or poorly controlled lysis Crosslink-before-lysis, rapid lysis, stringent washes, and RNA competition Interpret native-only signals cautiously
Spatial co-purification Chromatin, membranes, granules, ribosomes, or viral replication compartments Proteome matches compartment markers more than the target RNA Localization-matched controls, imaging, fractionation, and compartment marker tracking Interpret as neighborhood enrichment unless direct evidence exists
Stress or perturbation background Tag overexpression, transfection, infection, oxidative chemistry, or enzyme expression Proteome changes with delivery, dose, substrate, or labeler expression Expression titration, dose controls, stress readouts, and time courses Separate method-induced biology from target-specific RNP composition

The strongest control designs attack more than one layer. Reagent-only controls identify bead and probe binders. Unrelated RNAs identify general RNA affinity. Localization-matched controls identify compartment background. Mutant RNAs test sequence, structure, or modification dependence. RNase and nuclease controls test whether enrichment requires RNA. Crosslink-before-lysis controls reduce reassortment. Rapid lysis and stringent washing reduce post-extraction artifacts. Orthogonal capture using a different probe set or tag tests method-specific background. Candidate validation by immunoblotting, targeted proteomics, reciprocal capture, CLIP, imaging, and purified binding tests whether the discovery signal survives outside the original workflow.

RNase sensitivity is often misunderstood. If a protein disappears after RNase treatment, the association is RNA-dependent, but RNA dependence does not prove direct RNA binding. The protein may depend on another RBP, an RNA-mediated condensate, an RNA scaffold, or RNA-maintained chromatin structure. Conversely, RNase resistance does not prove RNA independence. Crosslinking, steric protection, incomplete digestion, insoluble material, or protein-protein retention can preserve a signal after RNase. RNase experiments are most informative when digestion efficiency is measured and when multiple nucleases or conditions are chosen to match the RNA structure being tested.

Another common mistake is to treat recurrent background lists as universal truth. A protein that appears in many RNA captures may be a contaminant in one experiment and a real broad RNP component in another. Poly(A)-binding proteins, hnRNPs, ribosomal proteins, helicases, and translation factors can be biologically relevant for many RNAs. The key question is whether their enrichment pattern follows the target RNA, a motif, a state, or a compartment better than it follows abundance and reagent background. Background knowledge should down-weight weak nonspecific signals; it should not automatically erase plausible biology.

Native and denaturing workflows can be combined productively. A native capture can identify a broad candidate neighborhood. A denaturing crosslinked capture can ask which proteins remain close enough for covalent stabilization. A purified binding assay can test direct contact. An imaging experiment can test spatial state. A perturbation experiment can test function. Agreement across these workflows is strong evidence. Disagreement is also informative: a protein seen only in native capture may be an indirect co-complex member or a weak client; a protein seen only after crosslinking may be a transient direct contact; a protein seen only in proximity labeling may be a compartmental neighbor.

For the running lncRNA example, native capture may recover an RNA-binding scaffold, chromatin-associated enzymes, abundant hnRNPs, and nucleolar proteins. Denaturing capture after formaldehyde may retain chromatin-proximal proteins but lose weak clients. UV-based capture may recover only a subset of direct RBPs. Imaging may show that some candidates colocalize with the RNA only during a cell-cycle stage. A domain deletion may remove one class of proteins but not another. The biological model should preserve these distinctions rather than force all candidates into one “lncRNA interactome.”

138.6. Integration with CLIP, imaging, and functional perturbation

RNA-centric proteomics is most useful as the discovery and composition arm of a larger evidence system. It says which proteins are enriched with an RNA-defined material. CLIP-family methods, imaging, biochemical reconstitution, structural methods, genetics, perturbation screens, and functional assays answer different parts of the problem. A mature RNP model usually needs several of these evidence classes because no single method simultaneously proves composition, direct binding, binding site, localization, stoichiometry, dynamics, and biological function.

CLIP and related protein-centric methods are natural partners. If RNA-centric capture identifies protein P as enriched with RNA R, CLIP can ask whether P contacts R in cells and where on R the contact occurs. A CLIP peak on the captured RNA does not prove that the same molecules were recovered in the capture, but it provides independent direct-contact evidence. Mutation of the CLIP-defined motif or structure can then test whether the proteomic enrichment depends on that site. Conversely, if a protein is enriched by capture but lacks CLIP evidence, it may still be a real indirect component, a low-efficiency crosslinker, a condition-specific binder not represented in the CLIP atlas, or a false positive.

Imaging adds spatial and temporal evidence. RNA fluorescence in situ hybridization can show where the target RNA resides. Immunofluorescence or tagged-protein imaging can test whether candidate proteins colocalize with the RNA. Live-cell systems can show whether protein recruitment follows RNA transcription, export, localization, stress, translation, or decay. Super-resolution imaging can distinguish broad compartment co-occupancy from closer spatial organization, although optical resolution still does not prove molecular contact. Imaging is especially important for nuclear architectural RNAs, localized mRNAs, viral replication compartments, and condensates because spatial neighborhoods strongly influence capture results.

Functional perturbation tests whether an association matters. The RNA can be depleted, mutated, truncated, relocalized, chemically modified, or rescued. The protein can be knocked down, knocked out, acutely degraded, inhibited, domain-mutated, relocalized, or rescued with RNA-binding-defective variants. A strong functional experiment asks whether perturbing the association changes the relevant output: RNA stability, splicing, translation, localization, chromatin state, viral replication, condensate formation, immune activation, cell phenotype, or therapeutic activity. Rescue is particularly valuable because it distinguishes direct mechanism from broad toxicity or stress.

Figure 138.5. Evidence ladder for RNA-protein association claims

Figure 138.5. Evidence ladder for RNA-protein association claims. Detection supports a weak candidate claim. Quantitative enrichment supports association. CLIP or purified binding supports direct contact. Imaging supports spatial context. RNA or protein mutation supports specificity. Rescue or perturbation supports functional mechanism. A claim should not be worded above the highest evidence tier actually tested.

Biochemical reconstitution remains the cleanest test of direct binding and mechanism. Purified RNA and purified protein can establish binding, affinity, stoichiometry, specificity, competition, and sometimes kinetics. Mutant RNAs, modified RNAs, truncated proteins, and RNA-binding-domain mutants can define the contact logic. Reconstitution cannot fully reproduce cellular context, but it can answer directness questions that capture proteomics cannot. For enzymes, reconstitution can connect binding to activity: a helicase may unwind an RNA structure, a nuclease may cleave a transcript, a modification enzyme may alter a nucleotide, or a translation factor may change initiation.

Integration also means scaling claims downward when evidence is incomplete. If a protein is enriched by antisense capture but not tested by CLIP, imaging, or perturbation, the appropriate claim is candidate association. If a protein has CLIP signal on the RNA but no functional perturbation, the appropriate claim is direct contact or likely binding, not biological requirement. If a protein is required for a phenotype but direct binding is untested, the appropriate claim is functional dependence, not necessarily direct RNA regulation. Good RNA biology often advances by precise intermediate claims rather than premature mechanism.

Box 138.3. Claim Language for RNA-Protein Associations

  • Box body (render-ready Markdown):

Use verbs that match the evidence. “Detected in an RNA capture” means the protein was observed in the recovered material. “Enriched with the captured RNA” means the protein increased over matched controls with quantitative support. “RNA-dependent” means the signal changed after RNase treatment or RNA perturbation, but it may still be indirect. “Proximal to the RNA” means the protein was labeled near an RNA-recruited enzyme during a defined window. “Directly binds the RNA” requires direct-contact evidence such as CLIP, crosslink mapping, purified binding, or a binding-site mutation with appropriate controls. “Regulates the RNA-dependent process” requires perturbation and preferably rescue. The most useful sentence often combines method and limit: “Protein P was enriched with formaldehyde-preserved RNA R chromatin material and remains a candidate indirect RNP component pending direct-contact tests.”

RNA-centric proteomics can also guide screens. A candidate list can seed CRISPR, RNAi, degron, or Perturb-seq experiments that test which proteins affect RNA abundance, localization, processing, translation, or downstream cell state. The candidate list can be intersected with CLIP atlases, domain annotations, disease variants, RNA modification readers, condensate proteomes, and compartment maps. Machine-learning or network approaches can prioritize proteins, but computational ranking should be treated as hypothesis generation. The experimental evidence remains the determinant of claim strength.

For synthetic and therapeutic RNAs, integration is critical because delivery and dose can dominate biology. A modified mRNA may recruit translation factors, decay factors, innate immune sensors, endosomal proteins, stress granule components, and delivery-particle-associated proteins. RNA-centric proteomics can identify these associations, but interpretation requires delivery controls, dose-response, time-course, RNA-quality measurements, innate-immune readouts, translation assays, and functional potency. For guide RNAs or RNA-targeting therapeutics, capture proteomics can reveal off-target RNP remodeling or protein sequestration, but clinical or engineering claims require evidence beyond enrichment.

The final product of an integrated study is not simply a list of proteins. It is a model that says which proteins contact the RNA, which proteins are indirect complex members, which proteins share a compartment, which proteins are required for function, which associations are condition-specific, and which observations remain unresolved. The model should include negative evidence and boundary cases. A candidate that fails purified binding but remains functionally required may act through another protein. A candidate that binds directly in vitro but fails imaging or perturbation may bind only under nonphysiological conditions. A candidate that appears only under stress may reveal a state-specific RNP rather than an artifact.

Figure 138.6. From RNA-centric discovery to an RNP mechanism

Figure 138.6. From RNA-centric discovery to an RNP mechanism. A robust study defines the RNA target, chooses capture chemistry, measures RNA recovery, quantifies protein enrichment, filters background, validates binding or proximity, perturbs RNA and protein features, and tests a biological output. The workflow is iterative because failed validation can reveal indirect association, state specificity, or method artifact.

Experimental Foundations and Evidence Standards

The evidence standard for RNA-centric proteomics begins with exact wording. “Protein P was detected in the RNA R capture” is a detection claim. “Protein P was enriched over matched controls” is a quantitative association claim. “Protein P directly binds RNA R” is a direct-contact claim. “Protein P regulates RNA R” is a functional claim. “Protein P is part of the RNA R condensate scaffold” is a structural or material-state claim. These claims require different evidence. Treating them as synonyms is one of the most common sources of overstatement.

RNA recovery metrics are foundational evidence. The target RNA should be enriched relative to input and controls, and recovery should be consistent across replicates. Off-target RNAs should be measured when sequence similarity, abundance, compartment, or probe design makes them plausible. Fragment size, probe coverage, crosslinking efficiency, lysis conditions, wash stringency, and elution method should be reported because they define what material could have been recovered. Protein evidence should include unique peptide support, replicate reproducibility, effect size, and behavior in negative controls.

Evidence becomes stronger when independent methods agree without sharing the same artifact. Antisense capture and aptamer capture share RNA enrichment logic but have different backgrounds. UV crosslinking and formaldehyde capture preserve different contact classes. CLIP provides protein-centric site evidence. Imaging provides spatial evidence. Reconstitution provides direct biochemical evidence. Functional perturbation provides causal evidence. The best studies use agreement and disagreement among these methods to refine the model rather than to force a single simplistic conclusion.

Biological Contexts Across Systems

Long noncoding RNAs are frequent targets because many lncRNAs are hypothesized to function by recruiting, scaffolding, decoying, or organizing proteins. Nuclear lncRNAs require special caution because chromatin, nuclear bodies, nascent transcription, and DNA-binding proteins can be co-captured. A lncRNA-associated chromatin regulator may bind the RNA directly, bind chromatin near the RNA, bind another RNP protein, or occupy the same nuclear compartment. Domain deletion, allele-specific analysis, imaging, CLIP, and chromatin perturbation can help separate these possibilities.

mRNAs present a state problem. The same mRNA can be newly exported, actively translated, localized, stored, decapping-prone, deadenylated, surveillance-targeted, or stress-granule-associated. Capture from the coding sequence may emphasize ribosomes and translation factors. Capture from the 3′ untranslated region may emphasize localization factors, miRNA machinery, poly(A)-binding proteins, and decay regulators. Time, stress, cell type, and translation inhibitors can alter the recovered proteome. Integration with ribosome profiling, polysome fractionation, RNA localization assays, and RNA stability measurements is often necessary.

Viral RNAs are attractive because viral genomes and transcripts recruit host factors for replication, translation, packaging, immune evasion, and decay resistance. Infection also changes the cell. Viral replication organelles, inclusions, nucleocapsids, membrane compartments, innate immune activation, and stress responses can dominate capture. Controls may include uninfected cells, replication-defective viral mutants, RNA-domain mutants, time courses, antiviral treatments, and compartment markers. The goal is to distinguish proteins that specifically support or restrict viral RNA biology from proteins that accompany the infection state.

Small RNAs and guide RNAs require scale-aware design. A short RNA may be mostly buried inside Argonaute, PIWI, Cas, RNase P, spliceosomal, or other protein-dominated machinery. Tags may disrupt function; antisense probes may compete with native base pairing; crosslink recovery may be low. RNA-centric proteomics can still reveal accessory factors or maturation intermediates, but protein-centric CLIP, structural biology, purified reconstitution, and genetic perturbation may provide cleaner mechanistic evidence.

Synthetic and therapeutic RNAs introduce chemical and delivery variables. Modified nucleotides, cap structure, poly(A) tail length, circularization, double-stranded RNA contaminants, delivery vehicles, concentration, and subcellular entry route can change protein association. RNA-centric proteomics can compare how these variables affect translation factors, decay factors, innate immune sensors, granule proteins, and delivery-associated proteins. Interpretation requires controls for transfection, lipid or polymer carrier, RNA purity, dose, time, and cell stress.

Computational analysis can prioritize candidates but cannot validate them by itself. RNA-binding domains, low-complexity regions, protein disorder, localization annotations, CLIP atlases, protein-protein networks, contaminant repositories, and sequence or structure predictions should guide follow-up experiments without turning familiarity into proof. Engineering applications use the same methods as design tools: RNA tags can recruit regulators, editors, nucleases, fluorescent proteins, or proximity enzymes, and guide RNAs can direct programmable proteins to RNA targets. Clinical translation is most plausible when capture proteomics explains a mechanism, such as protein sequestration by repeat RNAs, host-factor recruitment by viral RNAs, innate immune sensing of therapeutic RNAs, or RNP remodeling by RNA-targeted drugs. Disease and therapeutic claims still require relevant models, dose dependence, rescue, specificity, safety context, and functional outcomes.

Three assays can all be described informally as “mass spectrometry of an RNP” while producing incompatible outputs. Protein bottom-up analysis after RNA capture yields peptide-spectrum matches, protein groups, and relative or calibrated protein abundance. RNA mass spectrometry yields nucleoside composition, oligonucleotide sequence or modification evidence, termini, or intact-RNA mass. Native RNP mass spectrometry yields assembly-mass distributions, ligand or subunit states, collision products, and ion-mobility observables. The sample-preparation step that preserves or destroys covalent and noncovalent connectivity determines which question remains answerable.

Table 138.5. Three mass-spectrometry analytes hidden by the phrase RNP-MS. The phrase RNP mass spectrometry can refer to three different analytical objects. Digested proteins after RNA capture support peptide and protein-enrichment claims. Nucleosides, oligonucleotides, or intact RNA support RNA chemistry, sequence, terminus, modification, or product-mass claims. Native intact RNPs support assembly-mass, ligand-state, and stoichiometric claims subject to native-MS limitations. The preparative step that destroys or preserves connectivity determines the valid inference.

Analytical object Connectivity before MS Primary outputs What the output cannot establish alone Primary owner
Peptides from proteins recovered by RNA capture Proteins are denatured and digested; protein-proteoform and same-RNA co-occupancy connectivity is mostly lost Peptide-spectrum matches, protein groups, relative enrichment, calibrated peptide amount RNA chemical identity, intact proteoform, copies per RNA molecule, or intact-RNP stoichiometry Chapter 138
Ribonucleosides, RNase-generated oligonucleotides, or intact RNA RNA is digested selectively or preserved as an intact analyte, depending on the workflow Nucleoside composition, oligonucleotide sequence or modification evidence, termini, intact-RNA mass Which proteins were bound or the stoichiometry of an intact RNP Chapter 132
Native intact RNA-protein assembly Noncovalent assembly connectivity is intentionally preserved during solution preparation and ionization Assembly-mass distributions, oligomeric or ligand states, collision products, ion mobility, subunit stoichiometry Direct solution-state structure without gas-phase and transmission caveats; peptide-level protein inference Chapter 59

Recent Consensus

Current consensus is methodological. RNA-centric proteomics is strongest when it is treated as controlled enrichment evidence. Direct binding, indirect co-complex membership, proximity, compartmental co-enrichment, and functional requirement are distinct claim types. Endogenous capture and genomic tagging are often more physiologically relevant than overexpression or in vitro baits, but neither is artifact-free. Quantitative controls and replicate-aware analysis are essential. Condensate and chromatin-associated captures require special caution because spatial co-enrichment can masquerade as molecular binding.

For protein mass spectrometry after RNA capture, consensus practice is to expose the complete inference chain: protein release, denaturation, reduction, alkylation, digestion, peptide cleanup, LC-MS/MS acquisition, search database and modification space, peptide- and protein-level error control, treatment of shared peptides, quantification level, missingness, and contaminant behavior. Bottom-up data support peptide and protein-group claims; they do not automatically preserve proteoform identity, co-occupancy on individual RNA molecules, or RNP stoichiometry.

Consensus is also forming around integration. A credible RNP model usually combines RNA recovery metrics, protein enrichment, independent binding or proximity evidence, localization evidence, perturbation evidence, and rescue or reconstitution when possible. The precise combination depends on the biological question. A discovery study can stop at candidate generation if the language is modest. A mechanistic study must show why the association matters.

Open Questions, Controversies, Deprecated Models, and Common Misconceptions

Open questions:

  • How should membership in dynamic RNPs be defined when regulatory RNAs occupy multiple nascent, mature, translated, localized, stored, decaying, and stress-induced states?

Controversies:

  • Condensate composition remains unsettled. Enrichment in an RNA-containing body can indicate scaffold function, client partitioning, enzymatic action, transient passage, or purification artifact.
  • Noncanonical RNA-binding proteins remain controversial. Proteome-wide studies have identified many RNA-associated proteins lacking classical RNA-binding domains; some are true noncanonical binders, while others are indirect, state-specific, or chemistry-favored recoveries.

Common misconceptions:

  • “An antisense-capture hit is a direct RNA-binding protein.” The experiment usually shows enrichment with captured RNA material.
  • “RNase sensitivity proves direct binding.” It proves RNA dependence under the tested conditions, not direct contact.
  • “Denaturing capture is always more specific.” It is more stringent for noncovalent carryover but specific to the crosslinking or chemical handle.
  • “Failure to recover a known RBP disproves association.” Low crosslinking efficiency, poor probe accessibility, low abundance, state specificity, and peptide detection limits can all cause false negatives.
  • “A protein database search identifies intact proteins directly.” Bottom-up searches assign spectra to peptides and infer proteins or protein groups; shared peptides and unobserved sequence regions limit isoform and proteoform conclusions.
  • “A protein absent from the control has infinite enrichment.” Nondetection may reflect censoring or stochastic acquisition, so missing controls require sensitivity analysis or targeted remeasurement rather than automatic substitution with an arbitrary small value.
  • “Peptide intensity reports copies per RNA molecule.” Response factors, digestion, recovery, crosslinking, and heterogeneous RNA states prevent ordinary capture proteomics from yielding RNP stoichiometry without calibration and an RNA denominator.
  • “RNA-capture proteomics chemically characterizes the RNA bait.” Protein-centric LC-MS/MS analyzes captured proteins; nucleoside, oligonucleotide, and intact-RNA measurements are separate RNA-analyte workflows in Chapter 132.