This chapter explains how experimentalists map contacts between RNA molecules and proteins. The scope includes protein-centric methods such as RNA immunoprecipitation (RIP) and crosslinking immunoprecipitation (CLIP), RNA-centric methods such as interactome capture, proximity and subcellular RBP profiling, computational inference of peaks and binding sites, quantitative interpretation of occupancy, and evidence grading. The central lesson is that RNA-protein interaction maps are conditional measurements produced by chemistry, purification, sequencing, proteomics, and statistical models. A map can be extremely informative without being a direct photograph of regulation.
RNA-binding proteins (RBPs) influence nearly every stage of RNA life, including transcriptional coupling, splicing, polyadenylation, export, localization, translation, decay, editing, modification, immune sensing, and viral replication. An RBP is a protein that associates with RNA either through a direct RNA-binding surface or as part of a ribonucleoprotein (RNP) assembly. An RNP is the physical complex that contains RNA and protein, often with additional RNAs, cofactors, or transient remodeling enzymes. RNA-protein interaction mapping asks which proteins contact which RNAs, where the contact occurs, when the contact occurs, how many RNA molecules are occupied, and whether the interaction changes RNA fate.
No single method answers all of those questions. RIP enriches RNAs that remain associated with an immunopurified protein or complex. RIP is useful for asking whether an RNA is in a protein-containing assembly, but it does not by itself prove direct contact or map the contact site. CLIP-family methods use ultraviolet crosslinking, immunopurification, nuclease trimming, library construction, sequencing, and peak calling to infer candidate direct contact regions for one protein or protein class. HITS-CLIP, PAR-CLIP, iCLIP, eCLIP, irCLIP, colocalization CLIP, and multiplexed variants differ in crosslinking chemistry, library readout, controls, sensitivity, throughput, and analysis. RNA interactome capture reverses the direction of the experiment by purifying an RNA class, often polyadenylated RNA, and identifying associated proteins by mass spectrometry. Proximity labeling and subcellular methods identify proteins near an RNA, RBP, compartment, or RNP state, expanding atlas construction but weakening direct-contact inference.
The strongest interpretation separates association, direct contact, site localization, fractional occupancy, and functional consequence. A CLIP peak can support candidate contact when controls and crosslink-site evidence are strong. A peak is not automatically a regulatory element, a stoichiometric measurement, or a universal binding rule. Crosslinking is inefficient and biased; antibody specificity can dominate the result; RNase digestion and ligation shape fragment recovery; reverse transcription creates method-specific signals; and peak callers model data rather than reveal sites without assumptions. RNA structure also changes accessibility, so RBP maps should often be interpreted together with structure probing and functional assays [Rapakko et al. 2026].
This chapter uses the verified local bibliography honestly. Local anchors include a PAR-CLIP review [Ascano et al. 2012], a current review on RNA structure and RBP interplay [Rapakko et al. 2026], a subcellular colocalization CLIP method paper [Yi et al. 2024], a multiplexed RNA-protein interaction method paper [Wolin et al. 2025], and a PAR-CLIP variant paper [Hinze et al. 2018].
The first prerequisite is the distinction between RNA sequence, RNA molecule, RNP complex, and RNA function. An mRNA contains a cap, untranslated regions, coding sequence, splice junctions, a poly(A) tail, structures, modifications, and proteins. A binding site is a region contacted by a protein under a condition. A regulatory element is a region whose alteration changes an output such as splicing, localization, translation, decay, or immune sensing. A CLIP peak may overlap a regulatory element, but peak presence alone does not prove regulation.
The second prerequisite is immunoprecipitation logic. An antibody or tag enriches a target protein from a lysate. Molecules recovered with the target can be direct partners, indirect partners bridged through other components, post-lysis associations, nonspecific carryover, or true in-cell complexes stabilized by crosslinking. The interpretation depends on crosslinking chemistry, wash stringency, nuclease treatment, antibody specificity, negative controls, input controls, and biological replicates.
The third prerequisite is sequencing-library logic. Short reads are not the original RNA molecules. Reads are shaped by fragmentation, nuclease digestion, RNA-end chemistry, ligation, reverse transcription, PCR amplification, size selection, mapping, deduplication, and filtering. CLIP analysis therefore follows a chain of inference: contact creates a crosslink; digestion and purification recover fragments; library construction generates reads; reads align to a genome or transcriptome; peak callers infer enriched sites; biological interpretation connects candidate sites to mechanism.
The fourth prerequisite is proteomics logic. Mass-spectrometry-based RBP discovery detects peptides from proteins that survive capture and digestion and are abundant enough for identification. Missing a protein from an interactome capture experiment can mean absence, weak crosslinking, poor solubility, low abundance, peptide detection failure, or exclusion by the RNA-capture strategy. Proteomic enrichment is evidence, not a complete census.
As a running example, consider a splicing regulator that binds intronic RNA near an alternative exon. RIP can test whether the pre-mRNA is enriched with the regulator. CLIP can narrow the contact to upstream or downstream intronic sites. A motif model can test whether the sites contain the expected recognition sequence. RNA structure probing can test whether the motif is accessible. Mutating the motif or depleting the RBP can test whether exon inclusion changes. These are related but distinct claims.
CLIP-family methods begin with a biochemical premise: if ultraviolet light creates a covalent bond between an RNA base and a nearby amino acid side chain, then the RNA fragment remains attached to the protein during stringent purification. The method is protein-centric because the bait is a protein, antibody, tag, or tagged RBP complex. The output is an RNA map: reads, peaks, crosslink-induced mutations, truncation sites, motifs, and candidate binding sites for that target in the sampled state. This directionality matters. A CLIP experiment on an RBP asks where that RBP contacts RNA; it does not discover all proteins on a transcript.
A generic CLIP workflow has causal steps. Cells or tissues are irradiated, usually with ultraviolet light. The sample is lysed under conditions intended to preserve crosslinked adducts and reduce noncovalent background. The target protein is immunoprecipitated. RNase trims RNA so that fragments near the protein survive while distal RNA is removed. Protein-RNA complexes are separated, RNA is recovered, adapters or equivalent handles are added, cDNA is synthesized, products are amplified and sequenced, and reads are mapped. Each step can change the apparent map. Weak crosslinking reduces sensitivity. Excess RNase shortens fragments too much. Poor antibody specificity maps the wrong protein. Reverse transcription and ligation create sequence biases. Peak calling imposes statistical assumptions.
HITS-CLIP, or high-throughput sequencing CLIP, is the foundational conceptual member of the family. It uses ultraviolet crosslinking, immunoprecipitation, RNA trimming, and high-throughput sequencing to locate RNA regions associated with a target protein. HITS-CLIP made it practical to connect RBP binding to transcriptome-wide regulatory maps. A classic kind of example is an RBP that binds clusters in introns and untranslated regions, with binding positions compared to alternative splicing or mRNA stability changes.
PAR-CLIP, or photoactivatable-ribonucleoside-enhanced CLIP, changes the chemistry by incorporating photoactivatable ribonucleosides such as 4-thiouridine into newly synthesized RNA. Irradiation at longer wavelength promotes crosslinking, and reverse transcription can produce diagnostic conversions that help identify crosslink-centered sites. The chapter-local PAR-CLIP review by Ascano et al. supports this general logic and its use for RNA-protein interaction networks [Ascano et al. 2012]. Hinze et al. provide a local example of a PAR-CLIP variant designed to expand mapped interaction sites [Hinze et al. 2018]. The biological limitation is that analog incorporation defines the labeled RNA pool. PAR-CLIP is powerful in cells that tolerate labeling, but it may underrepresent stable old transcripts, primary tissues that label inefficiently, organisms with poor analog uptake, or states perturbed by the analog.
iCLIP, or individual-nucleotide-resolution CLIP, emphasizes reverse-transcription truncation at the crosslinked peptide-RNA adduct. In ordinary cDNA library construction, a premature stop can be treated as a failed cDNA. In iCLIP logic, that stop is informative because the reverse transcriptase stopped near the crosslink. Capturing truncated cDNAs converts a biochemical obstacle into a positional signal. This is especially useful when read starts or stops can be modeled as crosslink coordinates rather than treating the full fragment as the binding interval.
eCLIP, or enhanced CLIP, improves scalability and control design. eCLIP workflows are associated with standardized library handling, paired size-matched input controls, and analysis practices suitable for larger RBP atlases. The conceptual advance for readers is not merely a new acronym. eCLIP highlights that a protein-centric binding map must be compared with an input-like background because transcript abundance, fragmentation, and library bias can mimic enrichment.
irCLIP and related streamlined variants modify detection, adapter, or workflow design to improve sensitivity, convenience, reproducibility, or sample requirements. The names of CLIP variants should not be memorized as a linear ranking from old to new. Each variant makes choices about chemistry, cDNA recovery, input requirements, and controls. A variant optimized for low input may sacrifice some quantitative comparability. A variant optimized for crosslink-site localization may require more stringent assumptions about truncation or mutation signals. A variant optimized for throughput may depend heavily on batch correction and standardized controls.
Colocalization CLIP and other subcellular designs add location to the binding definition. An RBP and an RNA can be present in the same cell but separated by nucleus, cytoplasm, organelle, granule, membrane surface, dendrite, or viral replication compartment. Yi et al. provide a local anchor for mapping RNA-protein interactions with subcellular resolution by colocalization CLIP [Yi et al. 2024]. Such approaches are conceptually important because many RNPs are assembled and remodeled in compartments. The relevant biological question may be not simply whether an RBP can bind an RNA, but whether the RBP binds that RNA in a particular place.
Multiplexed CLIP-like strategies try to overcome the one-RBP-at-a-time bottleneck. Wolin et al. describe SPIDR as a multiplexed approach for RNA-protein interaction mapping and connect it to selective translational suppression during cell stress [Wolin et al. 2025]. Multiplexing can reveal coordinated RBP behavior, competition, and condition-dependent redistribution. The tradeoff is comparability. Different RBPs have different crosslinking efficiencies, expression levels, epitope accessibility, and motif properties. Multiplexed data therefore require barcode controls, batch controls, validation of bait identity, and careful evidence labels.
Figure 133.1 summarizes the CLIP-family measurement chain from contact chemistry to biological interpretation.

Figure 133.1. CLIP-family workflow and variant decision points. CLIP-family methods share the logic of covalent RNA-protein capture followed by sequencing, but method variants differ in chemistry, library readout, background control, subcellular resolution, and throughput.
Table 133.1 compares the main method families by capture direction, output, strengths, and caveats.
Table 133.1. Major RNA-protein interaction mapping methods. Method choice should follow the claim being tested. Association, direct contact, atlas discovery, spatial neighborhood, and functional consequence are different evidence targets.
| Method | Capture direction | Crosslinking or labeling chemistry | Primary output | Approximate resolution | Major strengths | Major caveats | Local citation status |
|---|---|---|---|---|---|---|---|
| RIP | Protein-centric capture of bait protein or complex | None, ultraviolet, or formaldehyde depending on design | Co-recovered RNAs measured by RT-qPCR, array, or sequencing | Transcript to complex level; no contact-site map | Simple association assay; useful for stable or large RNP assemblies | Direct contact, pre-lysis existence, and site position remain unresolved without added controls | |
| HITS-CLIP | Protein-centric RBP mapping | 254 nm ultraviolet zero-length crosslinking | Crosslinked RNA fragments, read clusters, candidate binding regions | Regional to narrow fragment clusters | Foundational transcriptome-wide direct-contact workflow | Crosslink yield, RNase trimming, antibody specificity, and peak calling shape the map | |
| PAR-CLIP | Protein-centric RBP mapping in labeled RNA pools | Photoactivatable ribonucleosides such as 4-thiouridine plus longer-wavelength UV | Read clusters with diagnostic conversion-enriched candidate sites | Often nucleotide-proximal when conversion signal is strong | Crosslink-site inference can be sharpened by diagnostic mutations | Requires analog incorporation; labeled pool and cell tolerance constrain interpretation | Ascano et al. 2012; Hinze et al. 2018 |
| iCLIP | Protein-centric RBP mapping | 254 nm ultraviolet crosslinking with truncated cDNA capture | Reverse-transcription stop positions near crosslinked adducts | Nucleotide-proximal truncation clusters | Converts RT stops into positional information | Truncation frequency depends on enzyme, peptide remnant, digestion, and sequence context | |
| eCLIP | Protein-centric RBP mapping and atlas construction | Ultraviolet crosslinking with standardized library and input-control design | Peaks compared with measured size-matched input background | Regional peaks plus method-specific crosslink signals | Scalable control framework for larger RBP atlases | Background controls reduce but do not remove antibody, abundance, and mapping biases | |
| irCLIP | Protein-centric low-input or streamlined CLIP variant | Ultraviolet crosslinking with infrared or simplified library handling | Candidate RBP contact regions from streamlined CLIP libraries | Method-dependent; regional to site-enriched | Can improve handling, sensitivity, or sample requirements | Convenience does not by itself strengthen evidence without matched controls | |
| Colocalization CLIP | Protein-centric mapping constrained by subcellular context | CLIP chemistry combined with subcellular colocalization strategy | Candidate contact sites assigned to compartments or local states | Regional or site-level with compartment qualifier | Tests where an RNA-protein contact occurs, not only whether it occurs | Fraction purity, localization markers, and lysis relocalization affect claims | Yi et al. 2024 |
| Multiplexed profiling | Multi-bait protein-centric mapping | Barcode, tag, or pooled CLIP-like chemistry depending on implementation | Comparative maps across many RBPs or conditions | Method-dependent; comparable only within controlled designs | Reveals coordinated RBP behavior and condition-dependent redistribution | Barcode assignment, bait comparability, batch effects, and validation are central | Wolin et al. 2025 |
| RNA interactome capture | RNA-centric capture of an RNA class | Crosslinking plus RNA purification, often oligo(dT) capture of poly(A) RNA | Candidate RNA-associated proteins identified by mass spectrometry | Protein inventory; usually no RNA contact-site resolution | Expands RBP discovery beyond canonical RNA-binding domains | Misses noncaptured RNA classes, inaccessible RNPs, low-abundance proteins, and exact sites | |
| Proximity labeling | Bait-, RNA-, or compartment-centric neighborhood mapping | Enzymatic biotinylation or related proximity-labeling chemistry | Proteins near a bait or compartment during a labeling window | Local neighborhood; radius and time-window dependent | Captures spatially organized RNP environments | Proximity is not direct RNA binding without direct-contact follow-up |
Box 133.1. Choosing a CLIP Variant by Evidence Need
Start with the claim, not the acronym. For candidate direct-contact sites for one RBP, prioritize validated bait capture, crosslink dependence, replicate concordance, and a readout that localizes crosslink sites or regions. For atlas comparability, prioritize standardized libraries, size-matched input, batch controls, and per-bait quality metrics. For compartment-specific binding, require localization evidence and controls for fraction purity or bait relocalization. For rare samples, a low-input or streamlined CLIP may be appropriate, but sample economy can reduce quantitative comparability. A newer CLIP variant does not repair a bad antibody, poor library complexity, missing control, or mismatched biological question. Report the variant as the measurement design, then state the claim it can support.
The first major boundary case is that a crosslinked site is not necessarily a high-affinity site. Crosslink yield depends on geometry and chemistry as well as occupancy. A low-affinity but well-positioned contact can crosslink efficiently, whereas a biologically important contact can be invisible if the reactive groups do not align. The second boundary case is that a missing CLIP peak is not proof of no binding. The method samples a subset of contacts recoverable under a workflow. The third boundary case is that a CLIP peak can be direct-contact evidence without being functional evidence. Function requires perturbation, natural variation, reporter assays, mutagenesis, kinetic changes, or other orthogonal tests.
RIP is the simplest entry point into RNA-protein association mapping. In RIP, an antibody or affinity tag captures a protein, and associated RNAs are detected by RT-qPCR, microarray, or sequencing. RIP can be performed without covalent crosslinking, after ultraviolet crosslinking, or after formaldehyde crosslinking. The method is attractive because it can be less technically demanding than CLIP and can preserve larger assemblies. It is especially useful when the question is whether an RNA is associated with a protein-containing complex, such as whether a viral RNA is recovered with an antiviral RBP, whether a long noncoding RNA associates with a chromatin regulator, or whether an mRNA class is enriched with a decay factor after stress.
The interpretation of RIP is association, not automatic direct binding. A protein can pull down an RNA because another protein bridges the interaction. A stable RNP can survive lysis even when the bait protein never touches the RNA. RNA released during lysis can bind proteins post-lysis. Highly abundant RNAs can appear enriched by nonspecific carryover. Antibody cross-reactivity can recover paralogous proteins. Beads and tags can add background. For these reasons, noncrosslinked RIP is usually a lower-resolution and lower-directness evidence class than well-controlled CLIP for direct-contact claims.
RIP remains scientifically valuable when the claim is matched to the method. If the biological question is whether a preassembled RNP contains a particular RNA, preserving the assembly may be more important than nucleotide resolution. If the complex is too fragile, too large, or too indirect for ultraviolet CLIP, RIP-like recovery may be the more appropriate assay. For example, some chromatin-associated lncRNA complexes, viral replication assemblies, or granule-associated mRNPs may be interpreted as assemblies rather than direct binary contacts. The correct language is then “associated with” or “recovered with,” not “binds at nucleotide X.”
Crosslinking chemistry controls the strength of a direct-contact claim. Ultraviolet light around 254 nm can create zero-length crosslinks between RNA bases and amino acid side chains already in close contact. Zero-length crosslinking is valuable because it supports direct-contact inference. It is also inefficient. Crosslink formation depends on nucleotide identity, amino acid side chain, RNA conformation, local geometry, irradiation dose, sample thickness, and protein environment. Therefore ultraviolet CLIP enriches a chemically biased subset of direct contacts.
PAR-CLIP chemistry uses photoactivatable ribonucleosides, commonly 4-thiouridine, incorporated into RNA before irradiation. The analog can improve crosslinking and create diagnostic mutation patterns during reverse transcription. This helps distinguish candidate crosslink sites from simple read accumulation. However, analog labeling is not neutral in every system. The labeled pool is biased toward RNAs synthesized during the labeling window, and some cell types, tissues, organisms, or stress states may respond to or poorly incorporate the analog. PAR-CLIP claims should specify the labeling condition and avoid implying that unlabeled RNA populations were measured equally.
Formaldehyde crosslinking has a different evidence meaning. Formaldehyde can bridge protein to protein and protein to nucleic acid across short distances. This can preserve indirect and higher-order assemblies that ultraviolet crosslinking misses. The cost is that direct-contact inference becomes weaker. A formaldehyde-supported RNP association may be exactly the right evidence for a large assembly, but it should not be described as zero-length RNA-protein contact. Evidence grading should distinguish ultraviolet-supported direct contact, photoactivatable-ribonucleoside-supported contact, formaldehyde-supported association, and noncrosslinked association.
RNA interactome capture reverses the direction of the experiment. Instead of immunopurifying one protein and identifying bound RNAs, an RNA class is purified and associated proteins are identified, usually by mass spectrometry. In a common design, cells are crosslinked, polyadenylated RNA is captured with oligo(dT), the sample is washed, and proteins associated with captured RNA are identified by peptide mass spectrometry. The output is a candidate RBPome for the sampled condition.
Interactome capture expanded the practical RBP universe because many recovered proteins lack classical RNA-binding domains. This finding is important but should be interpreted cautiously. Some proteins may contact RNA through noncanonical surfaces. Some may be indirect RNP components. Some may be condition-specific interactors. Some may be contaminants or proteins enriched because they are abundant and sticky under the capture conditions. Direct-contact validation, domain mapping, RNase sensitivity, orthogonal capture, and functional tests are needed before assigning a molecular binding mechanism to every candidate.
The capture chemistry defines blind spots. Oligo(dT)-based methods favor polyadenylated RNA and underrepresent nonpolyadenylated RNAs such as many rRNAs, tRNAs, snRNAs, snoRNAs, histone mRNAs, circular RNAs, many bacterial RNAs, and some viral RNAs. Strong RNA structure can reduce probe accessibility. Membrane-associated or insoluble RNPs can be lost during extraction. Proteins with poor peptide recovery may be missed by mass spectrometry. A protein absent from an interactome capture list is therefore not necessarily absent from RNA biology.
Figure 133.2 contrasts protein-centric, RNA-centric, and proximity-based experimental directionality.

Figure 133.2. Protein-centric, RNA-centric, and proximity-based RBP mapping. RNA-protein interaction methods answer different questions. Protein-centric CLIP localizes candidate contact regions for a target RBP, RIP tests association, RNA-centric interactome capture discovers candidate RNA-associated proteins, and proximity labeling defines local molecular neighborhoods.
Box 133.2. Capture Direction Is Not Binding Mechanism
Ask what was captured before naming the mechanism. RIP captures a protein or protein complex and asks which RNAs came along; the result is association evidence. CLIP captures covalent protein-RNA adducts for one bait; the result can support candidate direct contact at a region or site. RNA interactome capture captures an RNA class and asks which proteins are recovered; the result is an RBP discovery list conditioned by the RNA class and mass-spectrometry sensitivity. Proximity labeling captures proteins near a bait, compartment, or RNA-tethering system during a labeling window; the result is local-neighborhood evidence. A protein can appear in all four assay families for different reasons. Strong reports name the capture direction, biological state, and evidence grade, then reserve words such as direct binding, site, occupancy, and regulation for experiments that actually measured those properties.
For experimental design, the best method follows the claim. Use RIP when the question is assembly membership. Use CLIP when the question is candidate direct contact sites for a defined protein. Use interactome capture when the question is which proteins associate with an RNA class. Use RNA-centric capture or Chapter 138 methods when the question is which proteins bind a particular RNA molecule. Use proximity labeling when local neighborhood is the relevant property. Method choice should be driven by evidence need, not by acronym familiarity.
An RBP atlas is a structured inventory of RNA-associated proteins, RNA sites, cell types, conditions, compartments, and evidence grades. Atlas construction is difficult because RNP biology is not a static list. Proteins move between compartments, change phosphorylation or other post-translational modifications, bind different isoforms, respond to stress, and assemble into transient complexes. A useful atlas therefore needs method metadata: sample type, condition, crosslinking, capture bait, controls, sequencing depth or proteomic depth, confidence score, and claim type.
Proximity labeling addresses spatial and temporal questions that standard CLIP and interactome capture do not solve well. A labeling enzyme such as a biotin ligase or peroxidase can be fused to a protein, directed to a compartment, or recruited near an RNA. During a defined time window, nearby proteins are labeled and later enriched for mass spectrometry or sequencing-linked readout. The evidence means “near the bait under labeling conditions.” It does not automatically mean direct RNA binding. The labeling radius, enzyme expression, reaction time, substrate availability, local environment, and bait localization determine what is captured.
The strength of proximity labeling is that cellular neighborhoods matter. Nuclear speckles, stress granules, P-bodies, nucleoli, mitochondria, endoplasmic-reticulum surfaces, dendrites, viral replication compartments, and translating polysomes contain different RNP populations. A whole-cell interactome can average over those local states. Proximity methods can identify proteins near a compartment-specific RNA or RNP state, giving a more realistic atlas for spatially organized RNA biology.
The weakness is interpretive distance. A protein labeled near a granule marker may bind RNA directly, bind another protein, diffuse through the neighborhood, or be trapped by the labeling environment. A protein labeled near an RNA tether may contact the RNA, contact the tethering apparatus, or occupy the same compartment. Controls must include bait localization checks, inactive enzyme or no-substrate controls, unrelated bait controls, expression-level controls, and orthogonal validation for direct contact. For RBP atlases, proximity evidence should be tagged as local-neighborhood evidence unless direct-contact assays are added.
Subcellular CLIP occupies a related position between direct-contact mapping and spatial atlas construction. Yi et al. describe colocalization CLIP for subcellular resolution [Yi et al. 2024]. The key idea is that direct-contact evidence becomes more interpretable when the map is constrained to where the interaction occurs. A nuclear CLIP signal for a splicing factor has a different meaning from a cytoplasmic signal for the same protein. A localized neuronal mRNA bound in dendrites has a different regulatory context from the same transcript in the soma.
Multiplexed RBP profiling extends atlas construction across many proteins. Wolin et al. provide a local example through SPIDR, a multiplexed method that maps RNA-protein interactions and reveals a stress-linked mechanism of selective translational suppression [Wolin et al. 2025]. Multiplexing is attractive because RNP regulation is combinatorial. Splicing, localization, translation, and decay often depend on several RBPs competing or cooperating on the same transcript. A multi-RBP dataset can reveal patterns that single-protein maps miss.
However, multiplexed atlases require stronger metadata and normalization than single-bait experiments. Different baits vary in expression, tag accessibility, immunopurification efficiency, crosslink yield, and library complexity. A high signal for one RBP and a low signal for another may reflect biology or method efficiency. Batch effects can masquerade as regulatory networks. Atlas entries should therefore carry per-RBP quality metrics, replicate concordance, negative controls, input normalization, and validation status.
RBP atlas construction also depends on definitions. A narrow atlas may include only proteins with direct RNA-contact evidence. A broader atlas may include proteins recovered by RNA capture, proximity labeling, or RNP co-purification. Both definitions can be useful if labels are clear. The problem arises when all atlas entries are presented as equivalent “RNA binders.” A mitochondrial metabolic enzyme recovered by interactome capture, a canonical splicing factor with nucleotide-resolution CLIP peaks, and a proximity-labeled granule protein are different evidence objects.
Figure 133.3 shows an evidence-aware atlas design in which each entry records method, context, directness, resolution, and validation.

Figure 133.3. Evidence-aware RBP atlas schema. A useful RBP atlas records what was measured, in which biological state, by which method, and with which evidence strength. Direct-contact CLIP evidence, RNA-centric capture, proximity labeling, and indirect RNP recovery should not be collapsed into the same claim type.
RBP atlases should also include negative and boundary information. If an RBP was tested and not recovered under a condition, that absence is useful when the assay had enough sensitivity. If a protein binds RNA only after stress, infection, differentiation, or drug treatment, the condition qualifier is part of the claim. If an RBP binds only newly synthesized RNA in a PAR-CLIP labeling window, that temporal qualifier should travel with the record. Such qualifiers make the atlas more useful than a flat protein list.
Peak calling converts sequencing data into candidate binding regions. The input is usually aligned reads from CLIP, RIP, or a related method, often with control libraries. The output may be a genomic interval, transcript interval, crosslink site, cluster, score, false discovery estimate, motif, or differential binding call. The peak is a computational inference produced by a pipeline. It is not directly visible in the cell, and it is not automatically a functional regulatory element.
The first computational problem is mapping. CLIP reads are often short and may contain mutations, deletions, adapter remnants, or low-complexity sequence. Reads can map ambiguously to paralogous genes, repetitive elements, pseudogenes, rRNA repeats, mitochondrial sequences, transposable elements, or highly similar isoforms. Pre-mRNA reads may map to introns; mature mRNA reads may map across exon junctions; small RNA reads may map to many family members. The reference annotation and alignment policy can change the apparent distribution of binding.
The second problem is deduplication. PCR amplification can create many identical reads from one molecule. Unique molecular identifiers help distinguish original molecules from amplification copies, but they have their own error models. Over-aggressive deduplication can remove real biological abundance when many molecules start at the same crosslink site. Under-aggressive deduplication can inflate a peak. The correct strategy depends on library design.
The third problem is background. Highly expressed transcripts naturally produce more fragments. Some transcript regions are more accessible to nuclease digestion or ligation. Some sequences reverse transcribe more efficiently. Size-matched input controls, no-antibody controls, knockout controls, and replicate models help distinguish target-specific enrichment from background. eCLIP-style size-matched input logic is important because it treats the background as a measured library rather than a purely theoretical expectation.
Crosslink-induced mutation sites provide positional information beyond pileups. In PAR-CLIP, diagnostic conversions can mark candidate contact positions in compatible systems [Ascano et al. 2012]. In iCLIP-like workflows, cDNA truncation positions can indicate the peptide-RNA adduct. In other workflows, deletions, substitutions, or read-start enrichment may carry crosslink information. These signals are powerful because a narrow crosslink-site cluster is harder to explain by transcript abundance alone than a broad pileup.
Yet crosslink-induced signals are not universal. Some proteins do not produce strong diagnostic mutations. Some reverse transcriptases bypass lesions differently. Sequence context affects error rates. Peptide remnants can block, misincorporate, or be removed depending on proteinase digestion. Filtering thresholds can create apparent precision by discarding ambiguous reads. A missing crosslink-induced mutation does not prove absence of binding, and a mutation hotspot does not prove a regulatory effect.
Peak callers use different assumptions. Some identify enriched clusters relative to local background. Some model crosslink-induced mutation enrichment. Some use replicate reproducibility. Some integrate motif information. Some call differential binding between conditions. A peak caller suitable for a deeply sequenced eCLIP experiment may not fit a low-input irCLIP dataset, and a caller built for narrow crosslink sites may not fit broad RNP footprints. Method-specific reporting should state the caller, reference annotation, controls, replicate handling, filtering thresholds, and false-discovery approach.
Binding-site models connect candidate sites to recognition logic. A simple model says that the RBP recognizes a short sequence motif. A better model asks whether the motif is single-stranded or structured, whether neighboring motifs cooperate, whether the site lies in an intron, 3′ UTR, coding sequence, or noncoding RNA, whether the motif is conserved, whether RNA modifications alter recognition, and whether other RBPs compete nearby. Rapakko et al. emphasize that RNA structure and protein binding interact, so structure-aware interpretation is often necessary [Rapakko et al. 2026].
The best binding models explain negative cases as well as positive cases. If a motif appears thousands of times but only a subset is bound, the model should explain why the unbound sites are not occupied. Possible reasons include local RNA structure, lack of colocalization, low transcript expression, competing proteins, translation through the site, rapid RNA decay, or absence of the required cofactor. If a bound site lacks the canonical motif, the model should consider noncanonical motifs, structure recognition, indirect complex membership, annotation errors, or false-positive peaks.
Figure 133.4 illustrates the computational inference path from reads to peaks, crosslink sites, motifs, and functional hypotheses.

Figure 133.4. Computational inference from CLIP reads to biological hypotheses. A CLIP peak is an inferred interval produced by a pipeline. Crosslink-induced mutations, truncations, controls, and replicates can sharpen localization, but functional interpretation requires independent evidence.
Table 133.2 summarizes common analysis choices and artifacts that affect peak interpretation.
Table 133.2. Peak-calling and binding-model artifacts. Peak calling and binding-site modeling transform biased molecular libraries into candidate biological sites. Reporting the assumptions is part of the evidence.
| Analysis step | Assumption | Common failure | Control or mitigation | Recommended reporting |
|---|---|---|---|---|
| Adapter trimming | Adapter sequence is separable from genuine RNA fragment sequence | Over-trimming removes biological bases; under-trimming creates false mismatches | Use library-specific adapters, length filters, and trimming QC | Trimmer, adapter set, minimum length, and retained-read fraction |
| Alignment | Reads can be assigned to the intended genome or transcriptome reference | Short, mutated, or junction-spanning reads map incorrectly or are lost | Use appropriate genome, transcriptome, splice, and small-RNA references | Aligner, reference build, annotation version, mismatch policy |
| Multimapping | Ambiguous reads can be handled without distorting target distribution | Repeats, paralogs, pseudogenes, rRNA, and isoforms inflate or erase peaks | State discard, fractional, unique-only, or rescue policy; inspect abundant classes | Multimapper policy and affected RNA classes |
| UMI deduplication | Molecular barcodes separate original molecules from PCR copies | Barcode errors split molecules; aggressive collapsing removes true co-starting fragments | Correct barcode errors and match deduplication to library design | UMI length, error model, deduplication rule, complexity metrics |
| Background modeling | Control libraries approximate non-target recovery and library bias | Transcript abundance, nuclease accessibility, and ligation bias mimic enrichment | Size-matched input, no-antibody or knockout controls, replicate background models | Control type, normalization, statistical model, false-discovery threshold |
| Crosslink-induced mutation analysis | Diagnostic substitutions or deletions mark peptide-RNA adduct positions | Polymerase error, sequence context, filtering, or analog bias creates false precision | Compare with controls and require enrichment above local error patterns | Mutation type, thresholds, control rate, eligible chemistry |
| Truncation analysis | Reverse-transcription stops near crosslinks report contact positions | RT stops also arise from structure, damage, enzymes, or incomplete digestion | Capture truncated cDNAs and model stop enrichment against controls | Stop-site definition, RT enzyme, filtering, and replicate concordance |
| Motif discovery | Enriched peaks contain features that explain RBP specificity | Short motifs appear broadly and can be enriched by composition or annotation bias | Compare matched background regions and integrate structure or region context | Motif method, background set, region class, and negative-site behavior |
| Structure integration | Accessibility or structure measurements reflect binding-relevant RNA conformations | Structure probing and CLIP sample different states or perturb each other | Use matched condition, transcript abundance filters, and orthogonal validation | Structure assay, condition match, integration model, discordant cases |
| Differential binding | Normalized peak differences reflect biological redistribution | Expression changes, batch effects, or library depth masquerade as binding change | Include input abundance, replicates, batch design, and effect-size thresholds | Contrast design, normalization, covariates, and reproducibility criteria |
Functional inference requires a separate step. A peak near an alternative exon suggests a possible splicing regulatory site, but the site becomes functionally supported only when perturbing the RBP or site changes exon inclusion in a compatible system. A 3′ UTR peak suggests possible effects on stability, localization, translation, or miRNA cooperation, but the specific output must be measured. A viral RNA peak suggests host or viral RNP assembly, but infection-stage controls and RNA abundance controls are essential. Binding is a necessary part of many mechanisms, but binding alone is rarely sufficient evidence for mechanism.
Stoichiometry asks how many molecules are bound, or what fraction of a transcript population is occupied at a site. This question is harder than site detection. CLIP read counts are affected by RNA abundance, crosslinking efficiency, digestion, ligation, reverse transcription, PCR amplification, sequencing depth, mapping, and filtering. A high peak can represent high occupancy, high transcript abundance, favorable crosslink chemistry, or efficient library recovery. A low peak can represent low occupancy, poor crosslinking, low expression, or technical loss. Therefore read count alone is usually not an absolute occupancy measurement.
The distinction matters when two transcripts are compared. Suppose an RBP produces 500 reads at a site in a highly expressed housekeeping mRNA and 50 reads at a site in a rare developmental transcript. The first site may look stronger in a genome browser, but the second site could represent a larger fraction of available RNA molecules if the rare transcript is present at low copy number. Conversely, a weak site in a very abundant transcript may still represent many physical RNPs per cell. Quantitative interpretation therefore needs both numerator and denominator: recovered RBP-linked fragments and the size of the RNA population from which those fragments could have arisen. This is why input RNA abundance, spike-ins, unique molecular identifiers, and orthogonal abundance measurements are not cosmetic additions when the claim uses words such as high occupancy, low occupancy, saturated, or stoichiometric.
Fractional occupancy requires calibration. A quantitative design may need input RNA abundance, spike-ins, unique molecular identifiers, protein abundance, crosslinking controls, replicate models, and independent measurements such as imaging, electrophoretic mobility shift assays, reporter titrations, or targeted biochemical validation. Even then, the result is condition-specific. A transcript may be highly occupied during stress and weakly occupied during growth. A site may be occupied only during splicing, only during nuclear export, or only while the ribosome is absent.
Cell-type specificity is not a small correction to an atlas. It is often the biological phenomenon. Neurons contain long transcripts, dendritic localization programs, local translation, and granule transport. Immune cells rapidly remodel RNA decay and translation after activation. Stem cells and embryos use stage-specific RBPs and RNA stores. Cancer cells alter RBP expression, splicing programs, stress signaling, and RNA modification states. Viral infection introduces viral RNAs, host shutoff, innate immune activation, and viral RNPs. A map from one proliferating cell line can be a useful starting point but should not be treated as a universal binding map.
Compartment specificity is equally important. A splicing factor may bind nascent introns in the nucleus and a different RNA class in the cytoplasm. An RBP may shuttle to stress granules, P-bodies, mitochondria, nucleoli, nuclear speckles, or viral replication compartments. Whole-cell lysate averages these states. Fractionation, proximity labeling, colocalization CLIP, imaging, and compartment-specific controls help separate maps that otherwise look contradictory. Yi et al. support the principle that subcellular resolution can be built into RNA-protein interaction mapping [Yi et al. 2024].
Stoichiometry and cell specificity meet in dosage-sensitive biology. A modest change in fractional occupancy can have a large effect if the bound RNA controls a developmental switch, immune response, or neuronal synapse. Conversely, a strong peak on a highly abundant transcript may have little functional consequence if only a small fraction is affected or if redundant RBPs compensate. Quantitative claims should therefore avoid simple equations such as “more reads means more regulation.”
Table 133.3 lays out the evidence needed to distinguish detection, enrichment, fractional occupancy, and regulatory effect.
Table 133.3. Stoichiometry and cell-type-specific interpretation. Detection, enrichment, fractional occupancy, context specificity, and functional consequence require different evidence. Read depth alone is rarely enough for quantitative occupancy.
| Claim type | Minimum evidence | Quantitative caveat | Example | Preferred wording |
|---|---|---|---|---|
| Detected site | Reproducible reads, peak, truncation, or mutation signal above filtering threshold | Detection can reflect abundance, recovery bias, or favorable crosslink geometry | CLIP reads cluster in an intron near an alternative exon | The RBP has a candidate contact signal at this region in this sample |
| Enriched site | Signal exceeds matched input, negative control, or local background model | Enrichment is relative and does not equal percent occupancy | eCLIP peak remains above size-matched input in a 3′ UTR | The region is enriched for RBP-associated fragments relative to control |
| Fractionally occupied site | Calibrated numerator and denominator with abundance normalization or orthogonal quantification | No universal conversion from CLIP reads to occupied RNA molecules exists | Spike-in-normalized site estimates during stress versus growth | A calibrated fraction of this RNA population is occupied under the stated condition |
| Compartment-specific site | Direct-contact or proximity evidence tied to validated fraction or localization marker | Whole-cell averages can hide relocalization and fraction contamination | Colocalization CLIP assigns a splicing-factor contact to nuclear RNA | The contact is detected in the specified compartment under the assayed conditions |
| Cell-type-specific site | Comparable maps or targeted validation across defined cell types | Different expression levels and library efficiencies can mimic specificity | Neuronal RBP signal on a dendritically localized transcript | The site is detected in this cell type and not detected, or is weaker, in the comparator under matched controls |
| Functionally active site | Binding evidence plus perturbation of RBP or site changes an RNA output | Functional effect can be condition-specific, buffered, or indirect | Mutating an intronic motif changes exon inclusion after RBP perturbation | The site contributes to the measured regulatory output in this system |
Single-cell and spatial extensions are attractive but challenging. Most CLIP-family methods require enough material for crosslinking, immunoprecipitation, and library construction, making true single-cell direct-contact maps difficult. Low-input and indexed methods can move toward rare populations, but dropout, amplification, and antibody background become more severe. Spatial transcriptomic data can suggest where RBPs and RNAs coexist, but colocalization at tissue scale is not molecular contact. The strongest spatial claims combine expression, localization, proximity, and direct-contact evidence.
Cross-species comparisons add another layer. An RBP motif may be conserved while target transcripts change. An RBP family may expand in one lineage. Plants, fungi, bacteria, animals, organelles, and viruses have different transcript architectures and RNP machines. A mammalian 3′ UTR binding rule may not transfer to bacterial polycistronic RNA or plant stress granules. Comparative maps are most useful when they distinguish conserved recognition chemistry from lineage-specific RNA architecture.
Until those references are added, broad quantitative statements in this chapter should remain labeled as method-principle synthesis rather than citation-locked consensus.
Controls define the strongest claim that an RNA-protein interaction experiment can support. For RIP, essential controls include input RNA, nonspecific antibody or tag controls, RNase sensitivity when appropriate, validation of bait recovery, and orthogonal testing of candidate RNAs. For CLIP, essential controls include antibody specificity, input or size-matched input, biological replicates, library complexity, crosslink dependence, peak reproducibility, and appropriate negative regions or negative proteins. For interactome capture, controls include no-crosslink or RNase controls, noncaptured fractions, replicate proteomics, abundance correction, and validation of candidate RBPs. For proximity labeling, controls include enzyme localization, inactive or no-substrate conditions, unrelated bait controls, labeling-time controls, and direct-contact follow-up.
Antibody specificity is a recurring failure point. A cross-reactive antibody can recover related proteins and create a composite map. An antibody can work in western blotting but fail in immunoprecipitation. Epitope tags can improve capture specificity, but overexpression can create nonphysiological binding, and the tag can interfere with localization or RNP assembly. Endogenous tagging is often cleaner but can be difficult in tissues, primary cells, or organisms with limited genetics. A strong map documents bait identity and expression rather than assuming the antibody defines the target.
Lysis creates opportunities for artifacts. RNA and proteins that are separated in living cells can meet after membranes break. Highly abundant RNAs can bind exposed RNA-binding surfaces. Salt, detergent, magnesium, heparin, competitor RNA, temperature, and processing time can all alter complexes. Crosslinking before lysis reduces but does not eliminate these issues. Noncrosslinked RIP is especially vulnerable to post-lysis reassociation, so its claims should be phrased conservatively unless supported by additional evidence.
RNase digestion is both necessary and dangerous. Partial digestion improves CLIP resolution by trimming RNA to protein-protected fragments. Too little digestion creates broad peaks that merge nearby sites. Too much digestion can destroy real signal or favor unusually protected structures. RNases also have sequence and structure preferences. Reporting digestion conditions and fragment-size distributions is therefore part of interpreting resolution.
Library construction introduces biases that are easy to forget because they occur after the biological sample is gone. RNA-end chemistry affects adapter ligation. Structured fragments ligate and reverse transcribe differently. Crosslinked peptide remnants can stop reverse transcriptase or cause misincorporation. PCR amplification can dominate low-complexity libraries. Size selection can remove short or long products. These biases are not fatal, but they mean that a sequencing library is a filtered representation of recovered fragments.
Peak-calling artifacts often arise from abundant RNAs. Ribosomal RNA, mitochondrial RNA, small nuclear RNA, repetitive elements, and highly expressed housekeeping transcripts can dominate reads. Some are true RBP targets; others are background. Blacklisting all abundant RNAs can remove biology, but accepting every abundant peak can inflate false positives. The analysis should state how abundant and repetitive regions were handled and whether the same rules were applied to controls.
Evidence grading protects readers from overinterpretation. Association evidence means an RNA and protein or complex were recovered together. Direct-contact evidence means chemistry and purification support physical contact. Site-localization evidence means reads, truncations, or mutations identify a region or nucleotide. Occupancy evidence means the fraction or amount bound has been calibrated. Functional evidence means perturbation changes an RNA output through the site or interaction. These levels can be combined, but one level does not automatically imply the next.
This vocabulary is especially useful when multiple methods disagree, because it allows discordant results to be compared without forcing all assays into one binary binding category.
Box 133.3. Upgrading a Peak to a Mechanistic Claim
Treat each phrase as an evidence upgrade. A reproducible peak supports a candidate contact signal. Direct-contact language needs crosslinking and target-specific purification. Nucleotide-proximal language needs enriched truncation, conversion, deletion, or other crosslink-centered signal. Specificity claims need a model that explains bound sites and unbound motif-containing sites. Occupancy claims need a denominator, such as input abundance, calibration, spike-ins, UMIs, or orthogonal counting. Functional claims need perturbation of the RBP or site and measurement of a relevant RNA output, ideally with rescue or endogenous validation. When an upgrade is missing, use the weaker wording explicitly. “Recovered with,” “candidate contact region,” “localized crosslink signal,” “estimated occupancy,” and “contributes to regulation” are not interchangeable.
Table 133.4 provides a practical evidence-grade vocabulary for reporting RNA-protein interaction claims.
Table 133.4. Evidence-grade vocabulary for RNA-protein interaction claims. Strong reporting distinguishes what was measured from what was inferred. The same dataset can support one evidence level while leaving higher levels unresolved.
| Evidence level | Methods that can support it | Insufficient overclaim | Stronger phrasing | Validation upgrade |
|---|---|---|---|---|
| Association | RIP, RNP co-purification, formaldehyde-supported recovery, RNA-centric capture | This protein directly binds this RNA site | The RNA is recovered with the protein or complex under these conditions | Add ultraviolet CLIP, mutational validation, or reconstitution |
| Candidate direct contact | UV CLIP, PAR-CLIP, HITS-CLIP, eCLIP, irCLIP with target-specific controls | The site is a regulatory element | The region is a candidate direct-contact site for the assayed RBP | Add site perturbation and RNA-output measurement |
| Localized contact site | Crosslink-induced mutation, truncation, or narrow reproducible peak evidence | The diagnostic nucleotide is functionally required | Crosslink signatures localize the candidate contact near this position | Mutate the site and test binding plus function |
| Quantitative occupancy | Calibrated CLIP-like design, spike-ins, UMIs, RNA abundance, orthogonal quantification | Read count equals occupancy | Occupancy is estimated for this site under a specified calibration model | Validate with imaging, titration, biochemical binding, or targeted quantification |
| Functional consequence | RBP perturbation, site mutagenesis, reporter assay, endogenous editing, biochemical rescue | Binding alone proves regulation | Perturbing the RBP or site changes the measured RNA output | Demonstrate rescue, directness, and condition specificity |
| Atlas entry | CLIP atlas, interactome capture, proximity atlas, curated evidence-grade database | This protein is universally an RBP | The entry is supported by the listed method, context, and evidence grade | Add orthogonal method, cell-state metadata, and direct-contact validation |
| Local neighborhood | Proximity labeling, compartment-restricted labeling, subcellular profiling | The protein touches the RNA | The protein is near the bait or compartment during the labeling window | Pair with CLIP, RNA capture, or direct biochemical validation |
Several interpretation errors recur in this field. RIP enrichment is not proof of direct binding. A CLIP peak is not automatically a regulatory element. A diagnostic mutation is not automatically a functional nucleotide. Absence of a peak is not proof of absence. A motif under a peak is not sufficient to define specificity. Interactome capture is not a complete RBP census. Proximity labeling means nearby under labeling conditions, not necessarily touching RNA. Read counts are not direct stoichiometry without calibration.
The current consensus is therefore cautious but not pessimistic. RNA-protein interaction mapping has transformed RNA biology because it connects proteins, RNA regions, cell states, and regulatory hypotheses at scale. The field is strongest when maps are treated as evidence-rich measurements with known biases. The strongest studies combine direct-contact chemistry, stringent purification, matched controls, reproducible analysis, structure-aware interpretation, quantitative caution, and functional perturbation. The open frontier is not simply higher throughput. It is better integration of binding, structure, compartment, stoichiometry, dynamics, and causality.
Figure 133.5 summarizes the evidence ladder from association to direct contact, site localization, occupancy, and function.

Figure 133.5. Evidence ladder for RNA-protein interaction claims. Evidence strength increases when association evidence is combined with direct-contact chemistry, site-localizing readouts, quantitative calibration, and functional perturbation.
Current consensus can be stated as five principles. First, the method determines the claim. RIP, CLIP, interactome capture, proximity labeling, and multiplexed profiling measure overlapping but nonidentical properties. Second, direct-contact and site-localizing claims require chemistry and readouts that support those claims. Third, RNA structure and RNP context affect binding-site interpretation [Rapakko et al. 2026]. Fourth, RBP maps are cell-state and compartment specific, as emphasized by subcellular and multiplexed approaches [Yi et al. 2024; Wolin et al. 2025]. Fifth, function requires evidence beyond binding, usually perturbation, mutagenesis, reporter assays, endogenous editing, biochemical reconstitution, or convergent multi-omic evidence.
The field is also converging on more explicit control reporting. Size-matched input, replicate reproducibility, antibody validation, library complexity, and background modeling have become central to interpreting CLIP-family data. For atlas work, the strongest resources separate proteins with direct-contact evidence from proteins with RNA-centric capture evidence, proximity evidence, or indirect RNP evidence. For computational modeling, there is growing recognition that CLIP datasets are experimental labels with noise and bias, not perfect ground truth for biological function.
Open questions:
Controversies:
Deprecated or weakened claims: