This chapter owns experimental and computational technologies that convert RNA-RNA or RNA-chromatin proximity into sequencing evidence. It treats ChIRP-seq, CHART-seq, RAP, RNA-centric chromatin capture, GRID-seq-like RNA-DNA mapping, SPLASH, PARIS, COMRADES, LIGR-seq, chimeric-read analysis, bias controls, interaction-map integration, orthogonal validation, perturbation design, and inference limits. The central distinction is that these methods recover contacts under defined chemistry and sample handling; they do not by themselves prove direct base pairing, direct RNA-DNA binding, regulatory consequence, or a biological RNA-chromatin mechanism. Mechanistic recruitment models, PRC2-RNA specificity, enhancer/promoter RNA biology, chromatin remodeling, histone modification, and causal regulatory synthesis belong to Chapter 96.
RNA-RNA and RNA-chromatin interaction mapping methods are proximity-conversion assays. They preserve or enrich molecular neighborhoods, turn those neighborhoods into recoverable nucleic-acid molecules, and sequence the products. In RNA-RNA mapping, the characteristic output is often a chimeric read, meaning a sequencing read or read pair whose arms map to two RNA segments that were joined, copied, or otherwise linked during library construction. In RNA-chromatin mapping, the characteristic output is an enriched genomic interval, an RNA-by-genomic-locus matrix, or a set of DNA peaks recovered with an RNA or RNA population.
The reason these methods matter is that conventional RNA-seq measures abundance and sequence but not molecular neighborhood. RNA-centric capture can report where a defined RNA is enriched on chromatin; global RNA-DNA methods can produce RNA-by-locus matrices; RNA-RNA ligation methods can nominate intra- or intermolecular contact intervals. These outputs support mechanistic hypotheses, but the biological synthesis of how chromatin-associated RNAs recruit regulators or alter chromatin belongs to Chapter 96.
The reason these methods are difficult is that proximity is not mechanism. A chimeric RNA read can arise from direct base pairing, RNA-protein-bridged proximity, compartmental co-localization, random ligation after lysis, reverse-transcriptase template switching, PCR recombination, or ambiguous mapping. RNA enrichment at chromatin can reflect direct RNA-DNA hybrid formation, protein-mediated tethering, nascent transcription, local abundance, chromatin accessibility, nuclear compartmentalization, or capture carryover. The same dataset can therefore support a strong discovery claim and a weak mechanistic claim at the same time.
Current practice treats contact maps as evidence layers. A credible interaction is reproducible, control-sensitive, mappable, and robust to reasonable parameter choices. A credible direct duplex is supported by duplex-biased chemistry, sequence compatibility, structure-probing consistency, and ideally mutational disruption plus compensatory rescue. A credible RNA-chromatin regulatory mechanism requires perturbation of the RNA or contact, measurement of chromatin or transcriptional consequences, and separation of local transcriptional effects from direct action of the RNA molecule.
The reader should keep three concepts separate: structure, proximity, and function. RNA structure describes base pairs, helices, junctions, tertiary contacts, and RNP architecture. RNA proximity describes molecules or segments that are close enough for an assay to recover them. RNA function describes a biological consequence such as transcriptional regulation, splicing control, RNA stability, translation, replication, packaging, or chromatin modification. A proximity assay can support a structural or functional model, but a proximity assay alone does not establish the model.
Several examples recur throughout the chapter. Xist-like long noncoding RNAs provide a chromatin-localization example because the biological question is where a long RNA accumulates across a chromosome and how that accumulation relates to silencing. A promoter-associated or enhancer-associated RNA provides a local chromatin example because the RNA may simply be near its own transcription site. A viral RNA genome provides an RNA-RNA example because long-range intramolecular contacts can connect distant genomic regions. A cytoplasmic mRNA provides a cautionary example because abundant translated RNAs can produce contacts through local concentration or protein-bound states without a specific regulatory duplex.
The methods in this chapter sit between molecular biology and statistical genomics. On the molecular side, crosslinking, fragmentation, hybridization, ligation, reverse transcription, and PCR determine what enters the library. On the statistical side, alignment, filtering, normalization, clustering, replicate analysis, and false-discovery control determine what becomes a reported contact. Good interpretation requires both levels. A biochemically plausible contact with poor mapping remains uncertain, and a statistically strong contact with no biochemical controls remains mechanistically underdefined.
The local chapter bibliography contains verified references for a plant 3D genome architecture review by Ouyang et al. 2020 and a 2025 Bio-protocol method article by Ding et al. on simultaneous capture of chromatin-associated RNA and global RNA-RNA interactions. The chapter-local references do not yet include verified landmark papers for ChIRP-seq, CHART-seq, RAP, GRID-seq, SPLASH, PARIS, COMRADES, and LIGR-seq. This draft therefore treats those method families using general, non-fabricated method logic and flags the missing references where historical or protocol-specific claims need curation.
RNA-centric chromatin capture begins with a defined RNA and asks where that RNA is enriched on chromatin. The common workflow has six conceptual steps. First, cells or nuclei are fixed so RNA-containing chromatin complexes survive extraction. Second, chromatin is fragmented, usually to make genomic DNA recoverable at useful interval resolution. Third, biotinylated antisense oligonucleotides hybridize to the target RNA. Fourth, streptavidin beads recover the RNA and material crosslinked to it. Fifth, associated DNA, RNA, or protein is purified depending on the assay design. Sixth, sequencing or mass spectrometry identifies the recovered molecules. The core logic is simple: if a genomic region is consistently enriched when the RNA is captured, the region is a candidate RNA-associated chromatin site.

Figure 134.1. From Molecular Proximity to Sequencing Evidence. These assays convert selected molecular neighborhoods into sequencing signals; the output is not automatically a mechanism.
ChIRP-seq, or chromatin isolation by RNA purification followed by sequencing, is a prominent version of this logic. ChIRP-style designs typically tile many short biotinylated probes across the target RNA. A common design principle is to split probes into independent odd and even pools. A true RNA-dependent chromatin enrichment should appear in both pools, whereas a probe-specific off-target event may appear only in one pool. The odd-even design does not remove every artifact, but it makes probe-specific hybridization artifacts more visible. ChIRP-style experiments are often used for long noncoding RNAs because lncRNAs can be long enough to support many probes and because many lncRNA hypotheses involve chromatin localization.
CHART-seq, or capture hybridization analysis of RNA targets followed by sequencing, uses related antisense capture logic but is often described as placing special emphasis on empirical identification of accessible regions of the target RNA. The reason accessibility matters is that an RNA in a fixed RNP is not an unfolded string. Regions may be base-paired, protein-covered, chemically modified, or buried in a large complex. A probe can be perfectly complementary by sequence and still fail to bind in the actual sample. Empirical probe selection is therefore a biochemical recognition problem, not merely a sequence-design problem.
RAP, or RNA antisense purification, also uses tiled antisense hybridization but is associated with longer probe tiling and stringent hybridization conditions in several applications. Stringent conditions can reduce nonspecific binding and improve confidence that recovered material depends on the target RNA. The tradeoff is that stringent hybridization and washing can reduce recovery of fragile or low-abundance complexes, and long probes can create different off-target risks than short probes. ChIRP, CHART, and RAP are therefore not interchangeable acronyms. They share RNA-centric capture logic, but differ in probe design, hybridization, wash stringency, target classes, downstream readouts, and validation expectations.
The main reader-facing question is what an RNA-chromatin peak means. A peak is an enriched genomic interval recovered with the RNA capture, usually after comparison with input and controls. It may represent direct RNA-DNA hybrid formation, triplex-like nucleic-acid recognition, a protein bridge between RNA and chromatin, nascent RNA remaining near its transcription site, a shared nuclear compartment, a highly accessible sticky locus, or off-target probe recovery. Direct RNA-DNA binding is only one mechanism. A chapter on RNA-chromatin interactions can discuss mechanistic possibilities in detail, but a methods chapter must first protect the distinction between occupancy and mechanism.
Box 134.1. What an RNA-Chromatin Peak Can Mean
An RNA-chromatin peak says that a genomic interval was enriched with an RNA capture or RNA-DNA proximity workflow. It does not say, by itself, how the RNA reached that interval. Plausible explanations include direct RNA-DNA hybrid formation, triplex-like recognition, binding through an RNA-binding or chromatin protein, retention of nascent RNA near its own gene, colocalization in a nuclear compartment, or recovery of an accessible sticky locus. The first interpretation to write down should be the assay-level statement: “this interval was recovered with this RNA under these conditions.” Stronger wording needs stronger tests. Independent probe pools support target-dependent recovery. RNA depletion tests dependence on the RNA. RNase H or R-loop assays may support RNA-DNA hybrid models. Imaging tests whether the RNA and locus are near each other in intact nuclei. Perturbation and rescue are needed before calling the peak regulatory.
A concrete example clarifies the logic. Suppose a researcher asks whether an Xist-like lncRNA occupies a chromosome domain. An RNA-centric capture experiment can recover DNA intervals enriched with the lncRNA. Broad enrichment over one chromosome, reproducible in independent probe pools and reduced when the RNA is depleted, supports a localization model. However, the experiment does not automatically show which protein complexes bind the RNA, whether the RNA directly recognizes DNA, whether spreading is active or passive, or whether the occupancy causes silencing. Those questions require additional protein mapping, chromatin-state measurements, RNA-domain perturbation, imaging, and rescue logic.
Promoter-associated and enhancer-associated RNAs create a different boundary case. These RNAs are often produced at the locus where they are detected. If an RNA-centric capture recovers the encoding locus, the signal may reflect nascent transcription and local retention rather than a trans-acting chromatin-targeting mechanism. This does not make the signal uninteresting. Nascent RNA can influence local chromatin through transcription-coupled mechanisms, R-loop formation, co-transcriptional RNP assembly, or local recruitment of regulatory factors. But the interpretation should be local until evidence shows that a mature RNA molecule leaves its transcription site and acts elsewhere.
Controls for RNA-centric chromatin capture must match the failure mode. Independent probe pools address probe-specific off-targets. Scrambled or sense probes address nonspecific oligonucleotide and bead binding. Input chromatin controls reveal genomic regions that are abundant, accessible, or repetitive. RNA depletion or deletion tests whether recovery depends on the target RNA. Rescue with a probe-resistant RNA can test whether depletion effects are specific. RNase treatment, strand-specific perturbation, and R-loop mapping can help distinguish RNA-DNA hybrid models from protein-bridged models. Imaging can test whether a broad chromatin domain really contains the RNA in intact nuclei.
The evidence basis for ChIRP-seq, CHART-seq, and RAP should ultimately rest on verified method papers, replicate benchmarks, and well-controlled biological examples. The local reference file for this chapter explicitly lacks those verified landmark references. Therefore, this draft uses these methods as established method families but avoids assigning historical priority, numerical performance claims, or specific original-study conclusions. A future reference-curation pass should add verified primary papers and method reviews before finalizing method-specific history or claims about benchmark RNAs.
Table 134.1. Comparison of RNA-Centric Chromatin Capture Methods. The methods share antisense capture logic but differ in probe strategy, accessibility assumptions, and validation needs.
| Method family | Primary question | Capture logic | Characteristic design feature | Typical output | Major strengths | Major limitations | Required controls | Local reference status |
|---|---|---|---|---|---|---|---|---|
| ChIRP-seq | Where does a defined RNA associate with chromatin? | Crosslinked material is hybridized to biotinylated antisense probes that recover target RNA and associated DNA. | Many short tiled probes, often split into independent odd and even pools. | Enriched genomic intervals or DNA peaks linked to the captured RNA. | Probe-pool concordance helps expose single-probe off-targets; long RNAs can support many probes. | Probe accessibility, off-target hybridization, RNA abundance, accessible chromatin, and occupancy-without-mechanism remain issues. | Odd-even concordance, input chromatin, scrambled or sense probes, target RNA depletion, RNase or imaging checks when relevant. | Landmark ChIRP citation is missing locally; use as a method-family description only. |
| CHART-seq | Which chromatin intervals are recovered with empirically accessible regions of a target RNA? | Antisense hybridization captures the RNA and crosslinked chromatin from fixed material. | Probe design emphasizes experimentally accessible target-RNA regions. | Target-RNA-associated DNA enrichment across genomic intervals. | Accessibility testing can reduce ineffective probes and clarify the capture region. | Accessible probe sites may not represent all RNA domains; enrichment still does not prove direct binding or regulation. | Accessible-region validation, independent probes, input and sense-probe controls, target depletion, locus-specific validation. | Landmark CHART citation is missing locally; method-specific history needs curation. |
| RAP | What DNA, RNA, or proteins co-purify with a target RNA under stringent antisense purification? | Tiled antisense probes recover the target RNA from crosslinked material for DNA, RNA, or protein readouts. | Longer tiled probes and stringent hybridization or wash conditions in many applications. | Genomic occupancy peaks, co-recovered RNAs, or protein partners depending on the readout. | Stringency can reduce nonspecific recovery and supports multiple downstream assays. | Fragile or low-abundance complexes can be lost; long probes create their own off-target profile. | Negative probes, input material, target depletion, orthogonal capture or imaging, recovery and specificity checks. | Landmark RAP citation is missing locally; avoid protocol-performance claims. |
| Related RNA-centric chromatin capture variants | How do adapted workflows connect chromatin-associated RNAs, genomic loci, and RNA-RNA contacts? | Chromatin fractionation, antisense capture, proximity conversion, or ligation are combined according to the workflow. | Protocols may optimize reduced input, simultaneous molecule classes, or population-scale contact recovery. | RNA-by-locus matrices, locus enrichment profiles, or co-captured RNA pairs. | Flexible designs can connect RNA-chromatin occupancy with RNA-RNA interaction layers. | Normalization is harder across molecule classes; variant-specific biases limit direct comparison. | Input controls, no-probe background, abundance and accessibility covariates, replicates, perturbation-based validation. | Ding et al. 2025 is a local protocol anchor; variant-specific references remain incomplete. |
The second major family moves from single-RNA capture toward transcriptome-scale contact discovery. Some methods ask which RNAs associate with genomic DNA across many loci. Other methods ask which RNA segments contact other RNA segments across the transcriptome. The methods differ in chemistry, but they share a common ambition: convert many molecular contacts into sequencing pairs without requiring one candidate interaction at the start.
GRID-seq is commonly discussed as a global RNA-DNA proximity strategy. Conceptually, it produces links between RNA identities and genomic intervals. The output can be represented as a matrix in which rows are RNAs or RNA intervals and columns are DNA loci. A high matrix entry means that an RNA was repeatedly recovered near a genomic region under the assay conditions. Such data can reveal RNAs near their own genes, candidate trans associations between RNAs and distant loci, chromatin-associated noncoding RNAs, and RNA contact patterns that correlate with active or repressed chromatin. The 2025 Ding et al. protocol in the local references is relevant because it addresses simultaneous capture of chromatin-associated RNA and global RNA-RNA interactions with reduced input requirements, but the chapter still needs a verified GRID-seq landmark reference for final method history.
Population-scale RNA-DNA contact maps are powerful because they do not require the investigator to select one lncRNA. They are also hard to normalize. RNA abundance varies over orders of magnitude. Chromatin accessibility differs between promoters, enhancers, repeats, heterochromatin, and open gene bodies. Nascent transcripts are naturally close to their encoding DNA. Genomic copy number, mappability, repetitive sequence, and transcription rate all influence recovery. A highly expressed nuclear RNA can appear near many loci because even a small fraction of its molecules are recovered often. A low-abundance regulatory RNA can be missed even if its contact is biologically important.
SPLASH, PARIS, COMRADES, and LIGR-seq are usually grouped with RNA-RNA interaction mapping. The shared readout is chimeric RNA information: two RNA regions that were linked during the assay and then sequenced as a joint molecule or paired signal. Many such methods use psoralen or related chemistry to enrich duplex-like contacts. Psoralen intercalates preferentially into double-stranded nucleic-acid regions and can form covalent crosslinks after ultraviolet exposure. A psoralen-enriched assay therefore has a useful bias toward paired regions. That bias is informative when the biological question is RNA duplex structure, but it also means the assay sees a conditioned subset of interactions.
PARIS, often expanded as psoralen analysis of RNA interactions and structures, is conceptually useful for intramolecular and intermolecular duplex discovery. SPLASH, commonly described through psoralen crosslinking, ligation, and selection logic, similarly links RNA segments that are close in a duplex-biased context. LIGR-seq emphasizes ligation of interacting RNA followed by sequencing. COMRADES has been particularly associated with matched RNA crosslinking and deep sequencing logic and is often invoked for viral RNA structure in infected cells. These expansions and histories need verified citation curation in this repository before they are treated as final reference-backed statements; here they are used only to orient the reader to the assay family.
A concrete RNA-RNA example is a positive-strand viral RNA genome. The genome contains coding regions, untranslated regions, replication signals, packaging signals, and structured elements. Long-range contacts can bring the 5′ and 3′ regions into proximity, connect regulatory elements, or organize replication complexes. An RNA-RNA proximity assay can recover chimeric reads between distant genome intervals. If the contact appears reproducibly, overlaps a predicted duplex, is supported by structure probing, changes during infection, and affects replication when mutated, the contact becomes a strong mechanistic candidate. If the contact is a single ambiguous read in a high-abundance viral RNA, the evidence is weak.
Cellular transcriptomes create more contact classes. Intramolecular contacts occur within one transcript, such as a long-range stem in an untranslated region. Intermolecular contacts occur between different RNAs, such as a small RNA and a target mRNA, an antisense RNA and a sense transcript, or two RNAs sharing an RNP. Contacts can also occur between isoforms or between nascent RNA regions from neighboring transcription units. A method that recovers all of these contacts in one library creates a rich dataset but also a severe interpretation problem: each contact class has different background rates, mapping ambiguity, and biological expectations.
The relationship between GRID-seq-like RNA-DNA methods and SPLASH/PARIS/COMRADES/LIGR-seq-like RNA-RNA methods is not simply chromatin versus transcriptome. Some protocols can be adapted to recover multiple molecule classes, and some workflows combine chromatin-associated RNA capture with global RNA-RNA interaction mapping. The conceptual division remains useful because RNA-DNA pairs and RNA-RNA chimeras answer different questions. RNA-DNA proximity asks where RNAs sit relative to the genome. RNA-RNA proximity asks which RNA segments sit near each other. A regulatory model may need both: a chromatin-associated lncRNA may contact a genomic locus and also fold through intramolecular domains or pair with another RNA.
The current consensus is that these global methods are best viewed as discovery assays. They are excellent for generating candidate contact maps, comparing states, finding recurrent structures, identifying RNAs enriched near chromatin, and designing focused validation experiments. They are weaker as standalone proof of direct molecular mechanism. This distinction is especially important for large networks. Network visualizations can make an RNA appear to regulate many partners, but the edges may combine direct duplexes, shared compartments, abundance effects, and technical background.

Figure 134.2. RNA-Centric Chromatin Capture Workflows. Probe design, accessibility, and hybridization stringency shape the recovered chromatin map.
Table 134.2. Global RNA-DNA and RNA-RNA Contact-Mapping Assay Families. Method names should be interpreted through assay logic rather than treated as interchangeable acronyms.
| Method family | Molecule pair recovered | Chemistry or capture principle | Typical data object | Best use | Major caution | Citation status in local references |
|---|---|---|---|---|---|---|
| GRID-seq-like RNA-DNA mapping | RNA identities or RNA intervals linked to genomic DNA intervals. | Chromatin-associated RNA-DNA proximity is converted into sequenceable pairs; exact chemistry is protocol-specific. | RNA-by-genomic-locus matrix or RNA-DNA pair list. | Discovering chromatin-associated RNAs, cis retention, candidate trans loci, and state-dependent RNA-DNA neighborhoods. | RNA abundance, nascent transcription, chromatin accessibility, mappability, and 3D compartment background can drive signal. | Ding et al. 2025 supports simultaneous capture logic; verified GRID-seq landmark reference is still absent. |
| SPLASH-like psoralen-selected RNA-RNA mapping | Intramolecular or intermolecular RNA segments. | Psoralen enriches duplex-like RNA geometry before ligation, selection, and sequencing. | Chimeric reads grouped into RNA-RNA contact clusters. | Transcriptome-scale discovery of long-range and intermolecular duplex-biased contacts. | Psoralen-compatible contacts are a subset; protein-shielded, transient, or inaccessible proximities can be missed. | Verified SPLASH landmark reference is absent locally. |
| PARIS-like duplex mapping | RNA duplex arms within or between transcripts. | Psoralen crosslinking and duplex enrichment recover paired RNA fragments for sequencing. | Duplex groups or interval-level contact maps. | Mapping candidate structural contacts in abundant cellular or viral RNAs. | Ligation coordinates are not exact physical contact points; complementarity alone is not proof. | Verified PARIS landmark reference is absent locally. |
| COMRADES-like matched RNA mapping | Matched RNA-RNA contacts, often in target-enriched or infection-focused contexts. | Crosslinking, ligation, enrichment, and deep sequencing connect RNA arms sampled by the workflow. | Matched chimeric RNA pairs or contact maps for selected RNA populations. | Prioritizing viral RNA structures or RNA contacts in infected cells. | Target enrichment and high RNA abundance can dominate unless controls separate biology from recovery bias. | Verified COMRADES landmark reference is absent locally. |
| LIGR-seq-like ligation of interacting RNA | Proximal RNA regions joined into ligation products. | Interacting RNA fragments are ligated and sequenced across the junction. | Split-read junctions and clustered RNA-RNA contacts. | Finding candidate intra- and intermolecular RNA contacts for focused validation. | Ligation bias, reverse-transcriptase template switching, PCR recombination, and abundance effects can create false chimeras. | Verified LIGR-seq landmark reference is absent locally. |
RNA-RNA proximity ligation turns a physical neighborhood into a sequence junction. The simplest model is that two RNA segments are close, fragmentation exposes ends, a ligase joins those ends, and sequencing reads across the ligation junction. The real workflow is more complicated. A protocol may crosslink RNAs before fragmentation, enrich crosslinked duplexes, reverse crosslinks, repair ends, select biotinylated molecules, remove abundant background, reverse transcribe, amplify, and then computationally split-align reads. Each stage changes which contacts are visible.
The first interpretive step is to define the arms of a chimeric read. An arm is the sequence on one side of the junction that maps to a transcript or genomic interval. Arm mapping should record uniqueness, strand, orientation, edit distance, overlap with repeats, overlap with splice junctions, and compatibility with known transcript models. A read whose two arms map uniquely to plausible RNA intervals is more useful than a read whose arms map equally well to many paralogous transcripts or repeat copies. For intramolecular contacts, the distance between arms and their orientation matter. For intermolecular contacts, RNA identity, expression level, compartment, and shared sequence families matter.
The second interpretive step is to classify the contact. A cis intramolecular contact joins two intervals of the same RNA molecule. This class is often relevant to secondary or tertiary structure. A trans intermolecular contact joins two different RNA molecules. This class may represent antisense regulation, small-RNA targeting, viral-host interaction, or co-localization. A cis genomic but transcript-ambiguous contact can occur when arms map to the same gene but different isoforms or pre-mRNA regions. A repeat-mediated contact can be biologically real or impossible to place uniquely. Contact class should be explicit because the background model changes with the class.
The third interpretive step is clustering. A single chimeric read is a fragile observation. A cluster of reads connecting nearby intervals in independent libraries is stronger. Clustering allows for imprecision created by fragmentation and ligation while reducing overinterpretation of one read. However, clustering can also hide heterogeneity. Two nearby but distinct helices can merge into one broad interval-level contact. Abundant RNAs can create dense background clusters. An analysis should report the clustering resolution and sensitivity to parameter changes.
The fourth interpretive step is to ask whether a contact is chemically and structurally plausible. If a method is duplex-biased, the two arms might be expected to show complementarity. A predicted low-energy duplex can strengthen a base-pairing model, especially if the same intervals show reduced reactivity in SHAPE, DMS, or icSHAPE experiments. But complementarity is not proof. Short complementary sequences occur by chance in transcriptomes. Conversely, imperfect complementarity does not disprove proximity because protein-mediated contacts, kissing-loop interactions, pseudoknotted regions, and tertiary arrangements can produce proximity without a simple continuous helix.
The fifth interpretive step is to connect reads to biological context. An RNA-RNA contact between a small regulatory RNA and an mRNA is more convincing if the contact overlaps known seed or target regions and changes when the small RNA is depleted. A contact within a viral genome is more convincing if it is conserved across related viruses or affects replication when mutated. A contact between two abundant ribosomal or mitochondrial RNAs may reflect compartmental abundance unless the protocol and controls address that background. A contact in a nuclear sample should be interpreted differently from the same contact in total cellular RNA.
False chimeras require explicit handling. Reverse transcriptase can switch templates, especially at structured, damaged, or modified RNA. PCR can recombine related molecules. Random ligation can occur after lysis if abundant RNAs encounter each other in solution. Trans-splicing, read-through transcription, circular RNA junctions, and unannotated isoforms can create genuine RNA junctions that are not proximity-ligation products. Repeats and pseudogenes can cause an arm to be assigned to the wrong source. A well-designed pipeline removes or flags these classes rather than treating every split alignment as a contact.
The phrase “nucleotide-resolution contact” should be used carefully. A ligation junction has nucleotide coordinates, but those coordinates are not necessarily the exact nucleotides that touched in the cell. The junction reports RNA ends generated by fragmentation and end processing. The physical contact may have occurred nearby, or the ligase may have bridged a short distance within a constrained complex. Contact maps can be displayed at nucleotide resolution, but molecular certainty often remains interval-level.
A useful analysis record for each candidate contact includes: contact ID, RNA identities, arm coordinates, strand, class, read count, number of distinct molecules if unique molecular identifiers are available, replicate support, arm mappability, overlap with repeats, RNA abundance, predicted duplex score when relevant, structure-probing consistency, protein-binding context, control-library enrichment, and perturbation response. Such a record supports transparent evidence grading and makes downstream comparison possible.
Box 134.2. Reading a Chimeric RNA Junction
Treat a chimeric RNA read as a candidate observation, not as a finished conclusion. First ask where each arm maps and whether the placement is unique. Short arms that overlap repeats, paralogs, pseudogenes, unannotated isoforms, or splice junctions should be flagged. Next classify the junction: cis intramolecular contact, trans intermolecular contact, transcript-ambiguous junction, repeat-associated signal, or known RNA-processing product. Then ask whether the workflow could have made the junction without a cellular contact. Reverse transcriptase template switching, PCR recombination, trans-splicing, circular RNA junctions, read-through transcription, and post-lysis random ligation are common alternatives. Finally, place the junction in biological context. Replicate contact clusters, compatible expression, plausible compartment, structure-probing agreement, and perturbation response make the contact more credible. A single high-scoring split read with poor mappability remains weak evidence.

Figure 134.3. Chimeric Read Interpretation Workflow. A reproducible contact cluster is usually more meaningful than a single junction coordinate.
Bias enters RNA-RNA and RNA-chromatin mapping before the first sequencing read exists. A cell is not a uniformly accessible solution. RNAs differ in abundance, length, structure, modification state, protein coverage, compartment, processing stage, and decay status. Chromatin differs in accessibility, protein density, nuclear position, transcriptional activity, repeat content, and fragmentation behavior. A method samples this heterogeneous molecular landscape through a specific chemistry. Therefore, contact recovery is always conditional on the protocol.
Crosslinking bias is one of the most important constraints. Formaldehyde is useful for preserving protein-rich chromatin complexes because it can capture protein-nucleic-acid and protein-protein neighborhoods. This makes it valuable for RNA-chromatin capture, where many associations are expected to be protein mediated. The same property limits mechanistic resolution. Formaldehyde recovery of an RNA with a locus does not tell whether the RNA touches DNA, touches a protein that touches DNA, shares a condensate, or remains trapped near transcription machinery. Stronger fixation can improve recovery but may increase background and reduce accessibility.
Psoralen bias is different. Psoralen enriches duplex-like nucleic-acid geometry. In RNA-RNA mapping, this is a feature when the goal is to find paired regions. It is also a limitation because psoralen does not capture all RNA proximity. Protein-shielded helices, transient interactions, unusual geometries, short-lived contacts, or regions inaccessible to the compound may be underrepresented. A psoralen-based negative result therefore cannot be interpreted as absence of all interaction. It is absence or poor recovery under psoralen-compatible conditions.
Ligation bias begins with RNA ends. Ligation requires compatible end chemistry and physical access. Fragmentation can create 2′, 3′ cyclic phosphates, hydroxyls, phosphates, damaged bases, or ends blocked by proteins. Enzymatic repair can be incomplete. RNA ligases prefer certain end structures and sequence contexts. Structured regions may bring ends close or hide them. Abundant RNAs can create many ligation opportunities simply by mass action. As a result, raw chimeric read count is not a direct molecular ruler for contact frequency.
Capture bias is central for ChIRP-seq, CHART-seq, RAP, and related methods. Probe sequences differ in melting temperature, GC content, secondary structure, and off-target complementarity. A probe may bind an accessible loop but not a paired region of the same RNA. Repetitive RNA sequence can capture many genomic or transcriptomic regions nonspecifically. Highly abundant RNAs can contaminate capture libraries. Streptavidin beads and fixed chromatin can contribute sticky background. Independent probe pools and target-depletion controls are essential because no probe design algorithm can fully predict fixed-RNP accessibility and off-target behavior.
Mapping bias is especially severe in RNA biology. Many RNA reads derive from paralogous gene families, pseudogenes, transposons, ribosomal RNA fragments, mitochondrial transcripts, small nuclear RNAs, or low-complexity sequence. Spliced transcripts create exon-exon junctions that do not align as contiguous genomic sequence. Nascent RNAs include introns and unprocessed regions. RNA editing, RNA modification-induced misincorporation, genetic polymorphism, and sequencing error can shift alignments. When a chimeric read has two short arms, each arm may lack enough information for unique placement.
False contacts can arise through real biology that is not the intended biology. Two RNAs in the same nuclear body can be ligated without directly regulating one another. Two chromatin loci that are close in 3D genome space can produce apparent RNA-DNA associations because an RNA near one locus is crosslinked in a shared compartment. A highly expressed transcript can produce trans contacts because it fills a compartment. A stress condition can change fixation or extraction behavior, making a technical change appear like a biological network rewiring.
False contacts can also arise after lysis. If compartment boundaries are disrupted before contacts are fixed or if fixation is weak, abundant RNA fragments can meet in solution and ligate. If chromatin is overfragmented, material that should have remained spatially separated may redistribute. If library construction selects molecules by size or end chemistry, a subset of artifacts can be preferentially amplified. Spike-ins and mixing experiments can be useful here. For example, mixing cells or nuclei from distinguishable species before or after fixation can estimate cross-sample chimeras and post-lysis ligation background.
The practical response is not to demand a bias-free assay. Such an assay does not exist. The practical response is to match controls to the expected failure modes. RNA-centric chromatin capture needs independent probes, target-depletion controls, input chromatin, off-target assessment, and often imaging or locus-specific validation. RNA-RNA ligation needs no-ligase or altered-ligation controls where appropriate, crosslinker controls, replicate clustering, template-switching filters, abundance normalization, and targeted validation. RNA-DNA population maps need mappability masks, RNA abundance covariates, chromatin accessibility covariates, and conservative handling of trans contacts.
Do not overgeneralize from read depth. Sequencing depth improves detection of molecules that entered the library. It does not correct a biased capture, a missing crosslinking class, a faulty ligation condition, or an ambiguous alignment. More depth can make systematic artifacts more statistically significant. The quality of a contact map depends on chemistry, controls, replicates, normalization, and validation, not only on the number of reads.

Figure 134.4. Bias Sources Across the Contact-Mapping Pipeline. Bias is distributed across the full workflow, so controls must also be distributed.
Table 134.3. Interpretation Hazards and Control Strategies. Artifact control is a set of matched checks, not a single generic negative control.
| Hazard | How it appears in data | Affected assay families | Practical controls | Residual uncertainty |
|---|---|---|---|---|
| Off-target probe hybridization | Peaks or recovered RNAs track sequence-similar probes or appear in only one probe pool. | ChIRP-seq, CHART-seq, RAP, related capture variants. | Independent probe pools, scrambled or sense probes, off-target screening, target-RNA depletion. | Repetitive or highly accessible loci can mimic target-dependent recovery. |
| RNA abundance | Highly expressed RNAs dominate contact counts or create broad trans edges. | Global RNA-RNA mapping, GRID-seq-like RNA-DNA mapping, RNA-centric capture. | RNA-seq covariates, expression-matched backgrounds, spike-ins, replicate-normalized contact calling. | True biology of abundant RNAs can be hard to separate from mass action. |
| Chromatin accessibility | Open promoters, enhancers, repeats, or fragile regions are overrepresented. | RNA-chromatin capture and RNA-DNA proximity mapping. | Input chromatin, ATAC-seq or DNase covariates, mappability masks, matched negative loci. | Accessibility can be both technical bias and relevant chromatin state. |
| Formaldehyde-mediated indirect proximity | RNA appears at loci through protein bridges, compartments, or retained nascent complexes. | ChIRP-seq, CHART-seq, RAP, GRID-seq-like assays. | Fixation titration, RNase-sensitive tests, protein perturbation, imaging, orthogonal chemistry. | Formaldehyde recovery alone cannot distinguish direct RNA-DNA contact from bridged proximity. |
| Psoralen selectivity | Duplex-like contacts are enriched while transient or protein-shielded contacts are depleted. | SPLASH-like, PARIS-like, COMRADES-like, other psoralen workflows. | No-psoralen or UV controls, structure probing, orthogonal ligation chemistry, mutational tests. | A negative psoralen result does not prove absence of cellular proximity. |
| Ligation bias | Junction counts follow end chemistry, sequence, structure, or ligase preference. | LIGR-seq-like assays and other RNA-RNA proximity-ligation workflows. | No-ligase or altered-ligase controls, UMIs, calibrated spike-ins, end-repair monitoring, replicate clustering. | Raw chimeric read count remains an imperfect contact-frequency measure. |
| Reverse-transcriptase template switching | Split reads appear at structured, damaged, modified, or abundant RNA regions without ligation support. | Chimeric-read RNA-RNA mapping workflows. | RT-condition controls, artifact filters, independent validation, replicate and junction-support thresholds. | Rare true contacts and rare template-switch products can look similar. |
| PCR recombination | Apparent chimeras increase with amplification or low-input libraries. | Amplified RNA-RNA and RNA-DNA proximity libraries. | Limit PCR cycles, use UMIs, include no-template controls, require biological-replicate support. | Low-level recombination can remain after deduplication. |
| Repetitive mapping | One or both arms map to repeats, paralogs, pseudogenes, or low-complexity sequence. | All sequencing-based contact maps; strongest for trans and repeat-derived signals. | Mappability filters, repeat-aware annotation, longer arm requirements, multi-map reporting. | Some real repeat contacts cannot be assigned to a single genomic copy. |
| Unannotated isoforms | Read-through, splice, circular, or pre-mRNA junctions masquerade as proximity chimeras. | Chimeric RNA-RNA mapping and nascent RNA-rich libraries. | Updated transcript annotations, long-read support, junction classification, isoform-aware alignment. | Incomplete annotations leave a residue of uncertain chimeras. |
| Random ligation after lysis | Abundant RNAs from disrupted compartments join after extraction. | RNA-RNA proximity-ligation workflows with weak fixation or delayed denaturation. | In situ fixation, fast denaturing lysis, no-ligase controls, species-mixing before and after fixation. | Sparse post-lysis collisions may be difficult to eliminate completely. |
| Compartment background | RNAs or loci sharing nuclear bodies or compartments form edges without direct regulation. | Global RNA-RNA, RNA-DNA, and RNA-chromatin contact maps. | Imaging, fractionation context, Hi-C or spatial covariates, perturbation of candidate bridge factors. | Compartment membership may be biologically real but noncausal for a specific edge. |
An RNA contact map becomes more informative when integrated with orthogonal measurements whose biases differ enough that agreement is meaningful. For RNA-RNA contacts, relevant layers include structure probing, comparative covariation, mutational analysis, protein-binding maps, expression, and localization. For RNA-chromatin contacts, relevant layers include RNA abundance, nascent transcription, chromatin accessibility, histone marks, protein occupancy, 3D genome conformation, imaging, and perturbation. Integration prioritizes functional hypotheses; it does not establish a chromatin mechanism by accumulation of correlated datasets.
Integration with RNA structure probing is the most direct route from proximity to duplex models. SHAPE-MaP, DMS-MaPseq, icSHAPE, and related assays measure chemical reactivity patterns that reflect local nucleotide flexibility, accessibility, or base-pairing environment. If an RNA-RNA contact suggests that two intervals form a duplex, reduced reactivity in the paired regions can support the model. If the same intervals show high reactivity under conditions where the contact should form, the duplex model weakens. Disagreement can be biological, because a contact may form only in a subpopulation, or technical, because the probing reagent and the proximity method sample different states.
Comparative sequence analysis adds another layer. Conserved complementary changes across related RNAs can support a functional base-pairing model. Covariation is especially useful when contact positions are conserved despite sequence divergence. However, absence of covariation does not necessarily disprove a recent, species-specific, or poorly aligned interaction. Repeats and rapidly evolving lncRNAs may lack the comparative depth needed for strong covariance tests. Covariation should be treated as powerful when present and interpretable, not as universally available.
Expression data help distinguish contact frequency from contact specificity. Highly expressed RNAs generate more molecules available for capture or ligation. A contact that scales with RNA abundance may reflect mass action rather than regulated interaction. Nascent RNA measurements can identify contacts dominated by transcription sites. Isoform-resolved RNA-seq can show whether the contacted region exists in the expressed isoform. Single-cell or spatial RNA data can reveal whether two RNAs are co-expressed in the same cells or compartments. Without expression context, a contact between two RNAs may be mathematically possible but biologically irrelevant in the cell population of interest.
Chromatin data are essential for RNA-DNA proximity interpretation. ATAC-seq or DNase sensitivity can identify accessible loci that are prone to recovery. Chromatin immunoprecipitation can show whether loci contain proteins that might bridge RNA. Hi-C or related genome-conformation data can show whether two loci are already close in nuclear space. Nascent transcription assays can distinguish active transcription from steady-state RNA accumulation. DNA methylation, histone modification, and transcription-factor occupancy data can suggest whether an RNA-associated locus is active, poised, repressed, or repetitive. These layers do not prove RNA function, but they prevent naive interpretation of all RNA-DNA contacts as direct targeting.
Protein-binding information is often the missing bridge. Many RNA-RNA and RNA-chromatin contacts are mediated by RNA-binding proteins, chromatin proteins, helicases, transcription factors, or architectural complexes. CLIP-family data can identify proteins bound near RNA contact intervals. RNA-centric proteomics can identify proteins associated with a captured RNA. Chromatin proteomics and ChIP-seq can identify proteins at contacted loci. If a contact disappears when a bridging protein is depleted, the model shifts from direct nucleic-acid pairing to protein-mediated proximity. That shift is not a failure; it is a clearer mechanism.
Functional hypotheses require a prespecified molecular readout and an analysis that separates contact change from abundance, localization, transcription, accessibility, or cell-state changes. The relevant readout depends on the proposed model, but this chapter owns the assay and inference design rather than the resulting RNA-chromatin mechanism. Chapter 96 evaluates mechanistic conclusions about recruitment, chromatin remodeling, histone modification, enhancer-promoter regulation, and transcriptional consequence.
Machine learning can help integrate these layers, but the labels must be treated cautiously. A classifier trained on psoralen-enriched contacts may learn psoralen accessibility, RNA abundance, and mappability rather than universal RNA-pairing rules. A model trained on RNA-DNA contacts may learn open chromatin and transcription rate rather than RNA targeting. Useful models should report feature dependence, use held-out biological contexts, test robustness to abundance and mappability, and be validated by perturbation. Prediction is valuable for prioritization; it is not a substitute for molecular evidence.

Figure 134.5. Integrating Contact Maps with Orthogonal Evidence. Orthogonal agreement can move a contact from discovery signal toward a mechanistic model.
Table 134.4. Evidence Standards for Contact, Direct Binding, and Causality Claims. Different conclusions require different evidence thresholds.
| Claim type | Minimum evidence | Stronger evidence | Common overclaim | Validation examples |
|---|---|---|---|---|
| Proximity detected | Reproducible contact, peak, or pair enriched over input or negative controls with mappable arms or loci. | Independent probe set or chemistry, contact clustering, replicate robustness, control sensitivity. | Proximity is direct binding or function. | Independent capture, targeted ligation, RNA FISH-DNA FISH, replicate contact calling. |
| Direct RNA-RNA duplex | Proximity plus sequence compatibility or duplex-biased chemistry. | Structure-probing support, covariation when available, disruption and compensatory rescue. | A chimeric read proves base pairing. | SHAPE/DMS/icSHAPE, compensatory mutations, antisense disruption, biochemical reconstitution. |
| RNA-chromatin occupancy | Target-RNA-dependent genomic enrichment with input and off-target controls. | Independent probe pools, orthogonal capture, imaging, depletion and rescue. | Occupancy proves regulation or direct RNA-DNA binding. | ChIRP/CHART/RAP with independent probes, RNA FISH-DNA FISH, RNase-sensitive assays. |
| Direct RNA-DNA recognition | Occupancy plus evidence for RNA-DNA hybrid, triplex-like, or other direct nucleic-acid contact. | Strand-specific perturbation, RNase H or R-loop evidence when appropriate, sequence-motif mutation, purified binding assay. | Any RNA-chromatin peak is direct DNA recognition. | R-loop mapping, RNase H tests, RNA or DNA recognition-site mutation, in vitro RNA-DNA binding. |
| Protein-bridged contact | Proximity plus candidate RNA-binding or chromatin protein co-occupancy. | Bridge protein binds both sides and contact weakens after protein perturbation with rescue where feasible. | A protein bridge is the same as direct RNA-RNA or RNA-DNA pairing. | CLIP, ChIP-seq, RNA-centric proteomics, degron or knockdown of the bridge protein. |
| Regulatory consequence | RNA or contact perturbation changes transcription, chromatin state, processing, translation, decay, or localization. | Domain-specific perturbation preserves overall abundance and localization while changing the contact. | Correlation between a contact and expression proves regulation. | RNA depletion plus rescue, reporter or locus assays, chromatin accessibility and nascent transcription readouts. |
| Causal contact mechanism | Perturbing the specific contact changes the predicted molecular consequence without relying on global RNA loss. | Compensatory rescue, allele-specific or locus-specific logic, tethering tests, biochemical reconstitution. | A contact-network edge is a causal regulatory edge. | Helix disrupt-and-restore mutations, RNA-domain rescue, locus tethering, contact restoration tied to phenotype. |
The evidence ladder begins with detection. A detected contact is a candidate molecular relationship recovered by one assay. Detection becomes stronger when it is reproducible across biological replicates, robust to reasonable analysis choices, enriched over input or negative controls, and supported by independent probe pools or independent chemistry. Detection is the first rung, not the endpoint.
The next rung is orthogonal validation. For RNA-chromatin capture, validation can include RNA fluorescence in situ hybridization combined with DNA FISH, locus-specific capture, independent antisense probe sets, RNase-sensitive chromatin assays, R-loop mapping when an RNA-DNA hybrid is hypothesized, or imaging of nuclear compartments. For RNA-RNA contacts, validation can include targeted ligation assays, structure probing, antisense oligonucleotide disruption, mutational analysis, compensatory base-pair restoration, or biochemical reconstitution with purified molecules. Orthogonal validation should test the specific interpretation, not merely repeat the same bias.
Perturbation tests whether the contact is required or responsive. Depleting an RNA can show whether chromatin occupancy or RNA-RNA contacts depend on that RNA, but depletion has broad side effects. It can change transcription, processing, localization, RNA-binding protein availability, and cellular state. A stronger perturbation changes the putative contact domain while preserving overall RNA expression and localization. For a duplex, this might mean mutating one side of the predicted pair. For an RNA-chromatin contact, it might mean deleting a domain needed for localization while preserving transcription of the locus.
Rescue is the key step that separates direct contact models from indirect abundance models. In a duplex model, disrupting one side of a predicted helix should alter the contact and function; restoring complementarity with compensatory mutations should restore both if the helix is causal. In an RNA-chromatin model, reintroducing an RNA domain or tethering an RNA to a locus can test whether localization is sufficient or whether the native transcription context is required. Rescue is difficult for long RNAs because mutations can alter folding, protein binding, processing, and stability. Difficulty does not remove the need for careful causal logic.
Box 134.3. Designing a Rescue That Tests the Contact
A useful rescue experiment asks whether the proposed contact is the causal variable. For an RNA-RNA duplex, the clean logic is disrupt and restore: mutate one side of the predicted helix, measure loss of contact and function, then introduce compensatory mutations that restore pairing without restoring the original sequence. For an RNA-chromatin model, the analogous test may restore an RNA localization domain, tether the RNA to a locus, or reintroduce the RNA from a position that separates mature RNA action from transcription through the native locus. In either case, the rescue should measure molecular endpoints close to the model: contact recovery, RNA localization, chromatin occupancy, local structure, protein binding, transcription, processing, translation, or decay. Restoring total RNA abundance or cell growth alone is not enough, because broad depletion effects can mask which interaction mattered.
Allele-specific and locus-specific designs can separate local transcription, the mature RNA product, and a mapped contact. Promoter shutdown, RNA depletion, domain mutation, ectopic expression, locus tethering, and rescue perturb different variables and therefore support different inferences. This chapter defines those contrasts as validation designs; Chapter 96 owns the biological conclusions about cis or trans action and chromatin regulation.
Validation should measure an endpoint close to the mapped contact before using organismal or phenotypic outcomes. A perturbation should first be tested for contact recovery, local structure, RNA occupancy, abundance, localization, RNP binding, or another prespecified molecular output. A cell-growth or disease phenotype after RNA depletion is not enough to identify which mapped contact was responsible.

Figure 134.6. Evidence Ladder from Contact Detection to Supported Inference. Causal interpretation requires more than a contact edge.
The current consensus is that RNA-RNA and RNA-chromatin interaction mapping methods are discovery and evidence-layer methods rather than standalone mechanism tests. ChIRP-seq, CHART-seq, RAP, GRID-seq-like assays, SPLASH, PARIS, COMRADES, and LIGR-seq can prioritize molecular neighborhoods, candidate duplexes, lncRNA occupancy patterns, chromatin-associated RNAs, and state-dependent contact changes. They become strongest when interpreted with controls, replicates, orthogonal assays, and perturbation.
For RNA-RNA contacts, reproducible contact clusters are generally stronger than isolated chimeric reads. Direct duplex models require more than a junction: duplex-biased chemistry, sequence compatibility, structure-probing agreement, and mutational logic all add strength. For RNA-chromatin contacts, reproducible capture across independent probes or chemistries is important, but regulatory claims require tests that separate RNA abundance, transcription through a locus, local chromatin state, and direct action of the RNA molecule.
Negative evidence is also method conditioned. Failure to detect a contact in one protocol can reflect low abundance, poor mappability, protein shielding, crosslinker incompatibility, inaccessible probe regions, weak ligation, or cell-state specificity. Conversely, a highly significant contact can still be indirect or artifactual if it follows abundance, accessibility, repeat mapping, or compartment background. The consensus position is therefore neither skepticism that dismisses contact maps nor enthusiasm that treats every edge as a mechanism.
Open questions:
Controversies:
Common misconceptions:
Deprecated or weakened claims: