# Chapter 46. Evidence Framework for RNA Modification Claims, Terminology, and Artifact Control

## Scope Note

This chapter provides the evidence framework for interpreting RNA modification claims. It owns terminology, claim types, evidence tiers, stoichiometry, orthogonal-validation logic, assay constraints, writer/eraser/reader evidence, and artifact control. Shared modification-enzyme chemistry, substrate recognition, kinetics, specificity, evolution, and inhibitors belong to [Chapter 47](chapter1163.md); mark-specific pathways to Chapters [48](chapter1044.md) through [51](chapter1047.md); cross-mark biology to [Chapter 52](chapter1048.md); and measurement workflows to [Chapter 132](chapter1120.md).

## Executive Summary

The epitranscriptome is the set of chemical features added to, removed from, or retained on RNA molecules after or during RNA synthesis, together with the enzymes, binding proteins, detection methods, and biological contexts used to interpret those features. The concept is useful because RNA is not a four-letter polymer inside cells. Mature RNAs can contain methylated bases, methylated riboses, isomerized uridines, deaminated bases, complex tRNA hypermodifications, specialized caps, synthetic nucleosides, and chemically damaged residues. These modifications can affect folding, protein binding, decoding, RNA stability, innate immune recognition, localization, translation, processing, and decay. The concept is dangerous when it is treated as a single regulatory layer with uniform rules. A near-stoichiometric tRNA anticodon modification, a rare mRNA internal m6A peak, an inosine created by adenosine deamination in a duplex region, and oxidative RNA damage are chemically and evidentially different claims.

A modification claim should specify five things before biological interpretation begins: the chemical identity of the modified residue, the RNA species or transcript region, the site or resolution of the measurement, the modification fraction or stoichiometry when available, and the cellular context. Without those qualifiers, phrases such as "this transcript is methylated" or "m6A increases in cancer" can hide large differences in meaning. A transcript-level enrichment peak does not show that every molecule of the transcript carries the mark. A bulk nucleoside mass-spectrometry signal does not locate the residue. A direct RNA sequencing prediction is a model-dependent inference from electrical or base-calling features, not a chemical structure determination.

The strongest epitranscriptome evidence combines independent evidence classes. Regional enrichment, site-sensitive chemical or reverse-transcription signatures, bulk chemical identification, native-RNA signal models, and genetic or biochemical perturbation answer different questions and fail differently. Orthogonal validation means that independent principles converge on the same claim, not merely that one library or signal type is repeated. The protocols, benchmarking, calibration, and workflow comparisons for these technologies belong to [Chapter 132](chapter1120.md); this chapter asks what each output permits a reader to conclude.

The terms writer, eraser, and reader should be used as hypotheses unless supported by mechanism. A writer claim should show that the enzyme or complex modifies a defined substrate, ideally in purified or reconstituted conditions, and that perturbation changes site-specific modification in cells without confounding global RNA changes. An eraser claim should show direct removal or chemical conversion of a mark, not only altered total modification after stress. A reader claim should show preferential binding to the modified form, structural or biochemical recognition logic, and a downstream output that depends on the binding event. Reviews of mRNA modification proteins emphasize that these standards are clearest for some m6A machinery and less complete for many proposed marks and disease contexts.

Artifact control is not a technical afterthought in epitranscriptomics. It is part of the definition of the field. Antibody cross-reactivity, RNA fragmentation bias, incomplete chemical conversion, reverse-transcriptase misincorporation unrelated to the mark, sample degradation, rRNA or tRNA contamination, modified nucleotide carryover, batch-specific library effects, mapping ambiguity, model overfitting in direct RNA sequencing, and perturbation pleiotropy can all create false positives or false negatives. Current consensus therefore treats modification maps as evidence with resolution, sensitivity, specificity, and context limits rather than as direct inventories of biology.

## Concept Inventory

- **RNA modification:** a covalent chemical difference from the four canonical ribonucleotides adenosine, cytidine, guanosine, and uridine, or from the standard RNA end structures expected for a given RNA class. Examples include N6-methyladenosine, abbreviated m6A; 5-methylcytidine, abbreviated m5C; pseudouridine, often written psi in plain text; inosine, produced by deamination of adenosine; N4-acetylcytidine, abbreviated ac4C; ribose 2′-O-methylation; and complex tRNA modifications such as queuosine. The same chemical mark can have different meaning in different RNA classes. Pseudouridine in rRNA and tRNA is often ancient and highly occupied; pseudouridine in mRNA may be lower in occupancy and more context dependent.
- **Epitranscriptome:** the collection of RNA modifications and modification states across the RNAs of a cell, organism, virus, tissue, developmental stage, disease state, or experimental condition. The term is modeled partly by analogy to epigenome, but the analogy must not be pushed too far. RNA molecules are usually shorter lived than chromosomal DNA, many RNA modifications are installed during RNA processing or maturation, and many modified RNAs are products of specific RNA classes rather than long-term heritable chromatin states.
- **Stoichiometry:** the fraction of RNA molecules carrying a specified modification at a specified position under a specified condition. If 20 percent of molecules of a transcript carry m6A at one adenosine and 80 percent do not, the site has 20 percent modification stoichiometry in that sample. Stoichiometry is not the same as read count, peak height, or total nucleoside abundance. It requires a denominator: the amount of the same RNA molecules or same site that could have carried the mark.
- **Dynamic modification:** that modification fraction, site choice, or reader engagement changes across time, cell state, stress, development, infection, differentiation, treatment, or RNA life-cycle stage. Dynamic does not necessarily mean reversible removal from the same RNA molecule. A population-level change can arise because newly synthesized RNAs are modified differently, because modified RNAs decay faster or slower, because cell composition changes, or because an enzyme actively removes a mark.
- **Writers:** enzymes or enzyme complexes that install RNA modifications. Erasers are enzymes that remove or reverse modifications. Readers are proteins or RNA-protein complexes that bind a modified RNA differently from the unmodified RNA and produce or mediate a downstream consequence. The three terms are mechanistic shorthand. The appropriate evidence is biochemical activity for writers and erasers, selective binding for readers, and cellular causality for all three.
- **Orthogonal validation:** validation by methods with different failure modes. For example, an m6A site proposed by antibody enrichment becomes more credible if supported by a chemical or enzymatic site-resolution assay, mass-spectrometry evidence for the modified nucleoside in the relevant RNA class, loss of signal after writer disruption, and restoration by rescue. Repeating the same antibody enrichment with more sequencing depth is replication, not orthogonality.

## What to Know Before Reading This Chapter

The reader should know the basic structure of RNA. RNA is a polymer with a sugar-phosphate backbone, bases attached to ribose sugars, and polarity from the 5′ end to the 3′ end. A chemical modification can occur on the base, the ribose, the phosphate, or the end structure. The location matters because base modifications can change pairing and recognition, ribose modifications can change local geometry and nuclease sensitivity, and cap modifications can affect translation, innate immune sensing, and decay.

The reader should also distinguish an RNA species, a transcript isoform, and a nucleotide position. A total mRNA sample contains many molecules from many genes. A transcript such as a human mRNA can have alternative start sites, splice isoforms, and alternative polyadenylation sites. A modification call at a genomic coordinate may correspond to different transcript contexts depending on annotation. Claims about sites therefore require coordinate system, transcript model, strand, and isoform awareness, as introduced in [Chapter 18](chapter1017.md).

Finally, the reader should separate detection from function. Detection asks whether a modified residue is present. Quantification asks how much of it is present. Localization asks where it is on which RNA. Function asks whether the modification changes a molecular or cellular outcome. A single assay rarely answers all four questions. A careful epitranscriptome study makes the claim type explicit before drawing biological conclusions.

## 46.1. Epitranscriptome terminology and classification

Epitranscriptome terminology begins with a simple observation: cellular RNAs contain many covalent nucleoside variants beyond adenosine, cytidine, guanosine, and uridine. Reviews of common RNA modifications catalog dozens of recurrent marks across tRNA, rRNA, mRNA, snRNA, viral RNA, and other RNA classes. Some marks are abundant and ancient. Others are rare, condition-specific, or currently supported only in a subset of systems. The first classification question is therefore not "is this epitranscriptomic?" but "what exactly is modified, on which RNA, at what site, in what fraction of molecules, and with what evidence?"

Chemical classification groups modifications by the part of the nucleotide changed. Base methylations include m6A, m5C, and N1-methyladenosine. Ribose methylation usually refers to 2′-O-methylation, in which a methyl group is added to the ribose 2′ hydroxyl. Isomerization includes pseudouridine, in which uridine is rearranged so that the base is attached through a carbon-carbon bond rather than the usual nitrogen-carbon glycosidic linkage. Deamination includes adenosine-to-inosine editing and cytidine-to-uridine editing, treated in detail in [Chapter 50](chapter1046.md). Complex tRNA modifications can combine multiple atoms and biosynthetic steps; queuosine is an example whose biology is inseparable from tRNA identity and microbial or dietary metabolism. End modifications include the 5′ cap, cap-adjacent methylations, noncanonical caps, and tail-associated changes discussed in Chapters [26](chapter1025.md) and [29](chapter1028.md).

RNA-class classification is equally important. Transfer RNAs and ribosomal RNAs are heavily modified, often with high occupancy at conserved positions needed for decoding, folding, or ribosome function. Small nuclear RNAs and small nucleolar RNAs carry modifications linked to spliceosomal and guide-RNP maturation, as developed in [Chapter 27](chapter1026.md), [Chapter 42](chapter1039.md), and [Chapter 49](chapter1045.md). Messenger RNAs and many long noncoding RNAs can carry internal marks at lower and more variable occupancy. Viral RNAs and therapeutic RNAs introduce additional contexts: viral modifications can influence immune evasion and replication, whereas synthetic nucleoside substitutions in therapeutic RNA are intentionally engineered features rather than endogenous regulation. [Chapter 108](chapter1103.md) discusses immune discrimination of RNA, and [Chapter 153](chapter1137.md) discusses modified nucleosides in mRNA therapeutics.

Another classification axis is biological role. Some modifications tune RNA structure directly. Some alter codon-anticodon decoding or ribosome performance. Some recruit or repel proteins. Some protect RNA from nucleases or immune sensors. Some are biosynthetic maturation marks whose absence indicates failed RNA processing. Some are damage products produced by oxidation, alkylation, ultraviolet exposure, chemical handling, or sample storage. Damage products can be biologically meaningful when cells experience oxidative stress, but they should not be automatically placed in the same category as enzyme-programmed regulatory marks.

The word "epitranscriptome" can tempt readers to group all RNA modifications into one regulative layer parallel to chromatin epigenetics. That shorthand is useful for organizing methods and concepts, but it can obscure differences. A stable rRNA modification installed during ribosome assembly is not regulated in the same way as a transcript-specific m6A site on a stress-induced mRNA. A-to-I editing changes base-pairing interpretation and can recode proteins or alter double-stranded RNA recognition; it is often discussed alongside modifications because inosine is a noncanonical nucleoside, but the enzymology and evidence logic belong partly to editing biology. A synthetic N1-methylpseudouridine incorporated throughout an mRNA vaccine is a design variable, not an endogenous site-specific mark.

![Figure 46.1. Classification Axes for RNA Modifications](../assets/figures/chapter1043_figure1.png)

**Figure 46.1. Classification Axes for RNA Modifications.** RNA modification interpretation requires specifying what chemical change was measured, which RNA carried it, what biological role is proposed, and how strong the evidence is. Four independent axes structure the classification: chemical class (base methylation, ribose methylation, isomerization, deamination, cap or end modification, complex tRNA modification, synthetic substitution, or damage product), RNA class (mRNA, tRNA, rRNA, snRNA or snoRNA, lncRNA or circRNA, viral RNA, or therapeutic RNA), biological role (folding, decoding, protein binding, immune sensing, decay, processing, or damage marker), and evidence level (detected nucleoside, candidate region, candidate site, quantified site, reader-bound site, or functional site).

For this chapter, "RNA modification" is the broad chemical term and "epitranscriptome" is the systems-level term for modified RNA states and their machinery. A modification claim should avoid vague statements such as "the epitranscriptome controls translation." A better statement is: "In this cell type, depletion of a defined m6A writer complex decreases modification at mapped sites in a set of mRNAs, and the change is associated with altered reader binding and altered mRNA stability or translation." The second statement contains a chemical identity, method-aware site evidence, perturbation, effector hypothesis, and output.

The classification system also needs room for uncertainty. Many modification maps are candidate maps, not finished atlases. Reviews of detection methods emphasize that different methods disagree because they measure different molecular features, have different resolution, and fail in different ways. A responsible terminology should therefore distinguish "detected nucleoside", "candidate site", "high-confidence site", "quantified site", "differential site", "reader-bound site", and "functional site." These labels prevent an enrichment signal from being silently upgraded into a causal mechanism.

## 46.2. Interpreting stoichiometry, dynamics, and context-dependent modification claims

Stoichiometry is one of the most important and most often hidden quantities in epitranscriptomics. A modified site can be nearly fully occupied, partially occupied, or present in only a small subpopulation of RNA molecules. This matters because molecular interpretation changes with occupancy. If a tRNA anticodon-loop modification is nearly stoichiometric, the mature tRNA pool may behave as a chemically defined species. If an mRNA site is modified on 5 percent of molecules, then a bulk phenotype cannot be explained as if every copy of the transcript carried the mark unless the modified subpopulation has a special localization, translation state, or decay pathway.

A concrete example is an mRNA with a candidate internal m6A site. Antibody enrichment might show a peak over the region. That peak means fragments containing the region were enriched relative to input under the conditions of the assay. It does not directly say whether 5 percent or 80 percent of the transcript molecules carry the modification, whether the mark is on one adenosine or several nearby adenosines, whether one isoform is selectively modified, or whether a reader protein binds the modified molecules. A stoichiometric claim needs additional calibration or site-resolved quantitative evidence.

Population dynamics can be mistaken for molecular reversibility. Suppose a stress treatment increases the apparent modification level of an mRNA after one hour. That increase could result from faster installation on newly transcribed RNA, slower decay of modified RNA, faster decay of unmodified RNA, altered RNA isoform usage, altered cell-state composition, or active demethylation before the treatment followed by reduced demethylation after the treatment. A reversible enzyme may be involved, but the population-level measurement alone does not prove removal from individual RNA molecules. Reviews of dynamic m6A methylation make this distinction important when interpreting writers and erasers.

Context-dependent is a qualifier, not an explanation. A claim should name the cell type, organism, condition, RNA class, isoform, subcellular compartment, and time point that define its scope. The same enzyme, site, or mark may behave differently when substrate abundance, cofactors, localization, RNA structure, or competing RNA-binding proteins change. Cross-mark biological interpretation of those differences belongs to [Chapter 52](chapter1048.md); the evidential rule here is that support in one context does not license an unqualified general statement.

RNA life-cycle stage constrains causal claims. A mark detected only after export cannot directly explain an earlier splicing event unless an earlier, unmeasured pool or a separate enzyme function is demonstrated. Likewise, a condition-associated change can reflect altered synthesis, processing, decay, or cell composition rather than modification turnover. [Chapter 25](chapter1024.md) treats co-transcriptional processing, and [Chapter 52](chapter1048.md) synthesizes modification dynamics across biological states.

Every stoichiometric claim has a denominator problem. Bulk nucleoside abundance uses the nucleosides in the analyzed pool as denominator; regional enrichment is a relative signal rather than occupancy; and model-derived site probabilities depend on training, coverage, sequence context, and RNA integrity. Exact calibration and estimation workflows belong to [Chapter 132](chapter1120.md). Here the interpretation requirement is to state what population supplies the numerator and denominator and how uncertainty propagates into the biological claim.

Context dependence also means that negative results should be interpreted carefully. Failure to detect a mark in poly(A)-selected RNA does not exclude the same mark from rRNA, tRNA, nonpolyadenylated RNA, viral RNA, or low-abundance isoforms. Failure to detect a dynamic change in bulk tissue does not exclude changes in a rare cell type. Failure to observe a phenotype after writer knockdown does not exclude function if the perturbation was incomplete, compensated, toxic, or masked by redundant pathways. The evidence standard should match the biological question rather than assume one assay can close every case.

> **Box 46.1. Why Stoichiometry Changes Interpretation**
>
> - An mRNA site at 5 percent occupancy means only one molecule in twenty carries the mark; a bulk phenotype cannot be attributed to uniform modification of the transcript.
> - The same site at 50 percent occupancy may indicate a regulated partition between modified and unmodified pools, relevant to dynamic or condition-dependent biology.
> - At 95 percent occupancy the transcript behaves as a chemically defined species and population-average models apply more cleanly.
> - Low stoichiometry can still matter if the modified subpopulation is selectively localized, translated, or stabilized, but that must be demonstrated rather than assumed.

## 46.3. Evidence tiers and orthogonal validation logic

An evidence tier describes the strongest claim justified by the observations, not the prestige or scale of the technology. A discovery signal can support a candidate region or candidate site. A chemically selective or independently calibrated signal can support chemical identity, site localization, or a quantitative estimate within its validated scope. Enzyme activity, selective binding, perturbation, rescue, and site-specific phenotype evidence support progressively narrower mechanistic claims. A functional tier requires that the proposed consequence depend on the modification rather than only correlate with a signal or protein perturbation.

![Figure 46.2. Claim-Support Ceilings by Assay Class](../assets/figures/chapter1043_figure2.png)

**Figure 46.2. Claim-Support Ceilings by Assay Class.** Map each assay class from primary signal to supported claim types and unsupported upgrades. Antibody enrichment, conversion or reverse-transcription signatures, LC-MS/MS, direct RNA signal models, and genetic perturbation should appear only at the level needed to show their ceilings for chemical identity, site localization, stoichiometry, dynamics, and function. Refer workflow and benchmark details to [Chapter 132](chapter1120.md).

The tiers are multidimensional rather than a single staircase. Chemical identity, RNA identity, site localization, stoichiometry, dynamics, enzyme dependence, reader recognition, and function are separate claim axes. Strong support on one axis does not silently fill the others. Bulk chemical identification can be strong evidence that a nucleoside exists while saying little about its transcript location. A precise candidate site may still lack a validated occupancy estimate. A reproducible perturbation phenotype may remain indirect if RNA abundance or cell state changes first.

Orthogonal validation asks whether evidence with independent physical principles and failure modes converges on the same claim. Repeating the same antibody, reverse transcriptase, mapping model, or trained caller increases replication but does not remove the shared bias. Orthogonality is therefore claim-specific: evidence for chemical identity should not be presented as validation of function, and a functional rescue should not be presented as direct proof of chemical structure.

A useful validation matrix starts with the claim and works backward. A site claim needs evidence that distinguishes the candidate residue from nearby positions and confounding sequence features. A stoichiometry claim needs a defensible denominator and calibrated uncertainty. A dynamic claim needs time, composition, synthesis, and turnover alternatives considered. A writer or eraser claim needs catalytic dependence on a defined substrate. A reader claim needs modification-selective binding and a dependent consequence. A functional claim needs separation of the mark-dependent effect from RNA-abundance, protein-scaffolding, and cell-state effects.

Controls become meaningful through the alternative explanation they exclude. Positive and negative standards test whether the signal responds to the mark. Writer loss or catalytic-dead rescue tests enzyme dependence but can remain pleiotropic. Matched input tests whether differential RNA abundance explains enrichment. Modified and unmodified substrates test binding selectivity. Replicates estimate reproducibility, whereas spike-ins and held-out standards test calibration or model transfer. The detailed construction and benchmarking of these controls belongs to [Chapter 132](chapter1120.md).

![Figure 46.3. Evidence Tiers and Orthogonal Validation Matrix](../assets/figures/chapter1043_figure3.png)

**Figure 46.3. Evidence Tiers and Orthogonal Validation Matrix.** Show chemical identity, RNA identity, site localization, stoichiometry, dynamics, enzyme dependence, reader recognition, and function as separate claim axes. For each axis, contrast replication of one signal with independent evidence that shares the same entity and context but not the decisive failure mode.

Validation must preserve the biological denominator and context. Total m6A in poly(A) RNA does not validate an increase at one mRNA site; an mRNA enrichment assay does not validate a tRNA modification; total-RNA co-immunoprecipitation does not validate modification-selective reader binding. Evidence is orthogonal only when it addresses the same entity, condition, and claim axis with a genuinely different failure mode.

## 46.4. Assay-class limitations as constraints on modification claims

Assay classes appear here only because their blind spots set ceilings on claim wording. The operational chemistry, instruments, library workflows, calibration, benchmarks, and reporting requirements are taught in [Chapter 132](chapter1120.md). The interpretive task is to translate each primary signal into what it can and cannot establish.

Antibody-based methods are limited by specificity and resolution. An antibody raised against a modified nucleoside may cross-react with related chemical structures, bind sequence or structure contexts unevenly, or prefer exposed modifications over buried ones. Fragmentation determines resolution; large fragments create broad peaks that can contain many candidate residues. Immunoprecipitation conditions can favor abundant RNAs, structured fragments, or RNAs bound by proteins. Batch-to-batch antibody differences can matter. For these reasons, antibody enrichment is best treated as regional evidence unless paired with site-resolution information and appropriate controls.

Sequencing-based methods inherit the biases of RNA extraction, selection, fragmentation, adapter ligation, reverse transcription, amplification, alignment, and transcript annotation. Poly(A) selection excludes most rRNA, tRNA, many noncoding RNAs, and some degraded or nonpolyadenylated transcripts. Ribodepletion can leave residual rRNA fragments that dominate modification signal. Small and heavily modified RNAs such as tRNAs can be difficult to reverse transcribe, so absence from a library may reflect technical dropout rather than biology. Mapping is especially difficult for repeats, paralogs, pseudogenes, short reads, and transcript isoforms.

Chemical conversion methods add their own limitations. A reagent may not reach a residue buried in RNA structure or protected by protein. Reaction conditions can degrade RNA or create off-target modifications. Incomplete conversion can mimic low stoichiometry. Overconversion can create false positives. A chemical signature validated in synthetic oligonucleotides may behave differently in long cellular RNAs. Therefore conversion efficiency, RNA integrity, sequence context, and negative controls are part of the result, not optional methods details.

Mass spectrometry has a different strength and a different blind spot. LC-MS/MS can provide high-confidence chemical identification and quantitative abundance for modified nucleosides when standards, calibration, digestion, and chromatography are well controlled. It can distinguish chemical isomers in some workflows better than sequencing alone. However, bulk nucleoside LC-MS/MS usually cannot say which transcript or site carried the mark. Contamination is a major concern: a small amount of tRNA or rRNA, which can be highly modified, can distort an mRNA modification estimate. RNA purification, digestion completeness, isotope-labeled standards, and blank controls are therefore central to interpretation.

Direct RNA sequencing avoids reverse transcription for the primary readout, but it is not free of inference. Nanopore signal reflects several neighboring nucleotides in the pore, not a single isolated base. A modification may shift signal subtly and in a sequence-dependent manner. The same current anomaly can be caused by a different modification, RNA damage, secondary structure, motor behavior, pore chemistry, or base-calling error. Training models need ground truth, and ground truth is uneven across modifications. Reviews of nanopore modification calling emphasize that base-calling models are improving quickly but remain dependent on calibration, benchmarking, and transparent uncertainty.

Single-molecule claims from direct RNA sequencing also require coverage discipline. Low read depth can produce apparent molecule-level heterogeneity that is mostly sampling noise. RNA degradation can bias reads toward transcript ends or stable fragments. Modified RNA standards may not match endogenous sequence context. Enzyme knockdown controls may change RNA abundance and isoform usage, creating apparent signal changes unrelated to site chemistry. A direct RNA call is strongest when supported by matched unmodified or writer-depleted controls, synthetic or in vitro transcribed standards, independent validation, and model performance metrics.

A useful way to apply these limits is to ask what each assay class cannot see. Antibody enrichment cannot define exact chemical identity or stoichiometry alone. Bulk mass spectrometry cannot localize sites alone. Reverse-transcription signatures cannot always distinguish modification from structure or damage alone. Direct RNA sequencing cannot identify every modification without trained models and controls. Genetic perturbation cannot distinguish direct substrate loss from indirect cell-state changes alone. The governing rule is that claim language must not exceed the least-supported link in the inference chain.

**Table 46.1. Assay-Class Outputs and Unsupported Overclaims.** Each RNA modification detection method yields a primary signal that supports specific inferences; applying conclusions beyond those limits is a common source of false positives in epitranscriptomics.

| Method class | Primary signal | Strong inference | Common unsupported overclaim |
| --- | --- | --- | --- |
| **Antibody enrichment** | Enriched fragments | Candidate regions or sites with controls | Exact site and stoichiometry without validation |
| **LC-MS/MS nucleosides** | Modified nucleoside abundance | Chemical identity and bulk quantity | Transcript or site localization after complete digestion |
| **Chemical conversion** | Conversion-dependent sequence signature | Site candidates for compatible marks | Perfect stoichiometry without conversion calibration |
| **RT signature** | Stops, mismatches, or mutation patterns | Candidate sites with background model | Chemical identity without controls |
| **Direct RNA nanopore** | Native signal deviation or model call | Model-dependent candidate modifications | Unambiguous modification identity in all sequence contexts |
| **Writer perturbation** | Signal change after enzyme perturbation | Candidate enzyme dependence | Direct catalytic mechanism without rescue or activity evidence |

> **Box 46.2. Reviewer Questions for an RNA Modification Claim**
>
> - What chemical identity is claimed, and by which method?
> - What RNA species, isoform, and coordinate system are used?
> - What is the denominator for abundance or stoichiometry?
> - What controls exclude contamination and batch effects?
> - Does perturbation distinguish direct modification effects from cell-state changes?
> - Is function shown at the site, transcript, RNA class, or pathway level?

## 46.5. Writer, eraser, and reader evidence standards

![Figure 46.4. Writer, Eraser, and Reader Evidence Standards](../assets/figures/chapter1043_figure4.png)

**Figure 46.4. Writer, Eraser, and Reader Evidence Standards.** Writer, eraser, and reader are mechanistic assignments that should not be inferred from correlation alone. A writer claim requires purified catalytic activity, a catalytic mutant control, endogenous site-level changes after perturbation, and rescue by wild-type enzyme; an eraser claim requires direct chemical removal of the mark, product chemistry, and time-course logic that distinguishes demethylation from altered RNA turnover; a reader claim requires binding preference for the modified form, structural or biochemical recognition evidence, and a downstream output that depends on the modification at the specific site.

The writer, eraser, and reader vocabulary is useful because it connects chemical marks to mechanisms. A writer installs a modification, an eraser removes or reverses it, and a reader recognizes the modified state. The vocabulary emerged most strongly from m6A biology, where methyltransferase complexes, demethylases, and YTH-domain proteins provide well-studied examples. Reviews of mRNA modification proteins summarize the writer-reader-eraser framework while emphasizing that the strength of evidence varies across marks and systems.

A writer claim requires more than correlation between enzyme abundance and modification abundance. The strongest evidence has several layers. First, the protein or complex has plausible catalytic chemistry and conserved active-site residues. Second, purified or reconstituted enzyme modifies a defined RNA substrate in vitro, or a tightly controlled cellular system shows direct activity. Third, catalytic mutation reduces activity without simply destroying protein expression or complex assembly. Fourth, perturbing the writer in cells changes site-specific modification on endogenous RNA. Fifth, rescue by wild-type but not catalytic-dead enzyme restores the modification and any downstream phenotype. Sixth, substrate specificity is explained by RNA sequence, structure, cofactors, localization, or partner proteins.

An eraser claim is narrower than a "modification decreases or increases when the protein changes" observation. A true eraser should remove or chemically reverse a mark from RNA. For m6A-related demethylases, evidence standards include direct enzymatic demethylation, product chemistry, substrate specificity, catalytic mutants, and cellular site-level changes. An apparent increase in modification after knocking down a protein could result from altered writer expression, RNA turnover, stress, cell-cycle changes, or RNA composition shifts. Dynamic modification is compatible with eraser activity, but dynamics alone do not prove erasure from the same RNA molecule.

A reader claim begins with selective binding. The protein should bind a modified RNA more strongly, differently, or in a different structural mode than the unmodified counterpart. Binding should be tested with matched modified and unmodified substrates and quantified when possible. Structural evidence can show the modified base in a recognition pocket, as with many YTH-domain m6A readers. Cellular evidence should then connect binding to an output such as RNA decay, translation, localization, splicing, phase behavior, or immune recognition. A protein that co-immunoprecipitates with a modified RNA is not automatically a reader; it may bind another protein in the RNP, bind the RNA sequence independently of the mark, or associate after crosslinking artifacts.

Perturbation experiments are powerful but pleiotropic. Knocking out a writer can alter many RNAs and many cell states. A reader depletion can change expression of other RNA-binding proteins. Overexpression can force nonphysiological binding. Catalytic-dead rescue, domain-specific mutations, acute depletion, time-resolved experiments, and separation-of-function alleles help distinguish direct mark-dependent mechanisms from secondary effects. Quantitative proteomics of writer, eraser, and reader proteins can support context interpretation when protein abundance is itself variable.

The evidence standard also depends on RNA class. A tRNA modification enzyme may have a small number of high-occupancy substrates whose absence creates decoding defects. An mRNA writer complex may act on thousands of sites with variable occupancy, so phenotypes can be distributed and indirect. A viral RNA modification enzyme may act only during infection, and a therapeutic RNA design may use synthetic incorporation rather than endogenous writing. The same writer-reader-eraser vocabulary should not flatten these differences.

Reader and writer evidence also interact. If a reader phenotype disappears when the modification site is mutated without changing the encoded protein or RNA abundance, the case becomes stronger. If a writer depletion reduces reader binding at the same site and rescue restores it, the writer-reader link becomes stronger. If mass spectrometry, site mapping, and binding assays all agree, the mechanism becomes more credible. Conversely, if a writer perturbation changes RNA abundance before modification is measured, then apparent loss of a site may reflect loss of the RNA substrate.

> **Box 46.3. Do Not Overgeneralize From m6A**
>
> - The writer, eraser, and reader framework is best developed for m6A pathways, where methyltransferase complexes, demethylases, and YTH-domain readers provide mechanistic detail.
> - Many proposed marks and RNA classes lack equivalent biochemical or structural evidence; applying the same vocabulary does not imply the same level of support.
> - See Chapters [48](chapter1044.md) through [51](chapter1047.md) for mark-specific mechanisms and [Chapter 52](chapter1048.md) for cross-mark biology and disease synthesis.
>
> All four figure IDs, three table IDs, and three box IDs continue under registry revision 17. No visual ID is retired in this synchronization.

## 46.6. Artifact control and false-positive failure modes

**Table 46.2. Artifact-Control Checklist.** Common failure modes in epitranscriptome experiments, the typical source of each artifact, and the controls or mitigations that reduce its impact.

| Failure mode | Typical source | Control or mitigation |
| --- | --- | --- |
| **Antibody cross-reactivity** | Related modifications, sequence context, structure | Competition assays, orthogonal chemistry, antibody lot tracking |
| **RNA-class contamination** | rRNA or tRNA carryover in mRNA samples | Fraction purity markers, depletion controls, LC-MS/MS on purified fractions |
| **Reverse-transcription artifact** | Structure, damage, enzyme bias | Unmodified controls, multiple enzymes, chemical controls |
| **Mapping ambiguity** | Repeats, paralogs, isoforms | Unique-mapping filters, long reads, transcript-aware coordinates |
| **Batch effect** | Library preparation, antibody lot, pore chemistry | Randomized processing, replicates, batch modeling |
| **Model overfitting** | Direct-RNA modification callers | Held-out standards, knockout controls, versioned models |

False positives in epitranscriptomics often arise when a method-specific signal is interpreted as a chemical fact. An antibody peak can become a claimed modification site; a reverse-transcription mismatch can become an editing event; a nanopore current shift can become a named modification; a bulk mass-spectrometry peak can become an mRNA mark. Each upgrade requires additional evidence. The first artifact-control habit is to name the measured signal before naming the biological conclusion.

Sample handling can create or erase signals. RNA is chemically labile compared with DNA. Extraction conditions, pH, heat, metal ions, oxidative stress during processing, nucleases, and freeze-thaw cycles can damage RNA or enrich stable fragments. Highly modified and structured RNAs can survive treatments that degrade other RNAs, creating contamination in purified fractions. Small amounts of rRNA or tRNA contamination can be serious because those RNAs are abundant and modification-rich. Therefore RNA integrity, fraction purity, and contamination checks are essential when claiming low-abundance mRNA modifications.

Library construction can create false patterns. Fragmentation can make some regions overrepresented. Adapter ligation can be biased by end chemistry, structure, and sequence. Reverse transcription can stop at structured or damaged regions. Polymerase chain reaction amplification can distort relative abundance. Alignment can place reads incorrectly among paralogs, repeats, pseudogenes, and isoforms. Differential expression can masquerade as differential modification if input abundance changes are not modeled. A site that appears condition-specific may simply be better covered in one condition.

Perturbation artifacts are common. Writer knockdown can slow growth, change cell-cycle state, activate stress pathways, alter RNA abundance, or change RNA composition. An eraser overexpression experiment can expose RNAs to nonphysiological enzyme levels. Reader depletion can affect stability of other proteins or alter phase-separated compartments. A clean modification mechanism should separate direct catalytic or binding effects from downstream physiology. Acute perturbation, catalytic mutants, rescue, matched input RNA, and time-course designs help reduce these ambiguities.

Computational artifacts deserve equal attention. Peak callers, modification callers, and direct-RNA base-calling models require assumptions about background, coverage, sequence context, and training labels. A model trained on one organism, RNA class, or modification standard may not generalize to another. Benchmarking against synthetic controls, in vitro transcribed RNA, knockout or writer-depleted samples, and independent datasets is essential. Transparent reporting should include sensitivity, specificity, false-discovery control, coverage thresholds, replicate concordance, and the version of the model or software used.

A special false-positive mode is semantic inflation. "Associated with m6A" may mean a gene is near an antibody peak, a transcript changes after writer depletion, a reader protein binds somewhere on the RNA, or a disease sample has altered total m6A. These are different claims. Another semantic problem is treating all changes after perturbing a writer as modification-dependent. A methyltransferase complex can have scaffolding roles, RNA-binding effects, or indirect transcriptional effects independent of catalysis. A reader can bind unmodified RNA or proteins. The wording should preserve the evidence level.

**Table 46.3. Claim Wording by Evidence Level.** Modification claims should be phrased to match the evidence level of the underlying assay, avoiding language that implies stronger support than the method provides.

| Evidence level | Preferred wording | Avoid wording |
| --- | --- | --- |
| **Enrichment peak** | Candidate modified region | Modified site |
| **Site assay without occupancy** | Candidate or supported site | Fully methylated transcript |
| **LC-MS/MS bulk signal** | Modified nucleoside detected in sample | Site identified on transcript |
| **Writer knockdown signal change** | Enzyme-dependent signal | Direct writer mechanism |
| **Reader co-IP** | RNA association | Modification-dependent reader binding |
| **Perturbation phenotype** | Associated phenotype | Proven modification function |

## Recent Consensus

Recent consensus is cautious but not skeptical in the dismissive sense. RNA modifications are real, widespread, chemically diverse, and biologically important. The most established examples include many tRNA and rRNA modifications, cap modifications, A-to-I editing, and well-supported m6A pathways. The open questions concern how many proposed mRNA and noncoding RNA sites are reproducible, what fraction of candidate sites are highly occupied, which changes are direct responses rather than cell-state markers, and which disease associations are causal. Good artifact control does not weaken epitranscriptomics; it turns modification maps into reliable biology.

## Open Questions, Controversies, Deprecated Models, and Common Misconceptions

Open questions:

- Which combinations of independent evidence are sufficient to promote a candidate region to a chemically identified, site-localized, and quantified modification claim?
- How should uncertainty be propagated when modification fraction, transcript abundance, cell composition, and assay error all change between conditions?
- Which writer, eraser, and reader assignments outside the best-studied m6A systems have direct catalytic, structural, and site-dependent functional support?
- How portable are evidence thresholds across RNA classes, organisms, sequence contexts, and modification chemistries?

Common misconceptions:

- "A detected modification is automatically functional." Detection alone does not establish function; function requires evidence for stoichiometry, context, perturbation, and consequence.
- "Low-stoichiometry modifications are unimportant." Low occupancy can matter if the modified molecules have a specialized fate, but that specialized fate must be demonstrated.
- "High-confidence chemical identification is automatically a site map." Chemical identity and transcript-residue location are separate claims and require different evidence.
- "A site map is a mechanism." A site map identifies candidate residues; mechanism requires perturbation, reader or enzyme evidence, and a measured molecular consequence.
- "A writer, eraser, or reader label is enough to define mechanism." Enzymology, substrate specificity, binding, localization, and perturbation evidence are still needed.
- "A disease association is a therapeutic target." Disease association is only a starting point; causal biology and disease synthesis belong to [Chapter 52](chapter1048.md).
