Chapter 6. RNA Terminology, Nomenclature, Ontologies, Identifiers, and Entity Semantics

Scope Note

This chapter teaches the scientific semantics of RNA biology: how RNA entities are named, how names differ from identifiers, how identifiers differ from versions, how ontologies define relationships, and how scoped scientific claims can be recorded without losing evidence context. These topics may look administrative, but they are biological. A sentence about an RNA can mean a DNA locus, a primary transcript, a processed transcript isoform, a mature small RNA, a modified residue, a structural motif, a molecular family, a disease biomarker, a therapeutic target, or an assay signal. If the entity layer is unclear, the biological claim becomes unclear.

The chapter is not a catalog of RNA databases or a guide to record matching and data deposition. Chapter 19 treats the RNA resource landscape, accession scope, database comparison, and practical entity matching; Chapter 18 treats transcript models and annotation versioning; Chapter 46 treats RNA modification detection limits; Chapter 140 treats RNA family discovery and covariance models; and Chapter 144 treats workflow provenance, FAIR metadata, deposition, licensing, version monitoring, and maintenance. Chapter 6 provides the semantic foundation used by those chapters. The goal is to make the reader able to ask, for every RNA term, “What exact object is being named, under which convention, in which organism, under which version, and with what evidence?”

Executive Summary

RNA terminology is difficult because RNA objects are layered. A DNA locus may produce several primary transcripts. A primary transcript may be capped, spliced, cleaved, edited, tailed, or modified. A processed transcript may have multiple isoforms. A mature RNA may be localized, translated, degraded, packaged, bound by proteins, or used as a guide. A sequence may belong to a family whose members differ across organisms. A name may refer to any one of these layers. Therefore, an RNA name should not be treated as a complete scientific record.

Names and identifiers serve different functions. A name is a human-readable label such as XIST, MALAT1, U1 snRNA, let-7, 18S rRNA, m6A, or guide RNA. An identifier is a lookup handle such as a RefSeq accession, Ensembl transcript ID, HGNC gene symbol record, Rfam accession, miRBase entry, MODOMICS entry, DOI, PMID, or genome coordinate under a specified assembly. A good RNA statement often needs both: the name for readability and the identifier for disambiguation. A good statement also needs scope, because a stable identifier does not always define the exact molecular state being discussed.

Standardized nomenclature reduces ambiguity but does not prove biology. Standardized long noncoding RNA nomenclature helps distinguish lncRNA genes, transcript products, genomic contexts, and aliases, but a standardized lncRNA name is not evidence that the RNA has a molecular function. RNA base-pair nomenclature distinguishes interaction geometries, so that Watson-Crick, wobble, Hoogsteen-edge, sugar-edge, cis, and trans interactions are not collapsed into a vague phrase such as “base-paired”. Plant RNA-directed RNA polymerase nomenclature illustrates how names can be shaped by organism-specific pathway history.

Gene identifiers, transcript identifiers, isoform labels, and RNA family accessions are not interchangeable. A gene record refers to a genomic unit or curated gene concept. A transcript record refers to an RNA model or product. An isoform label distinguishes alternative transcript or product forms. A family accession, such as an Rfam accession, refers to a curated group of related RNAs built from sequence, structure, homology, or covariance evidence. Long-read RNA sequencing improves isoform discovery, but it also makes versioning and validation more important rather than less important.

Ontologies make relationship types explicit. An ontology is a controlled vocabulary plus defined relations among terms. For RNA biology, useful relation labels include transcribed-from, matures-to, part-of, has-modification, binds, cleaves, methylates, guides, regulates, located-in, associated-with, supported-by, and contradicted-by. Ontology-based RNA interaction knowledge graphs show why relationship labels matter: “RNA A binds protein B,” “RNA A is modified by enzyme B,” and “RNA A is associated with disease B” have different meanings and require different evidence.

Coordinates and processing states are part of an RNA entity’s scientific meaning. A genomic interval is interpretable only under a named assembly and strand convention; a transcript-relative coordinate depends on a particular transcript accession and version; and a mature-RNA coordinate may differ from precursor coordinates after cleavage, splicing, editing, or tailing. Likewise, the same sequence name can denote nascent RNA, precursor, mature product, aminoacylated tRNA, ribosome-bound mRNA, or a degradation intermediate. Organism, strain, cell type, compartment, developmental stage, condition, and assay must accompany identifiers whenever those qualifiers change what molecule or observation is meant.

Concept Inventory

  • Companion concept inventory defines the main vocabulary used in this chapter. A preferred name: the recommended readable label under a stated convention. An alias is an alternative label, but aliases can be exact synonyms, partial overlaps, historical names, or context-specific abbreviations. A deprecated name is an older label retained for search and interpretation but no longer recommended as the primary label. Entity resolution is the process of deciding whether two names, identifiers, records, or claims refer to the same biological object.
  • Gene identifier: a gene record or locus-level unit. A transcript identifier refers to a transcript model or RNA product. An isoform is one of multiple RNA or protein products produced from a locus by alternative transcription, splicing, cleavage, editing, polyadenylation, or other processing. An RNA family accession identifies a curated group of related RNAs rather than one molecule. An ontology defines terms and relationships. An evidence code labels the evidence type that supports an annotation or claim. A coordinate system defines the reference sequence, origin, orientation, and indexing convention used to locate a feature. A processing state specifies whether an RNA is nascent, precursor, mature, chemically modified, assembled, or degrading. A biological context qualifier states the organism, strain, cell type, compartment, developmental stage, disease state, or treatment that bounds an observation. A provenance record documents source, version, transformation, evidence, and citation purpose.
  • These definitions: not meant to replace prose. The terms are a scaffold for clearer prose. The reader should use them to notice when a sentence has silently moved from one layer to another.

What to Know Before Reading This Chapter

The reader should know the difference between DNA and RNA, the basic directionality of nucleic acids, and the idea that a gene can be transcribed into RNA. The reader should also know that RNA molecules are processed. A eukaryotic protein-coding pre-mRNA can be capped, spliced, and polyadenylated; a transfer RNA can be trimmed, modified, and charged; a microRNA can be processed from a longer transcript into a hairpin precursor and then into a mature regulatory RNA; a ribosomal RNA can be cleaved, modified, and assembled with proteins. These processing steps matter because names often attach to different stages of the same pathway.

The chapter uses several recurring examples. Long noncoding RNAs show how a locus, gene symbol, transcript product, genomic context, and functional claim can be confused. MicroRNAs show how a primary transcript, hairpin precursor, mature 5′ arm, mature 3′ arm, seed family, and target-regulatory claim require different terms. RNA modifications show why a chemical residue, an assay signal, an enzyme pathway, a site-specific modification claim, and a phenotype are different record types. RNA structure nomenclature shows why a contact must be named by geometry rather than by a generic word such as “pairing.”

The most important habit is to identify the layer before interpreting the claim. When a sentence says “this RNA is conserved,” ask whether the conserved object is a genomic locus, mature RNA sequence, secondary structure, Rfam family model, processing pathway, or function. When a sentence says “this RNA is upregulated,” ask whether the measurement is gene-level abundance, transcript-level abundance, isoform usage, mature small-RNA count, modification signal, or a disease panel score. When a sentence says “this RNA causes disease,” ask whether the evidence shows expression, variant association, perturbation, rescue, molecular mechanism, clinical utility, or therapy response.

6.1. RNA names, aliases, and organism-specific conventions

An RNA name is a readable label attached to an RNA-related object. Names are indispensable because scientists need short words for communication. But RNA names often carry history rather than a clean ontology. Some names come from discovery order, such as U1 small nuclear RNA. Some come from size, such as 7SK RNA. Some come from biochemical behavior, such as transfer RNA. Some come from phenotype, locus, organism, disease association, family membership, or method. Some names were assigned before the molecular object was fully understood.

Figure 6.1. RNA Entity Layers From Locus to Mature Molecule

Figure 6.1. RNA Entity Layers From Locus to Mature Molecule. A single DNA locus produces a cascade of distinct record layers: a gene identifier at the genomic level, a primary transcript, one or more processed isoforms, and mature RNA products that may carry chemical modifications and bind partner proteins. Each arrow in this pathway introduces a new record type with its own identifier namespace, versioning requirements, and evidence standards. The figure shows why a single readable name cannot stand in for entity-layer metadata: “MALAT1,” for instance, can refer simultaneously to a chromosomal locus, a gene record, an abundant nuclear-retained transcript, and a disease-biomarker label.

A preferred name is the label recommended under a chosen convention. Preferred names make text easier to read and search. A preferred name should be paired with a scope statement when ambiguity is possible. For example, “MALAT1” is a useful name, but it can refer to a human lncRNA gene, the abundant nuclear-retained transcript, related transcripts in other mammals, or a disease-associated biomarker claim. “let-7” can refer to a family of microRNAs, a particular mature microRNA sequence, a genomic hairpin, or a regulatory concept. “guide RNA” can refer to a CRISPR guide RNA, a mitochondrial RNA-editing guide RNA, a small nucleolar RNA guide, or a synthetic antisense guide depending on context.

An alias is an alternative label. Aliases are valuable for search because older papers, reagent catalogs, database records, and clinical reports may use different names. Aliases are also dangerous because not all aliases are exact synonyms. A historical name may refer to an activity before the responsible RNA was mapped. A gene name may be shared by a protein-coding locus and its mRNA. A family name may include several paralogs. A disease biomarker label may represent a panel rather than a molecule. An abbreviation can collide with an unrelated term in another field.

Organism-specific conventions add another layer. Human gene symbols are curated in a specific ecosystem, and HGNC tracks approved human symbols, names, aliases, and cross-references. Mouse, fly, yeast, plant, bacterial, archaeal, and viral communities have their own conventions. Capitalization, italics, locus tags, strain designations, segment names, and ortholog names may differ. A bacterial small RNA name may be tied to one strain’s genetic screen and may not map cleanly to another strain. A plant RNA-directed RNA polymerase name reflects plant RNA silencing history and may not align with animal polymerase terminology.

Long noncoding RNA nomenclature shows why names need rules. A lncRNA gene may overlap a protein-coding gene, lie antisense to another transcript, originate near an enhancer, or include repeat-derived sequence. A lncRNA name that looks like a protein-coding gene name can mislead readers into assuming a protein product. Standardized lncRNA nomenclature helps separate lncRNA genes, transcript products, genomic context, and aliases. The nomenclature is a communication standard, not functional proof. A named lncRNA can be functional, nonfunctional, condition-specific, or incompletely characterized. Functional evidence belongs to Chapter 5 and the lncRNA chapters.

Structure nomenclature gives a second lesson. In DNA-centered vocabulary, “base pair” often evokes Watson-Crick A-U or G-C pairing, with G-U wobble as a familiar exception. RNA structure is richer. Bases can interact through Watson-Crick, Hoogsteen, or sugar edges; glycosidic bonds can be in cis or trans orientation; and many noncanonical interactions stabilize motifs. The Leontis-Westhof classification gives RNA structural biology a precise vocabulary for these geometries. Without such vocabulary, a reader may merge chemically different contacts and misunderstand folding, recognition, or motif conservation.

Method acronyms should also be treated as names with context. “CLIP” may refer to several crosslinking and immunoprecipitation protocols. “SHAPE” may refer to a chemical probing principle or a specific mutational profiling implementation. “RNA-SIP” refers to RNA stable isotope probing, not to stable RNA molecules. A method label should be accompanied by enough protocol detail to tell what observable was generated.

Box 6.1. Common RNA Naming Traps

  • Gene symbol versus transcript ID. A gene symbol (e.g., MALAT1) refers to a locus-level record; a transcript accession refers to a specific RNA model. Using them interchangeably shifts the measurement layer from a gene-aggregate value to an isoform-specific value without notice.
  • Precursor miRNA versus mature miRNA arm. pre-miR-21 (the hairpin) and miR-21-5p (the mature guide strand) are different biological objects with different database records, different cellular abundances, and different functional contexts. A claim about one does not transfer automatically to the other.
  • RNA family accession versus individual RNA. An Rfam accession identifies a curated family model; it does not identify a specific transcript in a specific organism, tissue, or developmental stage. Family membership supports relatedness, not shared expression or identical function.
  • Chemical modification name versus detection signal. “m6A” names a chemical residue; an antibody-enrichment peak or a direct-sequencing signal names an assay output. The signal requires method-specific interpretation, threshold decisions, and orthogonal validation before it can be treated as a site-specific modification claim.
  • RNP complex name versus RNA component. “snRNP” refers to a ribonucleoprotein complex; “U1 snRNA” refers to the RNA component within it. Functional claims about the complex do not automatically apply to the RNA in isolation, and the RNA cannot be studied as if it were the full complex.
  • Method acronym collision. CLIP encompasses multiple crosslinking and immunoprecipitation protocols with different crosslinking chemistries and downstream biases. SHAPE names a chemical-probing principle but also a specific mutational-profiling implementation. RNA-SIP refers to stable-isotope probing of RNA, not to stable RNA molecules. Method labels should be accompanied by enough protocol detail to identify the observable generated.

Caution

Do not overgeneralize: a familiar RNA name is not an unambiguous entity. Before comparing claims, check organism, genomic locus, transcript model, mature molecule, modification state, family membership, and naming convention.

6.2. Gene, transcript, isoform, and family identifiers

RNA biology uses identifiers because names alone are too ambiguous. An identifier is a structured handle in a defined namespace. Identifiers vary in scope. A gene identifier points to a gene-level concept or locus. A transcript identifier points to an RNA transcript model or product. An isoform label distinguishes alternative products. A family accession points to a group of related RNAs. A genome coordinate points to a location under a specified genome assembly. A publication identifier points to a source, not to the biological entity itself.

Table 6.1. RNA Naming Layers and Common Hazards. Each row describes one entity layer in the RNA naming hierarchy, illustrating the recommended identifier, a common alias hazard, and an evidence or versioning consideration relevant to that layer.

Entity layer Example Best identifier Common alias hazard Evidence or version note Related chapters
Locus MALAT1 locus Genome coordinate (assembly-specific) Gene symbol used for both DNA element and transcript Evidence may be DNA-level; assembly version required Chapter 2, Chapter 18
Gene HOTAIR gene HGNC symbol + gene ID Name shared with protein-coding gene Gene record aggregates multiple transcripts Chapter 18, Chapter 19
Primary transcript XIST pre-mRNA Versioned transcript accession (NR_ or ENST) Gene name used when transcript is meant Transcript model changes across annotation releases Chapter 18
Transcript isoform TP53 isoform β Versioned transcript accession Isoform label shared across species Model support varies; long-read data can add isoforms Chapter 18, Chapter 19
Mature RNA let-7a-5p miRBase mature accession (MIMAT) Seed-family name applied to one arm Mature arm identity depends on processing context Chapter 19
Precursor RNA pre-miR-21 miRBase precursor accession (MI) Mature name applied to precursor Hairpin and mature products are different records Chapter 19
miRNA arm miR-21-5p vs miR-21-3p miRBase MIMAT accession per arm miR-21 used without arm designation Arm usage can shift; both arms may be functional Chapter 19
RNA family SSU rRNA family Rfam accession (RF number) Species-specific rRNA name used for whole family Family model updated per Rfam release Chapter 140
Modification site m6A at GGACU motif MODOMICS entry + transcript position Modification type used as site-specific claim Site evidence depends on detection method and stoichiometry Chapter 46
Structure motif GNRA tetraloop Rfam motif or Leontis-Westhof symbol “loop” used without specifying geometry Structure may be predicted, probed, or crystallographic Chapter 53
Protein-RNA interaction HuR binding VEGF 3′UTR RBP-binding database record Gene name used for the protein-RNA complex Interaction type (direct vs indirect) and method required Chapter 84, Chapter 85
Disease association HOTAIR in breast cancer Disease ontology term + study accession lncRNA name used as a disease-gene name Association type spans expression to clinical evidence Chapter 90

The gene-transcript distinction is the first source of confusion. A gene is a genomic unit that can be transcribed. A transcript is an RNA molecule or model derived from transcription and processing. A gene-level expression value may aggregate reads across multiple transcripts. A transcript-level value estimates one model or product. If a paper says “gene X is upregulated,” the data may actually support transcript abundance, exon usage, promoter usage, or an aggregate across isoforms. The statement is useful only when the measurement layer is clear.

An isoform is one of multiple forms produced from a locus. Isoforms can arise through alternative transcription start sites, alternative splicing, alternative cleavage and polyadenylation, RNA editing, differential processing, or alternative maturation. Some isoform labels are strongly supported by full-length sequencing, junction reads, proteomics, or targeted validation. Other isoform labels are transcript models with limited support. Therefore, an isoform identifier is not the same as proof that the exact molecule is abundant, stable, translated, or functional in every sample.

Long-read RNA sequencing has made isoform complexity easier to observe because reads can span longer transcript regions than short-read fragments. This is especially valuable in disease transcriptomes, where abnormal splicing, fusion transcripts, retained introns, and alternative terminal exons can be clinically relevant. Single-molecule transcriptomics further emphasizes that transcript molecules can vary in start sites, splice junctions, poly(A) tail features, modifications, and degradation state. These methods improve annotation, but they do not remove the need for versioned transcript models, orthogonal validation, and platform-bias awareness.

Transcriptome primers emphasize that gene sets and transcript sets are analysis objects, not natural constants. A gene set can change when a genome assembly changes, when an annotation release merges or splits genes, or when a transcript model is retired. A scientifically interpretable statement should state the annotation release, genome assembly, and identifier namespace when those choices change what the named entity means. Operational mapping procedures belong to Chapter 19 and Chapter 144.

RNA family identifiers answer a different question from transcript identifiers. Rfam curates RNA families using alignments, consensus secondary structures, and covariance models. An Rfam accession supports membership in a curated family under a model. It does not prove that a particular locus is expressed in a particular tissue, that the RNA is processed into a mature product, or that all family members share identical function. Family membership is strong evidence for relatedness, not automatic evidence for context-specific biology.

MicroRNAs illustrate the need to separate identifier layers. A primary microRNA transcript is processed into a precursor hairpin. The hairpin may produce mature RNAs from the 5′ or 3′ arm. Mature products can share a seed sequence with family members and can regulate overlapping target sets. miRBase records connect microRNA sequences, hairpins, mature products, names, and literature-supported annotation. A claim about “miR-21” should say whether it concerns the genomic locus, precursor hairpin, mature miR-21-5p product, a target relationship, or a disease biomarker association.

Figure 6.2. Name, Identifier, Alias, and Version Crosswalk

Figure 6.2. Name, Identifier, Alias, and Version Crosswalk. A worked crosswalk traces one RNA name through its semantic fields: preferred name, aliases, deprecated labels, identifiers from gene, transcript, mature-product, and family namespaces, organism, genome assembly or nomenclature version, and citation purpose. Using a lncRNA or miRNA example, the figure shows how the same readable label can legitimately denote a gene, transcript model, mature product, family, or disease-association concept, and why each relation requires an explicit scope statement and version tag. Resource selection and record matching hand off to Chapter 19.

Identifier systems can also be versioned. A RefSeq accession with a version suffix, an Ensembl transcript version, a GENCODE release, a genome assembly name, and a database release date can change interpretation. RefSeq curation standards and long-running reference-sequence practices are relevant because reference records are updated over time. When a paper predates a current annotation, a reader may need historical mapping rather than current-name lookup.

Caution

Do not overgeneralize: one locus does not equal one RNA, one RNA name does not equal one isoform, one isoform label does not prove mature molecule abundance, and one family accession does not prove one function.

6.3. Modification, structure, method, disease, and evidence ontologies

An ontology is a controlled vocabulary with defined relationships. A plain vocabulary can list terms such as mRNA, spliceosome, pseudoknot, N6-methyladenosine, CLIP-seq, cardiomyopathy, and evidence. An ontology states how terms relate. A mature mRNA is derived from a precursor. A modified residue is part of an RNA molecule. A methyltransferase can catalyze a modification. A protein can bind an RNA. A method can detect an assay signal. A disease association can be supported by a clinical study, a genetic association, a perturbation experiment, or a biomarker performance study.

Biomedical ontology work gives this principle a standards base rather than only a local editorial preference. The OBO Foundry describes shared principles for interoperable biomedical ontologies, the Gene Ontology and Sequence Ontology show how controlled terms and relations support biological annotation, and the Evidence Ontology gives a vocabulary for describing evidence types rather than leaving support implicit.

The relationship label is as important as the term. “Binds” is not the same as “regulates.” “Associated with” is not the same as “causes.” “Detected by” is not the same as “installed by.” “Located near” is not the same as “part of.” A graph or table that links an RNA to a protein without a relationship label hides the claim. A graph that says an RNA binds a protein, is methylated by an enzyme, guides a nuclease, and is associated with a disease makes four different claims with four different evidence standards.

RNA modification records need chemical, enzymatic, site-specific, and phenotypic layers. A modification such as N6-methyladenosine has a chemical identity. A database such as MODOMICS can connect modified residues, pathways, enzymes, structures, and biological contexts. A site-specific statement such as “this adenosine in this transcript is methylated in this cell type” requires detection evidence. A functional statement such as “this site changes translation through a reader protein” requires perturbation, reader binding, rescue, and mechanism-consistent readouts. Direct RNA sequencing and related reviews show why detection methods, model assumptions, and validation must be recorded separately from biological interpretation.

Structure records need similar separation. A secondary-structure model, a chemically probed structure, a cryo-EM density, a crystallographic base triple, a conserved Rfam covariance pattern, and a molecular dynamics ensemble are not the same object. A structure ontology should distinguish predicted, probed, resolved, conserved, and functionally tested structures. Base-pair nomenclature helps because structure claims often depend on specific contact geometry rather than a generic statement of pairing.

Method ontologies prevent method names from becoming conclusions. Ribo-seq detects ribosome-protected fragments under a protocol; interpretation as translation depends on controls and context. CLIP-seq recovers crosslinked RNA fragments associated with a protein or complex; interpretation as direct regulatory binding requires additional evidence. SHAPE-MaP reports nucleotide reactivity under chemical and analytic assumptions; interpretation as a cellular structure requires controls and model awareness. Direct RNA sequencing can preserve native RNA features but has platform-specific error and signal interpretation issues. A method term should name the observable and the limitation.

Disease ontologies are needed because disease names are also layered. A disease term can refer to a clinical syndrome, molecular subtype, histopathologic category, genetic condition, trial eligibility category, or patient-reported outcome. RNA disease claims may involve altered expression, splicing, modification, RNA-binding protein mutation, repeat expansion, viral RNA, circulating RNA biomarker, or therapeutic RNA response. A disease ontology term does not define mechanism by itself. It organizes the clinical or phenotypic context in which the RNA claim was made.

Evidence ontologies and evidence codes label how a claim is supported. The strongest evidence code for one claim may be irrelevant to another. Sequencing evidence can support detection and abundance. Imaging can support localization. Biochemical reconstitution can support direct activity. Genetic perturbation and rescue can support causality. Clinical trials can support efficacy under a trial design. The Evidence Ontology is useful because it treats evidence as a structured annotation object, not as a generic citation note. Chapter 5 explains evidence strength; Chapter 6 explains why evidence labels and citation purposes must be tied to entity and relationship labels. The citation-purpose vocabulary in Table 6.3 anticipates the fuller scientific-provenance treatment in Section 6.5.

Table 6.3. Citation Purpose Vocabulary. Each row defines one citation-purpose label, explains what it means, gives a representative example, identifies a common misuse, and lists the evidence metadata that should accompany a citation with that purpose.

Citation purpose Meaning Example use Common misuse Needed evidence metadata
Background Provides general context or historical framing Citing a review for RNA processing overview Treating a background source as primary evidence Topic relevance, year, review scope
Consensus review Summarizes agreed-upon findings across multiple studies Citing a systematic review or meta-analysis Using one review to resolve a genuinely contested claim Review scope, inclusion criteria, publication date
Primary evidence Original experimental data directly supporting the claim Citing the paper that first measured the effect Using a replication or confirmatory paper as if it were primary Experiment type, organism, method, controls
Method origin Identifies the source protocol or analytical technique Citing the original CLIP-seq protocol paper Citing the method paper as evidence of a biological result Protocol version, key assumptions, intended observable
Method benchmark Reports performance evaluation of a method or pipeline Citing sensitivity and specificity data for a sequencing workflow Using a benchmark result as evidence of biological truth Dataset, metrics, comparison baseline
Database source Identifies the database record used as a data input Citing miRBase for the sequence used in analysis Treating a database entry as independent experimental support Database version, curation rules, accession
Controversy Notes a conflicting or qualifying result Citing a study that challenges the prevailing model Omitting a controversy to strengthen the main claim Study design, sample size, effect size, context
Clinical/regulatory Provides evidence from clinical trials or regulatory records Citing a Phase II trial for therapeutic efficacy Using an animal-model paper as if it were clinical evidence Trial phase, population, endpoint, approval status
Historical landmark Credits the first description or naming of an entity Citing the original discovery of a noncoding RNA Treating a historical paper as a current methodological standard Year, organism, technique available at the time
Unresolved source need Marks a scoped claim that still requires a verified source before final use Flagging a narrow provenance gap without filling it with a loose citation Treating an unresolved flag as evidence for the claim Gap description, required claim type, resolution owner, review status

Ontology-based RNA knowledge graphs show the practical value of this discipline. A graph of RNA interactions must distinguish entity type, relationship type, evidence, and context; otherwise, binding, regulation, modification, disease association, and literature co-occurrence can collapse into one uninterpretable edge. Embedding methods can analyze such graphs, but the computation inherits curation errors, missing relationships, and ambiguous identifiers.

Caution

Do not overgeneralize: an ontology term is not evidence. An ontology makes the claim type explicit. Evidence comes from the experiments, observations, curation rules, and sources attached to the term and relationship.

Figure 6.3. Ontology Relationship Types in RNA Biology

Figure 6.3. Ontology Relationship Types in RNA Biology. The same pair of biological nodes is connected by a series of different labeled edges — transcribed-from, matures-to, binds, modifies, regulates, located-in, associated-with, supported-by, and contradicted-by — to show why relationship type matters as much as the nodes themselves. Collapsing these edges into a generic “related to” link would hide the difference between, for example, a direct biochemical binding interaction and a disease co-occurrence inferred from expression data. Each edge type also carries a different evidence standard and a different downstream implication for functional interpretation.

6.4. Coordinate systems, reference versions, processing states, and biological context qualifiers

An RNA coordinate is a position defined relative to a reference sequence. The reference may be a chromosome, genomic contig, transcript accession, precursor RNA, mature RNA, viral segment, or experimentally determined structure. A number without that reference is not a complete coordinate. The coordinate also needs an origin and indexing convention: some systems count the first nucleotide as one, whereas programming interfaces commonly count from zero; genomic intervals may be closed, open, or half-open; and strand orientation determines whether increasing genomic coordinates follow or oppose the RNA 5-prime-to-3-prime direction. Confusing these conventions produces one-nucleotide shifts, reversed intervals, and incorrect sequence extraction.

Genome coordinates require an assembly and, for many organisms, a strain or haplotype. A human interval on GRCh37 is not automatically the same interval on GRCh38. Sequence changes, gap corrections, alternate loci, and contig restructuring can move or eliminate a mapped feature. Conversion tools can project intervals between assemblies, but projection is an inference: duplicated sequence, structural variation, assembly gaps, and paralogy can make a mapping ambiguous or one-to-many. A lifted coordinate should retain both source and destination references and should not be described as exact when the mapped sequence or feature changed.

Transcript coordinates require a transcript accession and version because alternative transcription, splicing, and polyadenylation change the coordinate frame. Nucleotide 200 of one MALAT1 record need not represent nucleotide 200 of another record. A variant described in genomic coordinates may map to different transcript positions in different isoforms, and its functional label may change from exonic to intronic or untranslated depending on the selected model. RefSeq curation and transcriptome annotation practice therefore treat sequence accessions, version suffixes, annotation releases, and assemblies as complementary rather than interchangeable metadata.

Processing changes coordinates as well as molecular identity. A primary microRNA transcript contains a precursor hairpin, and cleavage releases mature 5-prime and 3-prime products whose positions are usually expressed in a mature-product frame. A tRNA gene contains sequences that may be removed during leader, trailer, or intron processing; the mature tRNA may acquire a non-templated CCA end and numerous modifications. Ribosomal RNAs are cleaved from longer precursors, and messenger RNAs acquire exon-exon junctions and poly(A) tails. A site should therefore be labeled as genomic, precursor-relative, transcript-relative, or mature-product-relative. Coordinate conversion without an explicit processing map can silently assign a modification, variant, or binding site to the wrong nucleotide.

Chemical state is another qualifier. The symbol m6A denotes N6-methyladenosine as a chemical species, but a statement about an m6A site also needs an RNA identity, nucleotide position, molecule state, detection method, and preferably an estimate of confidence or stoichiometry. Direct RNA sequencing, antibody enrichment, chemical conversion, and enzyme-assisted methods generate different observables and error profiles. A modification signal on a precursor does not necessarily imply the same occupancy on the mature RNA, because processing, localization, selective decay, or demodification can alter the population. MODOMICS links chemical identities to pathways and RNA contexts, but a database record does not replace sample-specific evidence.

Biological context qualifiers state where and when an observation applies. At minimum, organism and molecular entity should be explicit. Depending on the claim, strain, sex, genotype, cell type, tissue, developmental stage, subcellular compartment, disease state, environmental condition, infection status, treatment, and time point may also be necessary. The same transcript can use different isoforms across tissues; a microRNA can switch dominant arm; a viral RNA can differ among isolates and hosts; and an RNA modification can vary with growth state or stress. Context is not optional decoration when it changes the molecule population or mechanism.

Assays impose an observational coordinate system. Short-read RNA sequencing reports fragments aligned to a reference and summarized under an annotation. Long-read sequencing reports molecule-scale reads but still has platform errors, incomplete ends, and alignment choices. CLIP methods report crosslink-associated fragments, chemical probing reports reactivity, and imaging reports spatial signals under probe and segmentation assumptions. An assay coordinate is not automatically a molecular boundary: a peak summit is not necessarily the contacted atom, a read end is not necessarily a biological RNA end, and a modification call is not necessarily a stoichiometric chemical measurement.

A practical normalization record for an RNA observation therefore contains the readable name; external identifier and version; reference assembly or sequence; coordinate convention and strand; processing and chemical state; organism and sample context; assay; and mapping or transformation history. Missing information should remain explicitly unknown rather than being filled by a convenient default. The purpose is not bureaucratic completeness. It is to prevent apparently identical coordinates from merging observations that concern different sequences, molecule states, or biological systems.

Figure 6.5. Coordinate Frames Across an RNA Life Cycle

Figure 6.5. Coordinate Frames Across an RNA Life Cycle. The figure aligns five coordinate frames for a spliced messenger RNA and a microRNA: genome, primary transcript, spliced or precursor RNA, mature product, and assay-derived signal. Each frame labels reference sequence and version, strand, coordinate origin, processing boundaries, non-templated additions, and the points at which conversion can be ambiguous or one-to-many. A side panel contrasts exact sequence-preserving conversion with uncertain assembly lift-over and transcript-model projection.

Table 6.4. Coordinate and Context Checklist. Coordinate annotations are interpretable only when the reference assembly or transcript, strand, interval convention, feature version, and biological context travel with every value; conversion without those qualifiers can change the claimed RNA object.

Coordinate frame Required reference Required qualifiers Common conversion hazard RNA example
Genomic Assembly, contig, strand Organism, strain or haplotype, interval convention Assembly lift-over can be absent, ambiguous, or one-to-many Human splice-site variant
Transcript-relative Accession and version Isoform, origin, strand convention Alternative exons and transcript updates shift positions MALAT1 or protein-coding transcript site
Precursor-relative Precursor sequence record Processing stage, cleavage model Mature boundaries may vary by cell type or assay pre-miRNA hairpin
Mature-product Mature sequence record Arm or product identity, terminal additions, chemical state Non-templated ends and isomiRs alter length miR-21-5p or mature tRNA
Structure-relative Structure record or alignment Chain, residue naming, model or conformer Missing residues and alignment gaps break simple numbering rRNA or ribozyme structure
Assay-derived Raw-data reference and analysis version Protocol, sample, peak or call definition Peak summits and read ends are not exact molecular boundaries CLIP peak or modification call

Caution

Do not overgeneralize: a nucleotide number is not a universal address. Interpret it only after identifying the reference sequence, version, coordinate convention, strand, processing state, and relevant biological context.

6.5. Scientific provenance, evidence attribution, and citation semantics

A semantic relation connects one scientific object to another. A gene can be transcribed into a transcript; a precursor can mature into a product; an RNA can belong to a family; a residue can bear a modification; and a source can support or contradict a claim. These relations are useful only when their meaning is explicit. Identity, orthology, family membership, shared sequence, shared locus, historical alias, evidence relation, and loose topical relevance are not interchangeable. Practical cross-resource mapping belongs to Chapter 19.

Provenance is the record of where a statement came from and what happened to it. A publication identifier such as DOI or PMID identifies a source, but provenance asks more precise questions. Does the source provide primary experimental evidence, a consensus review, a database update, a method description, a benchmark, a historical origin, a clinical trial, a regulatory standard, or a conflicting result? Did a curator extract the claim from a table, figure, supplement, or database entry? Was the claim transformed by normalization, identifier mapping, coordinate conversion, or evidence grading?

Citation purpose is the reason a source is cited for a specific statement. The same paper can have several purposes in different locations. A database paper can support the existence and scope of a database resource but not every biological conclusion stored in that database. A review can support a consensus summary but not replace the primary evidence for a key experiment. A methods paper can support how an assay works but not prove every downstream biological interpretation made with that assay. A clinical trial can support efficacy in a defined population but not a molecular mechanism unless that mechanism was tested. Versioning is a special form of scientific provenance. RNA claims are often version-dependent because genome assemblies, transcript annotations, accession records, nomenclature rules, and classification schemes change. A coordinate without an assembly can be uninterpretable. A transcript accession without a version can refer to an updated model. A family assignment without the model release can change meaning after a covariance model is revised. Transcriptome data primers and reference-sequence curation records illustrate why versioned entity semantics matter. Operational version monitoring belongs to Chapter 144.

Figure 6.6. One RNA Identity Record Across Processing and Version Changes

Figure 6.6. One RNA Identity Record Across Processing and Version Changes. One generic RNA record is followed from genomic locus through primary transcript, processed isoform, mature RNA, and a modification-site claim while the coordinate frame and entity layer change at each transition. A persistent metadata envelope carries reference, coordinates, accession, version, processing state, preferred name, aliases, evidence, and citation purpose without implying that one accession identifies every layer. A database-update ledger records revised coordinates, processing models, names, and evidence; a synonym-conflict branch shows why a shared alias cannot justify merging precursor and mature-product records; and a source-to-claim lane preserves citation purpose and transformation history.

Scientific provenance also records absence and uncertainty. If a relationship is inferred rather than directly observed, the statement should say so. If naming authorities or classification systems assign different meanings, the conflict should be preserved rather than silently collapsed. If a claim lacks direct support, an adjacent or topically related citation should not be presented as evidence for it.

The minimum semantic context for an RNA claim includes the RNA entity layer, organism or system, identifier namespace, version when relevant, method or evidence type, claim type, citation purpose, evidence strength, and caveats. Citation Typing Ontology (CiTO) provides a model for treating citation intent as a typed relation rather than an undifferentiated link. For example, a claim about a microRNA should state whether the entity is a precursor hairpin or mature arm; name the organism; specify the relevant sequence or identifier; state whether the evidence is sequencing, perturbation, target-site mutation, reporter assay, or clinical association; and cite sources for the exact purpose they serve. Operational provenance capture and FAIR packaging belong to Chapter 144.

Box 6.2. Minimum Semantic Context for an RNA Claim

  • Organism or system: species, strain, cell line, tissue, or developmental stage in which the claim was tested or observed.
  • Reference and namespace version: required whenever a coordinate, transcript accession, or family assignment is cited, because positions and entity meanings change across versions.
  • RNA class: mRNA, lncRNA, miRNA, tRNA, rRNA, snoRNA, snRNA, circRNA, or other recognized category.
  • Entity layer: locus, gene, primary transcript, isoform, mature product, precursor, modification site, structure motif, or interaction partner.
  • Identifier and aliases: preferred name, alternative labels, deprecated names, and identifiers, each with a semantic relation label such as exact synonym, partial overlap, historical name, or abbreviation.
  • Molecule state: precursor, mature, modified, bound, localized, degraded, or translated — the state the evidence actually reports.
  • Method and evidence type: sequencing, biochemical reconstitution, genetic perturbation, imaging, computational prediction, clinical measurement, or other.
  • Citation purpose: primary evidence, consensus review, method origin, database source, background, controversy, or unresolved source need.
  • Evidence grade: experimental observation, computational prediction, clinical association, or expert opinion.
  • Caveat: detection limits, method assumptions, conflicting reports, isoform-model uncertainty, or version dependencies.

Caution

Do not overgeneralize: a bibliography is not provenance. Provenance maps specific sources to specific claims, versions, evidence types, transformations, and caveats.

6.6. Deprecated names, conflicting classifications, and semantic ambiguity

A deprecated name is an older label that should no longer be used as the preferred name under the current convention. Deprecated names should not be erased. They remain essential for literature search, reagent tracking, historical interpretation, clinical record review, and database mapping. A reader who searches only the current preferred name may miss older studies. A reader who uses only an old name may merge distinct entities or miss current classification.

Deprecated names should be stored with status and scope. “Deprecated” should not mean “false.” Some deprecated names refer to real entities under older conventions. Others reflect obsolete models. A name first assigned to a phenotype may later be resolved into a gene, transcript, protein complex, or pathway. A name first assigned to one RNA species may later be split into several paralogs or isoforms. A name first used for a method may later be replaced because the method family diversified.

Conflicting classifications occur when systems classify by different criteria. CRISPR-Cas systems can be classified by effector architecture, evolutionary relationship, guide RNA use, nuclease mechanism, target class, or application. A review of CRISPR-Cas classification emphasizes that system classification changes as new variants are discovered and mechanistic understanding improves. RNA-protein interactions can be classified by RNA class, protein domain, cellular compartment, binding mode, regulatory consequence, or method. RNA structures can be classified by sequence family, secondary motif, tertiary fold, topology, ligand, or function.

No single classification answers every question. If the question is evolution, family and homology may be central. If the question is mechanism, molecular architecture and catalytic activity may matter more. If the question is therapeutic design, target accessibility, delivery route, chemistry, safety, and regulatory category may be decisive. A classification should state its axis.

Semantic equivalence asks what relation holds between two labels or identifiers, not how a matching pipeline should operate. Exact identity is only one possibility. Two terms may instead denote overlapping loci, different transcript versions, precursor and mature products, orthologs, paralogs, members of one family, or a historical name and its later molecular interpretation. If two labels share only an abbreviation, they are semantically ambiguous rather than equivalent. Practical record matching and cross-resource reconciliation belong to Chapter 19.

Common RNA semantic hazards include gene-symbol collisions, transcript-versus-protein confusion, precursor-versus-mature-RNA confusion, species-specific ortholog names, noncoding RNA names that resemble protein-coding symbols, repeat-derived transcripts with ambiguous reference placement, paralogous small RNAs, and method acronyms that collide with molecule names. Clinical and therapeutic contexts add hazards because a target name may denote a gene, transcript, exon, splice junction, mutation, repeat expansion, protein product, pathway, or patient subgroup.

The safest prose preserves ambiguity where ambiguity matters. Instead of saying “this RNA is the same as that RNA,” say “these records share a historical alias but differ in mature product,” or “this Rfam family contains homologous RNAs, but the specific human transcript model is represented by a separate accession,” or “this disease association uses a gene-level label and does not identify the causal transcript isoform.” Such wording may feel slower, but it prevents wrong biological inferences.

Caution

Do not overgeneralize: name conflicts are not clerical problems. They can change literature interpretation, target selection, reagent design, clinical interpretation, and therapeutic development.

Box 6.3. Deprecated Names Are Not Trash

  • Old names remain indispensable for literature retrieval. Papers, patents, reagent catalogs, and clinical records filed under a deprecated label will not be found by searching only the current preferred name. Erasing old names from a naming record or text means losing that search path.
  • Preferred names should be explicitly marked as such, paired with the convention or database that defines them (e.g., HGNC for human gene symbols, miRBase for miRNA names), so a reader can locate the current authoritative record without guessing.
  • Aliases should be recorded with relationship labels rather than as a flat list. The relevant distinctions include: exact synonym (refers to the same object under a different label), partial overlap (shares some features but not all), historical activity name (named before the molecular identity was known), organism-specific ortholog label (valid in one species context), and abbreviation collision (the same abbreviation used independently in different fields). A flat alias list implies all entries are interchangeable, which can cause entities to be incorrectly merged.
  • A naming conflict in prose warrants an explicit warning when two names are circulating for different reasons: the original entity was split into two distinct molecules, a common abbreviation clashes with an unrelated term in another field, or an organism-specific label has been adopted outside its original taxonomic scope without adjustment.

Biological Contexts Across RNA Classes

Messenger RNA terminology often starts with a protein-coding gene symbol, but mRNA claims may concern promoter choice, transcription start site, exon inclusion, coding sequence, untranslated regions, poly(A) site, cap structure, codon usage, modification state, localization, ribosome occupancy, decay intermediate, or protein product. A gene-level label is not enough when the claim involves isoform-specific regulation, nonsense-mediated decay, untranslated-region elements, or therapeutic targeting.

Transfer RNA terminology must distinguish gene copy, mature tRNA, anticodon, amino acid identity, isoacceptor, isodecoder, modification state, charging state, fragment, and mitochondrial versus cytosolic origin. A tRNA gene count does not directly tell how much mature charged tRNA is available. A tRNA fragment may be produced by stress cleavage or processing and should not be assumed to be equivalent to the parent tRNA.

Ribosomal RNA terminology must distinguish rRNA gene arrays, precursor rRNAs, processed mature rRNAs, modification sites, ribosomal subunits, ribosome biogenesis intermediates, and organellar rRNAs. A statement about “18S rRNA” in a metatranscriptomic sample may mean a marker sequence for taxonomic profiling, not a mechanistic claim about ribosome assembly.

MicroRNA terminology must distinguish primary transcript, precursor hairpin, mature arm, seed family, target site, target gene, target transcript isoform, and phenotypic effect. Small RNA sequencing can detect a mature sequence, but target repression requires additional evidence. A mature miRNA family can share seed-mediated targeting patterns while individual family members differ in expression, processing, arm usage, and nonseed sequence.

Long noncoding RNA terminology must distinguish DNA locus function, transcriptional act, RNA molecule, RNA domain, cis regulation, trans regulation, chromatin association, protein binding, and disease biomarker status. A locus deletion can affect DNA elements independent of the RNA product. A knockdown phenotype can be direct, indirect, or off-target. A standardized name helps the discussion but does not resolve mechanism.

Viral RNA terminology must distinguish genome segment, positive or negative polarity, antigenome, subgenomic RNA, defective interfering RNA, replication intermediate, packaging signal, strain, isolate, accession, and host context. The same gene product name may appear across viral lineages with different genome organization. Viral nomenclature therefore needs sequence accession, strain or isolate, and segment or transcript context.

Therapeutic RNA terminology must distinguish modality, target, chemistry, delivery, pharmacology, and regulatory category. An antisense oligonucleotide, siRNA, mRNA vaccine, self-amplifying RNA, guide RNA, aptamer, RNA-editing guide, and RNA-targeting small molecule are different modalities. A “target” may mean a transcript to degrade, a splice site to redirect, an antigen to express, an enzyme to inhibit, or a disease pathway to modulate.

Computational RNA biology depends on precise identifiers. Alignment, quantification, differential expression, isoform discovery, structure prediction, RBP-binding prediction, modification calling, single-cell annotation, and network inference all depend on the entity dictionary supplied to the software. If transcript models differ, expression estimates can differ. If gene symbols are mapped incorrectly, enrichment analysis can be wrong. If a family model changes, homology search results can change.

Benchmarking also depends on nomenclature. A benchmark set for RNA structure prediction needs to state whether structures are experimentally resolved, chemically probed, predicted, or curated from comparative evidence. A benchmark set for RBP binding needs to distinguish direct binding sites, immunoprecipitation-enriched regions, motif predictions, and regulatory targets. A benchmark set for RNA modifications needs to distinguish chemical identity, site localization, stoichiometry, method, and validation.

Clinical translation raises the cost of ambiguity. A diagnostic report or therapeutic design record cannot rely on a loose RNA name. It needs patient or sample context, organism, assembly or reference sequence, transcript or exon identifier, variant nomenclature when relevant, assay method, clinical interpretation, and update status. An RNA therapy record needs target transcript, intended molecular effect, chemistry, dose form, delivery route, biodistribution, pharmacodynamic readout, safety signal, and evidence level.

Engineering and synthetic biology add design-version issues. A synthetic mRNA has a coding sequence, untranslated regions, cap, poly(A) tail, nucleoside modifications, purification state, formulation, and delivery vehicle. A guide RNA has scaffold, spacer, chemical modifications, nuclease context, target coordinates, off-target model, and delivery context. A riboswitch design has ligand, aptamer domain, expression platform, host organism, and performance conditions. The designed sequence is a versioned object, not just a name.

Recent Consensus

  • RNA names should be paired with entity layer, organism, and identifier when ambiguity is possible.
  • Preferred names, aliases, and deprecated names serve different functions and should not be collapsed.
  • Gene identifiers, transcript identifiers, isoform labels, mature RNA names, and family accessions represent different record layers.
  • Standardized nomenclature improves communication but does not establish biological function.
  • Ontologies are most useful when they distinguish entity type, relationship type, evidence type, context, and version.
  • RNA modification, structure, disease, and method terms should not be interpreted without method and evidence scope.
  • Coordinates and identifiers are interpretable only with their reference versions, coordinate conventions, processing states, and biological contexts.
  • Provenance records should state citation purpose because a source can support background, consensus, primary evidence, method origin, database scope, controversy, clinical evidence, or an unresolved source need in different contexts.
  • Deprecated names should be preserved for search and historical interpretation while clearly marking current preferred names and unresolved ambiguity.

Open Questions, Controversies, Deprecated Models, and Common Misconceptions

Open questions:

  • How should rapidly changing transcript annotations be harmonized across short-read, long-read, single-cell, spatial, and direct RNA sequencing datasets?
  • Which RNA records should require mature-product, isoform, cell-type, modification, localization, and disease-context qualifiers by default?
  • Which semantic relations best represent partial equivalence, such as overlapping loci, related isoforms, family membership, and historical aliases, without forcing false one-to-one identity?
  • How should RNA modification terminology connect chemical identity, detection confidence, stoichiometry, enzyme mechanism, and phenotype without overclaiming site function?
  • How should citation-purpose and evidence-code vocabularies be standardized across RNA biology without flattening field-specific evidence?
  • How should older literature be mapped to current names when genome assemblies, transcript models, and classification systems have changed?

Common misconceptions:

  • “A name uniquely identifies an RNA.” An RNA name can refer to a locus, transcript, isoform, mature product, family, modification, structure, or functional concept.
  • “A gene symbol and transcript ID are interchangeable.” A gene symbol usually refers to a gene-level record, while a transcript ID refers to a transcript model or product.
  • “A database accession proves biological function.” A database accession identifies a curated record; function depends on evidence and curation rules.
  • “A family accession means all members do the same thing.” Family membership supports relatedness under a model, not identical expression or function in every organism.
  • “An ontology edge is a fact by itself.” An ontology edge is a structured claim that still needs evidence, context, and provenance.
  • “Old names should be deleted.” Deprecated names should be preserved for search and interpretation while clearly marked as nonpreferred.
  • “A citation near a paragraph supports every claim in the paragraph.” Citation purpose should be assigned to specific statements or claim records.