RNA biology depends on databases, but an RNA database is not a neutral box of names. Each resource defines a particular kind of entity, applies a particular evidence standard, and changes over time. A record may describe an RNA family, a gene locus, a transcript model, a processed RNA product, a modified residue, a binding interaction, a disease association, a therapeutic molecule, or a literature-supported alias. These objects are related, but they are not interchangeable.
This chapter explains how public RNA resources divide the biological landscape into records, how accession scopes differ, how resource coverage and curation models should be compared, and how related records can be matched without falsely declaring them identical. Rfam, miRBase, MODOMICS, RNAcentral, NONCODE, long noncoding RNA resources, circular RNA resources, HGNC, SILVA, RNA-binding protein resources, disease resources, and therapeutic resources illustrate different record models. The central lesson is simple but important: before using a database record or cross-reference, ask what object the resource represents, what evidence and coverage rules created the record, which release supplied it, and what ambiguity remains.
Chapter 6 owns the meaning of scientific terms, nomenclature rules, identifier layers, coordinate semantics, ontology classes and relations, and evidence-attribution semantics. Chapter 12 explains comparative RNA family inference, and Chapter 18 explains transcript models, isoforms, genome browsers, and annotation versioning. Chapter 144 owns operational provenance capture, FAIR metadata, licensing and reuse, repository deposition, workflow execution, version monitoring, and maintenance. This chapter retains only the short semantic and operational bridges needed to compare resource records and resolve them across databases.
RNA databases make RNA biology searchable, comparable, and reproducible. They also impose boundaries. A family database such as Rfam describes RNA families through curated alignments, consensus secondary structures, covariance models, and stable family accessions. An Rfam hit supports membership in a model-defined family, but it does not by itself prove that the sequence is expressed, processed into a mature RNA, or functional in a particular cell. A microRNA registry such as miRBase distinguishes hairpin precursors from mature microRNA products, so a name must state whether it refers to a locus, a precursor, a 5′ mature arm, a 3′ mature arm, a seed family, or a target relationship. A modification resource such as MODOMICS defines modified residues, pathways, and enzymes, but a modification record does not prove that a specific transcript carries that modification at a specific site under a specific condition.
Database identifiers operate at several record levels. A gene record, transcript record, mature-product record, family record, and assembly-specific feature record can all describe related biology without describing the same object. HGNC/Genenames.org, GENCODE, Ensembl, RefSeq, Rfam, and miRBase therefore provide complementary anchors rather than interchangeable identifiers. The nomenclature and coordinate semantics behind these layers are developed in Chapter 6; the database question here is which resource record the reader needs.
Specialized databases store different evidence types. SILVA supports ribosomal RNA sequence classification and microbial taxonomy, but it is not a general RNA nomenclature authority. RNA-binding protein interaction resources may store motif preferences, in vitro binding, cellular crosslinking enrichment, direct regulation, or functional consequences; those evidence types must be separated. Disease association resources may represent expression correlation, variant association, biomarker evidence, perturbation support, mechanistic evidence, or literature co-occurrence; a disease association should not be read as a mechanism unless the supporting evidence justifies that interpretation.
Cross-database mapping is therefore a biological and computational task, not a simple table join. A link can mean identity, same family, same locus, same sequence, same publication, related transcript, orthology, shared alias, or obsolete-to-current replacement. A defensible match names both resources and releases, the source and target accessions, the relationship asserted, the evidence used, and any unresolved alternatives. Operational capture, licensing, deposition, and long-term monitoring of those mappings belong to Chapter 144.
Entity resolution is the process of deciding which database record or set of records is meant by a name, accession, sequence, coordinate, or literature mention. The comparison must use organism, molecule class, entity level, sequence or coordinate version, family membership, evidence, source release, and ambiguity notes. When context is insufficient, entity-resolution analysis should preserve unresolved alternatives rather than force a one-to-one mapping that the evidence does not support.
Before reading this chapter, the reader should know three distinctions from earlier chapters. First, a name is not the same as an identifier. Names are human-readable and often reused; identifiers are assigned within a namespace and can be versioned. Second, a gene is not the same as a transcript, and a transcript is not necessarily the same as a mature RNA product. Third, an annotation is not a biological mechanism. A database may record that an RNA has a predicted family membership, a detected expression pattern, a reported binding site, or a disease association, but the biological interpretation depends on evidence.
Chapter 12 provides the family logic needed for Rfam. A covariance model combines sequence and secondary-structure information to detect homologous RNA families. Such models are powerful for RNAs whose function depends on secondary structure, but a model hit remains an inference about relatedness. Chapter 18 provides the transcript-model logic needed for RefSeq, Ensembl, GENCODE, and genome-browser records. A transcript identifier belongs to a particular annotation system and release; different resources may model the same locus differently.
This chapter uses four recurring examples. The first is let-7, a microRNA name that can refer to a family, a precursor hairpin, a mature arm, a seed class, a species-specific locus, or a regulatory literature topic. The second is U6, a small nuclear RNA name that can refer to a spliceosomal RNA, a gene family, a promoter used in expression vectors, or a normalization control. The third is N6-methyladenosine, commonly abbreviated m6A, which can refer to a chemical modification, a transcript site, a mapping signal, a writer-enzyme pathway, or a regulatory field. The fourth is tRNA-Leu, which can refer to an isotype, anticodon group, genomic locus, mature tRNA, mitochondrial tRNA, modified molecule, or disease-associated gene.
These examples show why database literacy is a core RNA skill. A researcher designing primers, annotating a genome, interpreting an RNA-binding protein dataset, comparing lncRNA catalogs, or checking an RNA therapeutic target must know which level of entity is being used.
An RNA database record is a scoped claim. It says that, under the rules of the resource, a particular object has been recognized and described. That object might be a family, sequence, transcript, product, modification, structure, interaction, or association. The same RNA biology can appear in several resources because each resource asks a different question.
For example, a structured noncoding RNA sequence in a bacterial genome might appear as a genomic feature in an annotation browser, as a family hit in Rfam, as a sequence in RNAcentral, as a taxonomic marker if it is ribosomal RNA, and as a literature mention in a review. Those records are linked by biology, but none of them automatically replaces the others. The family hit helps with homology. The genome annotation helps with coordinates. The sequence integration record helps with cross-resource lookup. The literature record helps with interpretation.
Table 19.1. RNA Resource Scope and Evidence. Major RNA database types differ in what entity their records describe and what biological claims those records support or cannot support alone.
| Resource type | Example resources | Main entity | Evidence stored | What the record does not prove |
|---|---|---|---|---|
| Family databases | Rfam | RNA family (alignment, covariance model) | Curated alignment, covariance model scores, seed sequences | Expression, maturation, or function in a given organism |
| Sequence integration resources | RNAcentral | Noncoding RNA sequence | Integrated sequences from expert databases | Biological activity or processing state |
| MicroRNA resources | miRBase | Hairpin precursor, mature microRNA product | Sequences, arm annotations, precursor-to-mature links | Which arm is functionally dominant or what targets are regulated |
| Modification databases | MODOMICS, RNA Modification Database | Modified nucleoside, biosynthetic pathway | Chemical structure, enzyme, pathway annotations | Site-specific modification in a given transcript, cell type, or condition |
| rRNA databases | SILVA | rRNA sequence, taxonomic classification | Aligned rRNA sequences, phylogenetic trees | Applies to rRNA only; not a general RNA nomenclature authority |
| lncRNA resources | NONCODE, LNCipedia | lncRNA locus, transcript model | Expression support, transcript models, functional annotations | Molecular mechanism without perturbation evidence |
| circRNA catalogs | circBase, circAtlas | circRNA back-splice junction | Junction coordinates, expression context | High abundance, stability, translation, or regulatory function |
| RBP interaction databases | POSTAR, eCLIP repositories | RNA–protein binding region | CLIP enrichment, in vitro binding, motif preferences | Regulatory consequence of binding |
| Disease association databases | Manually curated disease regulators | RNA locus or product associated with phenotype | Expression correlation, variant association, literature support | Mechanistic causation without perturbation or functional evidence |
| Therapeutic RNA databases | Clinical trial registries, drug databases | Therapeutic RNA molecule | Sequence, chemistry, clinical status | Complete delivery, manufacturing, or regulatory attributes in one record |
Rfam is a curated database of RNA families. A typical Rfam entry contains a family accession, a name, a curated seed alignment, a consensus secondary structure, a covariance model, and annotations that help users interpret homologous sequences. The resource was developed to make RNA families searchable and comparable, and later releases expanded its coverage across many RNA classes, metagenomic sequences, viral RNAs, and microRNA families.
The key point is that Rfam is not simply a list of names. Rfam embodies a model of family membership. A covariance model can identify sequences that preserve both sequence patterns and structural constraints. This is especially useful for RNAs whose function depends on secondary structure, such as riboswitches, small nucleolar RNAs, signal recognition particle RNA, and many viral RNA elements.
The evidence boundary is equally important. An Rfam hit supports the statement that a sequence matches a family model under defined thresholds. It does not prove that the RNA is transcribed in the sampled organism. It does not prove that the RNA is processed into a mature product. It does not prove that the RNA performs the canonical function of the family in every biological context. It also does not always define the exact boundaries of a mature RNA. Those claims require expression evidence, processing evidence, functional experiments, or class-specific annotation.
miRBase illustrates a different database problem: one RNA locus can produce several named entities. A microRNA gene is transcribed into a primary transcript, processed into a hairpin precursor, and then processed into mature microRNA products. Mature products can come from the 5′ arm or the 3′ arm of the hairpin. A mature microRNA can also be grouped by seed sequence, which is important for target recognition but is not the same as the precursor hairpin or locus.
miRBase records connect microRNA names, hairpin precursors, mature sequences, and nomenclature history. This makes miRBase essential for interpreting microRNA literature. It also makes miRBase a good lesson in entity levels. A statement that “miR-21 is up-regulated” may need clarification: did the assay measure a mature arm, a hairpin precursor, a primary transcript, or a family-level signal? A statement that a microRNA “targets” an mRNA is still another type of claim, requiring target-site prediction, reporter assays, Argonaute binding evidence, perturbation evidence, or other functional support.
let-7 is a useful running example. “let-7” can refer to a founding microRNA family, a set of related genes in a species, a mature product such as let-7-5p, a less abundant arm such as a 3′ product, a seed family, or a regulatory program. A database record that resolves one of these levels should not be used as if it resolves them all.
RNA modifications add chemical identity to sequence identity. A database that stores RNA modifications must represent modified nucleosides, chemical structures, biosynthetic pathways, enzymes, and sometimes organism-specific distributions. MODOMICS catalogs RNA modification pathways and related molecular information, with recent updates extending modification, enzyme, and pathway coverage. The earlier RNA Modification Database helped establish modified nucleosides as searchable database records.
The basic database distinction is between a chemical modification and a site-specific biological claim. N6-methyladenosine, m6A, is a chemical entity: an adenosine with a methyl group at the N6 position. A MODOMICS-style record can connect that chemical entity to enzymes, pathways, and RNA classes. A site-specific claim is narrower: it states that a particular adenosine in a particular transcript is methylated in a particular cell type, developmental stage, stress condition, or disease state. That claim requires mapping evidence and must carry the limitations of the mapping method.
This distinction applies beyond m6A. A tRNA modification record describes a chemical residue and pathway; it does not prove that every tRNA of a given isotype carries the same modification in every organism. A ribosomal RNA modification record may depend on species, organelle, developmental condition, and detection method. Modification databases are therefore powerful anchors for chemical terminology, but they should not be treated as universal site maps unless they explicitly store site-level evidence.
RNAcentral, NONCODE, long noncoding RNA databases, and circular RNA resources address the problem of scattered RNA sequence and annotation records. RNAcentral is designed as an integrative resource for noncoding RNA sequences from expert databases. NONCODE and related long noncoding RNA resources focus on lncRNA loci, transcripts, expression, and functional annotation. circRNA resources focus on back-splice junctions, circular RNA candidates, expression contexts, and sometimes predicted functions.
These resources are necessary because many noncoding RNA classes are fragmented across genome annotations, class-specific databases, tissue studies, disease studies, and literature. A lncRNA may have one name in a genome annotation, another in a disease paper, and several transcript isoforms in different releases. A circRNA may be named by host gene, back-splice junction coordinate, or database-specific accession. A sequence integration resource can reduce fragmentation, but it cannot remove all ambiguity because different source resources may disagree on transcript boundaries, expression support, function, or evidence grade.
RNAcentral directly addresses sequence integration by aggregating noncoding RNA records from expert resources and linking them to genes and literature. NONCODEV6 and LncBook 2.0 illustrate lncRNA catalogs with explicit species, transcript, expression, and annotation scopes rather than universal lncRNA identifiers. circAtlas 3.0 and CircNet 2.0 illustrate circRNA resources that organize back-splice junctions, standardized circRNA names, species coverage, cancer context, and regulatory-network annotations. These sources make the same caution stronger rather than weaker: lncRNA and circRNA records remain release-, organism-, coordinate-, and evidence-specific.
A record in any of these resources should be read according to its scope. A family record does not prove expression. A sequence record does not prove function. A precursor record does not identify the mature product unless the mature product is separately stated. A back-splice junction supports a circRNA candidate but does not by itself prove high abundance, stability, translation, or regulatory activity. A lncRNA transcript model supports an annotated transcript, but it does not by itself prove a molecular mechanism. A modification record defines chemistry and pathways, but it does not prove a condition-specific site.
This is not skepticism for its own sake. It is the normal discipline of using databases scientifically. A database record should sharpen a question, not replace the evidence needed to answer it.
Box 19.1. A Database Record Is a Scoped Claim
- An Rfam record states that a sequence matches a curated family model; it does not prove expression, maturation, or function in the queried organism.
- A miRBase record distinguishes hairpin precursor from mature arm; using one level to describe the other conflates distinct molecules with different sequences and functions.
- A MODOMICS record defines a chemical modification and its biosynthetic pathway; it does not locate the modification in a specific transcript, condition, or cell type.
- An HGNC record provides an approved human gene symbol; it does not define every transcript isoform, mature product boundary, or nonhuman ortholog.
- A SILVA record classifies an rRNA sequence within a taxonomy; it is not a general RNA nomenclature authority.
- Before using any database record, ask: what entity does it describe, what evidence supports it, and what does the evidence not prove?
An accession locates a record within a resource namespace. Its meaning comes from the resource’s record model: an Rfam accession identifies a family record; a miRBase accession may identify a precursor hairpin or mature product; a RefSeq accession identifies a reference sequence record; an Ensembl or GENCODE identifier locates an annotation object; and a MODOMICS identifier refers to modification-related information. The same RNA can therefore be represented by several valid accessions that answer different questions.
Table 19.2. Resource Record and Accession Scopes. Database accessions identify different record types, carry different release behavior, and answer different biological questions.
| Resource record | Example | Entity level | Release behavior | Main ambiguity | Related chapter |
|---|---|---|---|---|---|
| HGNC symbol | MALAT1, MIR21 | Gene locus | Stable approved symbol; can be revised by nomenclature authority | One symbol may cover multiple transcripts, isoforms, and mature products | Chapter 18 |
| RefSeq accession | NR_002819.4 | Transcript model | Numeric version suffix incremented on model change | Different releases may model the same locus differently | Chapter 18 |
| Ensembl transcript ID | ENST00000456789.6 | Transcript model | Decimal version suffix; release-specific | May differ from RefSeq model for the same locus | Chapter 18 |
| GENCODE transcript ID | ENST00000456789.6 | Transcript model | Shares Ensembl versioning by release | Transcript boundaries can change across annotation releases | Chapter 18 |
| Rfam accession | RF00001 | RNA family | Stable accession; families can be merged or retired | Family membership does not determine locus, expression, or function | Chapter 12 |
| miRBase name or accession | hsa-mir-21, MIMAT0000076 | Precursor or mature microRNA | Names and arm annotations can change across releases | Name may refer to precursor hairpin, mature arm, or seed family | Chapter 19 |
| MODOMICS modification abbreviation | m6A, m5C | Modified nucleoside | Generally stable chemical abbreviation | Abbreviation may refer to chemistry, a site claim, or an assay signal | Chapter 19 |
| SILVA accession | AJ297673 | rRNA sequence record | Accession stable within release; records updated across releases | Scoped to rRNA; not a general noncoding RNA identifier | Chapter 19 |
| Genomic coordinate | chr1:100000–100100 (hg38) | Genomic position | Assembly- and release-specific | Loses meaning without assembly name, strand, and coordinate convention | Chapter 18 |
| Organism-database record | FlyBase, WormBase, TAIR, SGD, or PomBase accession | Taxon-specific gene, transcript, or feature | Resource-specific release and curation | Similar labels across taxa do not establish orthology or identity | Chapter 6 |
Consider a human microRNA locus. HGNC can provide a gene-level record, Ensembl or GENCODE can provide genomic and transcript annotation records, miRBase can provide precursor and mature-product records, Rfam can provide a family record, and disease resources can provide association records. None is a universal master record. Selecting the accession begins by deciding whether the task concerns a locus, transcript, precursor, mature product, family, or association.

Figure 19.1. One RNA, Many Resource Record Types. The same RNA biology can be represented by family, gene, transcript, precursor, mature-product, interaction, and disease-association records. For let-7, Rfam can supply a family record, genome resources can supply locus and transcript records, miRBase can supply precursor and mature-arm records, and specialized resources can supply target or disease-association records. The visual should emphasize resource name, record type, accession scope, and the biological question each record can answer. Nomenclature semantics are a prerequisite from Chapter 6, not the subject of the figure.
The vocabulary used to distinguish genes, transcripts, isoforms, products, families, coordinates, and aliases is developed in Chapter 6. Here those distinctions function as a prerequisite for choosing the correct database record. The practical test is not “Which identifier is best?” but “Which resource record has the scope needed for this claim?”
HGNC curates approved symbols, names, previous symbols, aliases, and cross-references for human genes. GENCODE, Ensembl, and RefSeq represent gene, transcript, and reference-sequence objects through their own curation and release models. These records overlap, but overlap is not identity. Transcript models can differ in exon boundaries, sequence, completeness, support, and version, while a gene record can encompass several transcript records.
Family and mature-product resources introduce further levels. One Rfam family can encompass many loci and organisms. One microRNA precursor can yield two mature arms, and identical mature sequences can sometimes arise from more than one locus. A family accession is suited to homology and family coverage; a mature-product accession is suited to the processed regulatory molecule. Neither can substitute for the other in an assay or mapping result.
MALAT1 illustrates the record-level distinction. A gene-level database record can anchor the human locus, an annotation database can define a transcript model on a particular assembly and release, and a disease resource can record an association. A statement about sequence-specific perturbation needs the transcript record; a statement comparing disease curation needs the association record. The shared display name does not make the records equivalent.
Database records exist within release histories. Some accessions remain stable while their annotation changes; some records add version suffixes after sequence changes; some are split, merged, retired, or redirected. A record cited without its resource and release can therefore be impossible to reconstruct precisely. The minimum information for interpreting a database-derived match is the resource name, release or access date, accession and version when present, organism, and relevant assembly or sequence context.
Version scope is especially important for sequence-dependent work. A RefSeq accession, Ensembl transcript ID, and GENCODE transcript ID may refer to overlapping models without having identical sequences or exon boundaries. A circRNA back-splice junction depends on the genome assembly and coordinate representation used by its source catalog. A lncRNA boundary may differ among annotation releases. A tRNA locus may differ between nuclear, mitochondrial, and specialized annotation resources. Chapter 18 develops transcript-model and browser versioning in depth; this chapter asks how those versions constrain cross-resource matching.
Organism databases such as FlyBase, WormBase, The Arabidopsis Information Resource, the Saccharomyces Genome Database, and PomBase combine organism-specific annotation, curation, and cross-references. Their identifiers should be interpreted within the taxon and data model they serve. A similar-looking human, mouse, fly, worm, yeast, plant, bacterial, or viral label does not establish orthology or record identity.
Names and aliases are useful candidate generators because older literature and current databases often use different labels. They are unsafe as automatic merge keys. A short label can be reused across organisms or refer to different entity levels. The resource-matching response is to use the label to retrieve candidates and then filter by taxon, molecule class, record type, sequence, coordinates, and evidence. Detailed rules for preferred names, aliases, and deprecated terminology belong to Chapter 6.
A database query should begin with the intended output. Family discovery points toward Rfam. Mature microRNA sequence lookup points toward miRBase. An integrated noncoding RNA sequence lookup may point toward RNAcentral. Exact transcript design points toward a versioned annotation or reference-sequence resource. A modification pathway question points toward MODOMICS, whereas a condition-specific modification-site question requires a site resource and method-level evidence. A disease or therapeutic question requires resources whose records expose the relevant evidence and product levels.
This order prevents a common inversion in database use: starting with whichever identifier is easiest to find and then treating its record as if it answered the biological question. The accession should follow the entity requirement, not define it after the fact.
RNA databases are increasingly specialized because RNA biology is experimentally diverse. A sequence database cannot fully represent RNA modifications. A family database cannot fully represent RNA-protein binding. A disease database cannot fully represent therapeutic chemistry. Specialized resources solve these problems by storing evidence fields appropriate to their domain. The danger is that users may read every specialized record as if it carried the same strength of evidence.
A good database user asks three questions. What entity does the record describe? What evidence supports the record? What claim does the evidence not support by itself?
RNA modification resources store chemical entities, modified nucleosides, pathways, enzymes, and sometimes organism- or RNA-class-specific information. MODOMICS is the main local source anchor for this chapter. Structural reviews of m6A and m6Am methyltransferases also show how modification biology connects enzymes, substrates, active sites, and RNA context.
The evidence scale spans several levels. At the chemistry level, a modified nucleoside can be defined by chemical structure. At the pathway level, enzymes and intermediates can be identified by biochemistry, genetics, and structural biology. At the site level, a residue in a transcript can be mapped by sequencing, mass spectrometry, antibody enrichment, mutational signatures, or other assays. At the functional level, a modification can affect splicing, export, translation, decay, localization, stress response, or immune sensing, but those functions require perturbation and interpretation.
The common error is to treat a modification word as a complete claim. “m6A” may mean the chemical modification, a database entry, a transcriptome-wide mapping signal, a site in a particular transcript, an antibody-enrichment peak, a writer or eraser pathway, or a broader regulatory topic. The database record should specify which level is being used.
RNA structure resources and structure-associated records address shape rather than only sequence. A record may describe an experimentally solved three-dimensional structure, a secondary-structure model, a probing-derived constraint set, a consensus structure in an alignment, or a computational prediction. These records should not be merged without method labels.
An experimentally solved structure depends on construct boundaries, organism, ligand, ions, protein partners, crystallization or cryo-electron microscopy conditions, resolution or model confidence, and sometimes mutations introduced for stability. A secondary-structure model depends on thermodynamic assumptions, comparative covariance, chemical probing constraints, or prediction algorithms. A transcriptome-wide probing dataset depends on cell type, reagent accessibility, reverse-transcription behavior, read depth, and analysis pipeline.
For entity resolution, a structure record should be linked to the RNA sequence, construct boundaries, source organism, accession, and biological form. The RCSB Protein Data Bank provides a broad structural-data anchor for experimentally determined biological macromolecule structures, including RNA-containing entries, but each entry still has construct, ligand, resolution, method, and assembly context that must travel with the citation. A riboswitch aptamer structure, for example, may not include the full expression platform. A ribosomal RNA fragment in a structural model may represent a processed mature molecule, not the primary transcript.
RNA-binding protein interaction databases collect evidence about proteins that bind RNAs. Such resources are valuable because RNA regulation often depends on ribonucleoprotein complexes. A review of RNA-protein interaction database resources emphasizes the diversity of data types in this area. Large-scale maps of human RNA-binding protein sequence, structure, and context preferences illustrate how motif preference and cellular binding evidence can be assembled at scale. More recent large-scale interactome mapping continues to expand the evidence base.
The key distinction is between binding and regulation. In vitro binding shows that a protein can bind an RNA sequence or structure under assay conditions. Motif discovery identifies sequence or structural preferences. CLIP-like assays identify cellular crosslinking enrichment near binding sites, but crosslinking efficiency, antibody specificity, library construction, and peak calling affect interpretation. Perturbation experiments show whether changing the protein affects RNA abundance, localization, splicing, translation, or decay. Direct mechanism requires still more evidence.
Therefore, an RBP database record should state the assay type and evidence level. “Protein X binds RNA Y” may mean purified protein binding in vitro, motif match, cellular crosslinking, co-immunoprecipitation, proximity labeling, or functional regulation. These are related but not equivalent.
Disease association resources connect molecular entities to phenotypes, diseases, or clinical contexts. The RNA field needs these resources for lncRNAs in cancer, microRNAs as biomarkers, repeat-derived RNAs in neurological disease, mitochondrial tRNA variants, RNA modification enzymes in disease, viral RNAs, and therapeutic targets. However, disease association is an evidence category, not a mechanism by itself.
An association can arise from differential expression, genetic variation, copy-number change, epigenetic regulation, single-cell cell-state differences, perturbation experiments, animal models, patient cohorts, clinical trials, or literature co-occurrence. A manually curated database can be rigorous while still storing heterogeneous evidence. FerrDb V2 is a useful example of a manually curated disease-associated regulator database, but it is not RNA-specific and should not be used as a direct source for RNA-specific disease database claims.
For RNA entity resolution, disease records should state whether the associated object is a gene, transcript, mature RNA, variant, modification site, RNA-binding protein, pathway, biomarker signature, or therapeutic target. LncRNADisease illustrates an RNA-specific disease resource that curates lncRNA-disease causality information, while FerrDb V2 illustrates a broader disease-associated regulator database whose records are useful for boundary comparison but not RNA-specific by design. A mature microRNA biomarker and its host gene are not the same disease entity. A lncRNA expression signature and a validated lncRNA mechanism are not the same evidence level.
Table 19.3. Evidence Fields for Database Records. Database records draw on a range of evidence types that differ in what they establish and what they cannot prove without additional data.
| Evidence field | What it supports | What it cannot support alone | Example method | Common overinterpretation |
|---|---|---|---|---|
| Sequence match | Presence of a related sequence in a database or genome | Expression, function, or mature product identity | BLAST alignment | Treating any match as proof of RNA production |
| Covariance model hit | Membership in a model-based RNA family | Expression, precise mature boundaries, or cellular function | Infernal cmsearch | Treating family membership as proof of expression or canonical activity |
| Curated literature | Association of a name, sequence, or entity with published evidence | Mechanism or causation without perturbation evidence | Database curator annotation | Treating curated association as experimental validation |
| Direct sequencing | Presence of an RNA sequence in a sample | Function, modification state, or interaction partner | RNA-seq, small RNA-seq | Treating detection as proof of biological role |
| CLIP enrichment | Crosslinking-enriched binding regions of an RBP in cells | Direct regulation or functional consequence | eCLIP, PAR-CLIP | Treating crosslinking peaks as proof of regulatory outcome |
| In vitro binding | Biochemical affinity between protein and RNA under assay conditions | Cellular binding or regulatory function in vivo | EMSA, filter-binding assay | Treating in vitro affinity as proof of in vivo interaction |
| Modification mapping | Candidate modified sites in a transcriptome | Site identity in every transcript version or condition | m6A-seq, BS-seq | Treating transcriptome-wide enrichment as precise single-site identification |
| Structural biology | Three-dimensional or secondary-structure model under defined conditions | Conformational state in all cellular contexts | X-ray crystallography, cryo-EM, SHAPE | Treating one structure as the only biologically relevant conformation |
| Disease association | Co-occurrence or correlation between an RNA entity and a disease phenotype | Mechanistic causation or therapeutic relevance | Differential expression, GWAS | Treating expression correlation as proof of disease mechanism |
| Clinical trial record | Formal testing of a therapeutic agent in a specific indication | Broad efficacy or mechanistic understanding of the RNA target | Phase I/II/III clinical trial | Treating trial registration as proof of therapeutic activity |
| Manual curator review | Expert judgment about record quality, evidence, and annotation | Replacement for primary experimental evidence | HGNC symbol curation, miRBase annotation | Treating curated annotation as equivalent to experimental proof |
Therapeutic RNA resources are especially demanding because an RNA therapeutic is not only a sequence. It may include chemical modifications, backbone chemistry, conjugates, delivery vehicles, formulation, target transcript, allele specificity, route of administration, pharmacokinetics, biodistribution, immune activation, clinical indication, regulatory status, and manufacturing attributes. Small interfering RNAs, antisense oligonucleotides, messenger RNA vaccines, guide RNAs, aptamers, splice-switching oligonucleotides, and RNA editors require different fields.
A therapeutic database record should therefore distinguish modality, molecule, target, chemistry, delivery system, evidence stage, and regulatory status. A target transcript identifier without versioning can be unsafe for exact sequence design. A disease indication without clinical evidence level can be misleading. A chemical modification field without site positions cannot support manufacturing or immunogenicity interpretation.
General drug resources can anchor some therapeutic records, but they do not eliminate RNA-specific entity-resolution requirements. DrugBank 5.0, for example, provides a curated drug database context for drug products, targets, and pharmacological information, yet an RNA therapeutic record still needs modality, sequence or chemistry, target transcript version, delivery system, indication, trial or approval status, and regulatory provenance before it can support design-level or clinical claims.
Cross-database mapping links records across resources. The word “mapping” sounds mechanical, but every mapping asserts a relationship. The relationship may be exact identity, same sequence, same genomic locus, same transcript model, same mature product, same family, same organism-specific ortholog, same disease association, same publication, same alias, or obsolete-to-current replacement. If the relationship type is not stated, users may treat related records as identical.

Figure 19.2. RNA Database Scope Map. Major RNA resources differ in the type of entity they describe and the evidence they store, so the same RNA biology can appear across several resources without any single record replacing the others. Family databases such as Rfam describe evolutionary and structural RNA families through covariance models and curated alignments; sequence integration portals such as RNAcentral aggregate noncoding RNA sequences across resources; microRNA registries such as miRBase separate hairpin precursors, mature products, and arm annotations; modification resources such as MODOMICS record chemical entities, pathways, and enzymes; RNA-binding protein databases collect binding evidence of heterogeneous types; disease and therapeutic databases connect molecular entities to clinical contexts; and transcript browsers and rRNA resources such as SILVA address still different entity scopes. Selecting the correct resource requires knowing which entity and evidence type is needed before querying.
One-to-one mapping is the simplest case. A database may cross-reference another resource’s accession for the same record. One-to-many mapping is common in RNA biology. One Rfam family can map to many genomic loci. One gene can produce many transcripts. One precursor microRNA can produce more than one mature arm. One tRNA isotype can correspond to many genomic tRNA genes. Many-to-one mapping is also common. Multiple historical aliases can map to one preferred symbol. Several database accessions may be merged when an annotation is revised. Uncertain mapping occurs when names, sequences, or coordinates are insufficient to choose one record.

Figure 19.3. Cross-Database Mapping Is Not Always Identity. Cross-database links between RNA resources represent relationships of different types, and conflating them produces errors in annotation and analysis. A one-to-one link may indicate exact record identity or only shared family membership. A one-to-many link arises when one Rfam family maps to many genomic loci or when one gene produces multiple transcript models. A many-to-one link arises when multiple historical aliases converge on one preferred symbol or when scattered mature arm names consolidate under one miRBase entry. Uncertain mappings arise when the same short alias names unrelated entities in different organisms or when an obsolete accession points to a partially replaced record. Each link should carry a relationship label—identity, same family, same sequence, related transcript, ortholog, or alias—rather than being treated as simple equivalence.
Sequence matching is powerful but limited. Exact mature RNA sequence can identify a product, but identical mature sequences can be produced by multiple loci. A short microRNA sequence may not distinguish paralogous precursors. A tRNA anticodon sequence may not distinguish all tRNA genes. A guide RNA spacer sequence may map to multiple genomic targets depending on mismatch rules.
Coordinate matching is also conditional. A genomic coordinate depends on assembly, chromosome naming convention, strand, and coordinate system. Back-splice junctions in circRNA catalogs are especially sensitive to annotation versions and alignment pipelines. Transcript coordinates depend on transcript model. A modification site reported as a transcript coordinate must be mapped to the correct transcript version before comparing with genome coordinates.
Family mapping is different from identity mapping. An Rfam family accession can connect many homologs. A miRNA seed family can connect mature products with shared seed sequences. A snoRNA family can include paralogs with related guide functions. Family mapping is useful for search and evolutionary interpretation, but it should not be used to claim that all family members have the same expression, processing, targets, or disease role.
Record provenance identifies the source resource and release from which a candidate match came. For cross-database entity resolution, the minimum useful record includes source and target resource names, release or access date, accessions and versions, organism, assembly when relevant, relationship type, evidence used to assert the match, and confidence or unresolved status. These fields permit a reader to inspect whether a claimed identity was actually sequence identity, locus overlap, family membership, historical continuity, or another relationship.
Rfam releases show why source context matters: family models, coverage, and annotations change over time. HGNC records can change their approved symbols, aliases, and cross-references. Annotation resources can revise transcript boundaries or sequence versions. A match must therefore describe which records were compared rather than asserting that two unqualified names are equivalent.
Deprecation, splits, and merges are themselves relationship types. A retired accession may point to a replacement; one old record may split into several current records; several previous records may merge. These mappings preserve historical interpretability but do not imply that every old and new record has identical boundaries. Entity resolution should retain the source record, destination record, direction of replacement, and reason for uncertainty.
This chapter uses release and source fields only to evaluate biological record matching. The practical systems for capturing computational provenance, pinning software and data environments, validating transformations, depositing datasets, recording licenses and reuse permissions, citing databases and software, monitoring resource drift, and maintaining workflows belong to Chapter 144. FAIR principles and persistent-identifier guidance provide an operational framework for those tasks, but they are not a substitute for the resource-specific biological comparison developed here.
A useful database is one whose intended coverage matches the question. Coverage includes organism or taxonomic range, RNA class, entity level, evidence class, historical depth, and update state. Rfam emphasizes model-defined families; miRBase emphasizes microRNA precursor and mature-product records; MODOMICS emphasizes modifications, pathways, and enzymes; RNAcentral integrates noncoding RNA sequence records; NONCODE and LncBook emphasize lncRNA catalogs; circAtlas and CircNet emphasize circRNA junctions and associated annotations. A missing record means little unless the resource claims to cover that entity type and taxon.
Selection should therefore be question-first. A family search favors resources with explicit homology models and family boundaries. Exact transcript design favors a versioned sequence or annotation record. A mature microRNA assay favors a product-level record. A cross-species rRNA classification task favors a resource such as SILVA. An RBP mechanism question needs records that expose assay type rather than only a reported pair. A disease question needs curation that distinguishes association from causal evidence.
Databases combine manual curation, automated import, expert-resource integration, prediction, and community submission in different proportions. Manual review can interpret literature and flag boundary cases, but it cannot eliminate curator judgment or source error. Automated integration scales across many records, but inherited cross-references can propagate stale or overbroad mappings. Computational prediction expands coverage, but predicted family membership, transcript boundaries, back-splice junctions, binding sites, or modification sites are not equivalent to experimental validation.
The database paper and record documentation should reveal which evidence enters the resource, which thresholds are applied, what curators review, and what fields are imported from partners. Two resources that both list lncRNAs may differ substantially because one prioritizes transcript catalog breadth while another integrates expression or disease annotations. Two circRNA resources may differ in species, assemblies, detection pipelines, junction filters, and network inference. Resource comparison must examine those differences rather than ranking databases by record count alone.
Entity resolution begins with a mention, identifier, sequence, coordinate, or user query. The workflow should proceed causally rather than jump directly to one answer.
First, classify the input. Is it a name, accession, sequence, coordinate, chemical abbreviation, gene symbol, transcript ID, mature RNA name, family name, disease term, or therapeutic product name? A string such as “let-7” is a name, not a complete entity. A string such as “m6A” is a chemical abbreviation or field label depending on context. A coordinate is assembly-dependent.

Figure 19.5. Ambiguity-Resolution Workflow for RNA Mentions. Entity resolution begins with a mention, accession, sequence, or coordinate and proceeds through a defined comparison rather than jumping to a single answer. The first step classifies the input by the resource record type it may denote. The second step generates candidate resources and namespaces—HGNC for human gene records, miRBase for microRNA precursor and mature-product records, Rfam for family records, MODOMICS for modification records, and SILVA for rRNA sequence records. The third step applies organism, molecule-class, sequence, coordinate, and evidence filters. The fourth step ranks candidates and preserves unresolved alternatives. The output is a source-qualified mapping that names the records, releases, relationship type, evidence, confidence, and remaining ambiguity. Operational implementation is handed to Chapter 144.
Second, generate candidate namespaces. A human gene symbol should be checked against HGNC. A microRNA name should be checked against miRBase and relevant genome annotations. A structured RNA family name should be checked against Rfam. A modification abbreviation should be checked against MODOMICS or another modification authority. A ribosomal RNA sequence may require SILVA or related rRNA resources. A transcript accession may require RefSeq, Ensembl, or GENCODE.
Third, apply context filters. Organism is often the strongest filter. Molecule class is next. A name found in a microRNA resource and a protein-coding gene resource may require context from the sentence, assay, organism, or dataset. Sequence can resolve some cases but not all. Coordinates can resolve some cases if assembly and transcript version are known. Evidence fields can separate predicted association from validated mechanism.
Fourth, rank candidates and preserve alternatives. If the context supports one record, record the match and provenance. If several candidates remain plausible, store them as unresolved alternatives with reasons. If two records are related but not identical, label the relationship. “Same mature sequence from different loci” is more informative than forcing one locus. “Same family, different product” is more accurate than claiming identity.
Fifth, report the result as a source-qualified mapping. The output should state the selected record or candidate set, source databases, releases, identifiers, relationship type, evidence, confidence, and unresolved ambiguity. This makes the biological inference inspectable. The operational implementation and maintenance of such mappings belong to Chapter 144.
Alias collision occurs when the same alias names different entities. Organism mismatch occurs when a name from one species is applied to another without orthology evidence. Precursor/mature collapse occurs when a microRNA hairpin, primary transcript, and mature arm are treated as one object. Family/locus collapse occurs when an Rfam family or miRNA family is treated as one genomic locus. Transcript/gene collapse occurs when a gene symbol is treated as a transcript ID. Coordinate drift occurs when assembly or transcript versions are ignored. Modification/site confusion occurs when a chemical modification is treated as a proven site. Association/mechanism confusion occurs when disease or RBP association is treated as functional causation.
Table 19.5. Entity-Resolution Failure Modes. Common failure modes in RNA entity resolution arise when identifier type, entity level, organism, evidence type, or relationship type are not recorded or are conflated.
| Failure mode | Example | Why it misleads | Control | Related concept |
|---|---|---|---|---|
| Alias collision | “let-7” used for family, precursor, and mature arm | Same string maps to non-equivalent entities with different sequences and functions | Store organism, molecule class, and entity level with each alias | Aliases as search terms, not merge keys |
| Organism mismatch | Mouse miR-21 treated as equivalent to human miR-21 | Ortholog sequences, expression patterns, and targets can differ | Record taxon identifier and require orthology evidence before equating | Ortholog vs. paralog distinction |
| Precursor/mature collapse | miRNA hairpin accession used in place of mature arm identifier | Sequence, length, function, and regulatory targets differ between forms | Store and report entity level: precursor vs. mature product | miRNA processing hierarchy |
| Family/locus collapse | Rfam family accession treated as one genomic locus | Family encompasses many loci across organisms | Query locus-specific records separately from family records | RNA family record scope |
| Transcript/gene collapse | Gene symbol used as if it identifies a unique transcript sequence | One gene may have many isoforms with different sequences and exon structures | Retrieve transcript ID and version, not only gene symbol | Transcript model versioning |
| Coordinate assembly mismatch | circRNA junction coordinates from hg19 compared with hg38 data | Coordinates shift between assemblies; junctions may not lift over cleanly | Record assembly name, strand, and coordinate convention | Genome assembly versioning |
| Obsolete accession | Deprecated Ensembl transcript ID used in a downstream pipeline | Record may have been merged, split, or retired without notice | Check deprecation status and follow replacement links | Database update cycles |
| Modification/site confusion | “m6A” treated as a proven site in every adenosine of a transcript | Chemical modification name does not specify transcript, position, or condition | Record transcript version, method, position, and condition for site claims | Chemical entity vs. site-specific claim |
| Association/mechanism confusion | Disease-associated lncRNA treated as a causal driver | Association evidence does not establish biological mechanism | Record evidence type and distinguish correlation from perturbation result | Evidence levels in disease databases |
Each failure mode has a practical control. Store organism and taxon. Store molecule class. Store entity level. Store sequence and coordinate versions. Store evidence type. Store relationship type. Store deprecation status. Store unresolved alternatives. These controls are simple, but they prevent many downstream errors in annotation, literature review, assay design, and computational analysis.
Suppose a user asks for “let-7 targets in human cancer.” The string “let-7” first generates candidate entities: a microRNA family, human let-7 gene loci, precursor hairpins, mature 5′ products, mature 3′ products, seed-family groupings, and target relationships. The phrase “targets” suggests mature microRNA-mediated regulation rather than precursor transcription. The phrase “human cancer” suggests Homo sapiens and disease context, but it does not identify a particular mature product or cancer type.
Box 19.2. let-7 Is Not One Entity
- Family: the let-7 microRNA family groups related genes across animals with conserved sequences and seed regions.
- Locus: in humans, multiple distinct let-7 gene loci produce independent primary transcripts.
- Precursor: each locus produces a hairpin precursor, such as hsa-mir-let-7a-1, with its own miRBase accession and sequence.
- Mature 5′ arm: the dominant mature product, let-7a-5p, is processed from the 5′ side of the hairpin.
- Mature 3′ arm: a less abundant arm, let-7a-3p, comes from the opposite hairpin side and has different targets.
- Seed class: microRNAs grouped by shared seed sequence may share target recognition without sharing a locus, precursor, or arm designation.
- Target evidence: claims about let-7 targets require specifying which mature product, which target assay, and which organism.
- Using “let-7” without qualification can refer to any of these levels and causes incorrect database queries, merged assay results, and overstated biological conclusions.
A careful workflow would retrieve human mature let-7 products from miRBase or another curated microRNA resource, keep precursor and mature records separate, record seed-family relationships, and then query target databases or literature with evidence filters. It would not treat the family name as one molecule. It would also distinguish predicted targets from Argonaute binding evidence, reporter validation, perturbation evidence, and disease-context evidence.
Suppose a paper states that “m6A is increased after stress.” The mention could refer to global m6A abundance, antibody-enrichment signal, a set of transcript sites, methyltransferase activity, a reader-protein pathway, or a field-level interpretation. Entity resolution begins by identifying m6A as a modification chemical entity and then examining the assay. If the evidence comes from liquid chromatography and mass spectrometry, the claim may concern global modified nucleoside abundance. If the evidence comes from antibody enrichment, the claim may concern relative enrichment peaks. If the evidence comes from a site-resolved method, the claim may concern specific sites, but method-specific false positives and resolution still matter.
Box 19.3. m6A Is a Chemical Name, a Site Claim, and a Field
- Chemical entity: N6-methyladenosine (m6A) is a modified adenosine with a methyl group at the N6 position; this definition belongs to chemistry and modification resources such as MODOMICS.
- Transcript site: an m6A site claim states that a specific adenosine in a specific transcript is methylated under a specific condition; this requires site-resolved mapping evidence with method and confidence fields.
- Mapping signal: antibody-enrichment methods detect enrichment peaks over a window, not single-nucleotide positions; signal width and specificity depend on antibody, library construction, and analysis pipeline.
- Writer pathway: METTL3/METTL14 and related enzymes define a biosynthetic pathway; pathway records do not prove which transcripts carry the modification under any given condition.
- Field label: “m6A research” refers to a broad area of study, not a single database entity.
- A database record should specify which level of m6A is being described and carry the appropriate method and evidence fields.
The database record should therefore separate chemical identity, site list, assay method, transcript version, condition, and evidence confidence. A MODOMICS modification entry supports the chemical and pathway vocabulary; it does not substitute for the site-specific experiment.
“tRNA-Leu” can refer to all leucine isoacceptor tRNAs, a specific anticodon group, one genomic tRNA gene, a mature tRNA molecule, a mitochondrial tRNA, a modified tRNA species, or a disease-associated mitochondrial gene. The context determines the entity. A codon-decoding discussion may need isotype and anticodon. A disease discussion may need mitochondrial gene identity and variant. A mass-spectrometry discussion may need mature molecule and modification state. A genome annotation discussion may need locus, assembly, and predicted tRNA boundaries.
The lesson is that biological vocabulary often names categories first. Database curation must then choose the exact level required for the task.
Database records are only as strong as their evidence and curation rules. Evidence can enter an RNA database through manual expert review, computational prediction, sequence alignment, covariance models, literature curation, high-throughput sequencing, structural biology, biochemical assays, genetic perturbation, clinical observation, or integration from another resource.
Manual curation is valuable because curators can interpret literature, detect synonym problems, apply nomenclature standards, and record caveats. Manual curation is not infallible; curators can inherit source errors or make judgments that later evidence changes. Computational prediction is valuable because it scales to genomes and transcriptomes. Prediction is not equivalent to validation. A predicted structured RNA family hit, lncRNA transcript, circRNA junction, RBP motif, or modification site should carry method and confidence.
High-throughput assays require artifact control. Small RNA sequencing can reveal mature microRNAs, but ligation biases and mapping ambiguity affect quantification. Long-read sequencing can improve transcript models, but coverage and error profiles matter. CLIP-like assays can locate RBP binding regions, but crosslinking, antibody specificity, and peak calling affect interpretation. Modification mapping can identify candidate sites, but chemistry, antibodies, reverse-transcription signatures, and statistical pipelines affect false positives. Disease association studies can reveal correlations, but mechanism requires additional perturbation or functional evidence.
The database user’s task is not to reject high-throughput evidence. It is to use the correct verb. A record may “predict,” “annotate,” “match,” “detect,” “associate,” “curate,” “validate,” or “demonstrate” a claim. These verbs should not be collapsed.
Entity-resolution problems differ across RNA classes. Ribosomal RNA records often connect sequence, taxonomy, and mature rRNA processing. SILVA is a major rRNA sequence and alignment resource for classification, but it should not be treated as a general authority for all RNA names. Transfer RNA records require isotype, anticodon, genomic locus, mature sequence, intron status, organelle origin, and modifications. Small nuclear and small nucleolar RNAs often have gene families, paralogs, guide relationships, and processing products. MicroRNAs require precursor and mature-arm distinctions. Long noncoding RNAs require transcript model, locus, strand, expression evidence, and functional evidence. Circular RNAs require back-splice junction, host gene context, assembly, abundance, and validation.
Organism context matters just as much. Human gene nomenclature has HGNC authority. Model organisms have their own nomenclature authorities. Bacterial and archaeal small RNAs may be named locally by discovery papers or organism-specific databases. Viral RNA elements may be strain-specific and may change as viral taxonomy changes. Mitochondrial and chloroplast RNAs have organelle-specific gene names and processing systems.
Developmental stage, cell type, disease state, and experimental condition can also define the relevant entity. A transcript detected in one tissue may not be expressed elsewhere. A modification site may be condition-dependent. An RBP interaction may occur only in a compartment or stress state. A therapeutic target may depend on disease allele, splice isoform, or delivery tissue. Database records should not be read outside the context their evidence supports.
RNA database reasoning has practical consequences. In genome annotation, incorrect mapping can create duplicated or missing RNA genes. In RNA-seq analysis, outdated transcript IDs can misassign reads. In single-cell analysis, gene-symbol aliases can merge unrelated features or split the same feature into multiple rows. In CRISPR and RNA-guided technologies, a guide RNA or target transcript must be resolved to exact sequence and genome assembly. In antisense or siRNA design, transcript version and variant context can determine whether a reagent hits the intended target. In mRNA therapeutic design, sequence, codon usage, untranslated regions, modifications, and delivery context all require precise identifiers.
Clinical interpretation raises the stakes. A mitochondrial tRNA variant must be mapped to the correct gene and variant nomenclature. A microRNA biomarker must specify the mature product and assay. An RNA therapeutic record must separate drug product, target, chemistry, delivery, indication, and regulatory status. A disease association should not be promoted to a mechanism without evidence. These are not clerical details; they affect diagnosis, trial interpretation, and experimental reproducibility.
Computational analyses should treat record matching as a scientific inference. Stripping version suffixes, ignoring organism, or collapsing aliases may increase the nominal match rate while decreasing biological truth. A defensible comparison records match confidence, unresolved alternatives, source releases, and relationship types. Text-mining systems such as PubTator Central illustrate the need to normalize extracted mentions to candidate records, but the semantic definitions of ontology classes and identifier layers belong to Chapter 6, while execution, deposition, licensing, FAIR implementation, and maintenance belong to Chapter 144.
Box 19.4. Cross-References Need Relationship Labels
Cross-references between databases can mean very different things, and unlabeled links are a source of systematic annotation error. The following relationship types should be distinguished:
- Identity: the two records describe the same object with the same boundaries and evidence.
- Same family: the records belong to the same homology group but may differ in organism, locus, or processing state.
- Same sequence: the records share a sequence but may originate from different loci or carry different annotations.
- Related transcript: the records describe overlapping but not identical transcript models.
- Ortholog: the records describe functionally related genes in different organisms, with orthology evidence required.
- Historical alias: one record uses a name that a different resource has updated, retired, or reassigned.
- Obsolete replacement: one accession has been deprecated and the cross-reference points to its current successor.
Storing relationship type with every cross-reference prevents users from treating related records as identical.
Box 19.5. Entity Resolution Should Say “Unresolved” When Needed
Entity resolution workflows should preserve ambiguity rather than force a one-to-one assignment when evidence is insufficient. Forcing a weak mapping produces false certainty and can propagate errors through downstream databases, analyses, and publications. A well-designed entity record should carry an explicit unresolved state with reasons—for example: “two candidate loci remain after organism and class filtering; arm annotation is ambiguous without assay context.” Preserving unresolved alternatives allows later evidence to close the gap correctly and signals to users that an additional experimental step or contextual check is needed before the name can be used with confidence. An unresolved record is more accurate than a wrongly resolved one.
The current practical consensus is that RNA databases must be interpreted through resource scope. Rfam records support model-based family membership, not automatic expression or function. miRBase records require careful separation of precursor hairpins, mature products, arm annotations, and functional target claims. MODOMICS and related modification resources define chemical and pathway knowledge, but site-specific modification claims require separate evidence. HGNC provides essential human gene-symbol curation, but gene symbols remain distinct from transcripts, mature products, and nonhuman orthologs.
There is also broad agreement in practice that database-derived claims must identify their source records. Release, version, accession, evidence field, and access date can affect interpretation. Cross-references require relationship labels. Deprecated accessions remain useful for reconstructing older studies but should not be treated as current records. Ambiguity should be carried explicitly when context is insufficient.
The unsettled areas are not whether these principles matter, but how uniformly databases implement them. Different resources have different coverage, data models, curation mixtures, evidence codes, and release practices. Some fast-moving RNA classes, especially lncRNAs, circRNAs, RNA modification sites, RBP interactions, and therapeutic RNAs, still have uneven evidence standards and fragmented identifiers.
Open questions:
Common misconceptions: