This chapter explains what it means to describe the RNA gene repertoire of a genome, organism, organelle, virus, mobile element, cell type, or environmental sample. An RNA gene repertoire is not a single universal count. It is a biologically scoped inventory of loci, genetic elements, and RNA products that must state whether it is counting protein-coding genes, noncoding RNA genes, precursor transcripts, mature RNA molecules, processed fragments, repeat-derived products, viral genomes, organellar genes, or database entries.
Chapter 14 is a bridge between the evolutionary chapters before it and the annotation chapters after it. Chapter 12 introduced homology, covariance, and comparative inference. Chapter 13 introduced mobile elements and RNA-linked genome innovation. Chapter 15 will treat promoters, terminators, operons, and transcription units as RNA-production architecture. Chapters 16 through 19 will then expand repeats, organelles, transcript models, genome browsers, databases, and identifier systems. The purpose here is to give the reader a durable conceptual map: which RNA classes exist across life, why repertoires differ among biological systems, why copy number and fragments complicate counting, and why absence from a catalog is not the same as biological absence.
An RNA gene repertoire is the set of RNA-producing genetic features that a specified biological system carries or expresses. The word “specified” is essential. The RNA repertoire of a complete bacterial genome, a fragmented metagenomic contig, a human nuclear genome, a mitochondrion, a plastid, an RNA virus, and a retrotransposon are different kinds of inventories. A useful repertoire claim states the biological context, the annotation release or assembly, the counted object, and the evidence threshold. It should be clear whether the count refers to genomic loci, transcription units, precursor RNAs, mature RNAs, processed fragments, functional motifs, viral RNA genomes, or database records.
Protein-coding genes belong in an RNA repertoire because their immediate products are RNAs. A messenger RNA is not merely a temporary intermediate between DNA and protein; it has ends, untranslated regions, splice isoforms, chemical modifications, localization signals, RNA-binding proteins, decay routes, and translation states. Noncoding RNA genes are genes whose primary functional products are RNAs rather than proteins. “Noncoding” is therefore a statement about protein-coding capacity, not a proof of function, mechanism, or importance. Many noncoding RNA classes are ancient and essential, including rRNAs, tRNAs, RNase P RNA, and spliceosomal RNAs. Other noncoding categories, such as lncRNAs and circRNAs, contain transcripts with diverse mechanisms and uneven levels of functional evidence.
RNA repertoires vary across life. Bacterial genomes are compact but include mRNAs, rRNAs, tRNAs, cis-regulatory leaders, riboswitches, small regulatory RNAs, antisense RNAs, CRISPR arrays, toxin-antitoxin RNAs, phage RNAs, and mobile-element products. Archaeal genomes combine bacterial-like genome organization with transcription and RNA-processing features related to eukaryotes, and they carry guide RNAs, modification RNAs, CRISPR RNAs, tRNAs, rRNAs, and other stable RNP components. Eukaryotic genomes add introns, spliceosomal RNAs, snoRNAs, miRNAs, piRNAs in many animal germlines, lncRNAs, circRNAs, enhancer-associated RNAs, promoter-proximal transcripts, and extensive repeat-derived transcription. Organelles carry reduced but specialized repertoires whose products often depend on nuclear-encoded processing factors. Viral and mobile-element RNAs cross ordinary gene boundaries: retroviral genomic RNA is genome, packaging substrate, and reverse-transcription template, and viroids are infectious circular RNAs without protein-coding capacity.
Even ancient RNA classes are not fixed inventories. tRNA copy number and anticodon repertoires vary across cellular life, demonstrating that a conserved class can have lineage-specific repertoire structure. Bacterial small RNA repertoires can vary within a species, as shown by strain-level small RNA differences in Aggregatibacter actinomycetemcomitans. Duplicated RNA genes can remain redundant, divide ancestral functions, acquire new expression patterns, or decay into pseudogenes. Pseudogene-derived and repeat-derived RNAs require special caution because sequence similarity and repetitive origin make mapping and functional interpretation difficult.
Observed repertoires are shaped by methods. Standard small RNA library preparation can miss modified RNAs because modifications interfere with adapter ligation, reverse transcription, or amplification. PANDORA-seq showed that modification-aware workflows can expand the detectable regulatory small RNA repertoire. Integrated small RNA annotation tools can organize labels, but class assignment still depends on library design, mapping strategy, modification bias, and database versions. Metadata also matters: lineage-specific or environment-specific claims require reliable information about organism, strain, tissue, developmental stage, cell type, treatment, infection state, and sampling protocol.
The main lesson of this chapter is conservative and constructive. RNA gene repertoires are broader than protein-coding gene sets, richer than transcript abundance tables, and more conditional than static genome annotations. At the same time, not every RNA read, noncoding transcript, fragment, repeat-derived molecule, or disease-associated expression change is a new functional RNA gene. A strong repertoire annotation separates detection from function, locus from product, precursor from mature RNA, repeat mapping from unique evidence, and true absence from non-detection.
The most important reader-facing terms are summarized here so that the chapter can be read independently.
The first prerequisite is the distinction between a locus, a transcript, and a mature RNA. A locus is a genomic interval. A transcript is an RNA molecule synthesized from a template. A mature RNA is the processed molecule that persists after capping, splicing, cleavage, trimming, editing, modification, export, loading into a ribonucleoprotein complex, or decay pathway selection. A single locus can produce multiple transcripts. A single transcript can produce multiple mature RNAs. A mature RNA can be further cleaved into fragments.
The second prerequisite is the distinction between detection and interpretation. RNA sequencing reads show that molecules compatible with a protocol were sampled. Reads do not automatically define gene boundaries, prove a distinct RNA class, or show biological function. A 22-nucleotide read can be a miRNA, a siRNA, a piRNA-size molecule, a tRNA fragment, a degradation product, or an ambiguously mapped repeat read. A long polyadenylated read can be a protein-coding transcript, lncRNA, readthrough product, retained intron, processed pseudogene transcript, or assembly artifact. Chapter 5 introduced evidence standards and artifact control; Chapter 14 applies those standards to gene repertoires.
The third prerequisite is that different genetic systems encode different kinds of RNA. A bacterium may organize genes in operons. A eukaryote may splice a precursor transcript and produce alternative isoforms. A mitochondrion may encode a compact set of RNAs but depend on nuclear-encoded factors for processing. A retrovirus may package RNA as genome and use reverse transcription to move genetic information into DNA. A viroid may be an infectious RNA with no protein-coding genes. The word “gene” remains useful, but each system requires its own boundary rules.

Figure 14.1. What Is an RNA Gene Repertoire? An RNA gene repertoire is not a single count. A genome can be counted by loci, primary transcripts, processed products, RNA elements, or annotation records, and different counts answer different biological questions. The figure separates genomic loci, primary transcripts, mature RNAs, processed fragments, RNA elements, and database entries so that readers can see why the same genome yields several valid but different repertoire counts.
Table 14.1. Major RNA Gene and RNA Product Classes. Overview of major RNA gene and RNA product classes with their products, biogenesis, distribution, evidence, and the boundary problems that complicate counting.
| Class | Main product | Typical biogenesis | Organism distribution | Main evidence | Common boundary problem |
|---|---|---|---|---|---|
| mRNA | Translatable messenger RNA | Transcription, capping, splicing, polyadenylation | All cellular life | RNA-seq, ribosome profiling | Coding vs noncoding isoforms, uORFs |
| rRNA | Ribosomal RNA | Precursor cleavage and modification | All cellular life and organelles | Conserved structure, homology | rDNA array copy number, collapse |
| tRNA | Mature transfer RNA | Leader/trailer removal, CCA addition, modification | All cellular life and organelles | Conserved sequence and structure | Locus vs anticodon vs mature pool vs fragment |
| snRNA | Spliceosomal small nuclear RNA | Pol II/III transcription, snRNP assembly | Eukaryotes | Conserved sequence, snRNP association | Multiple paralogs and variants |
| snoRNA | Modification guide RNA | Often intron-encoded, box C/D or H/ACA | Eukaryotes and archaea | Conserved boxes, target complementarity | Host-intron overlap, snoRNA-derived fragments |
| miRNA | Mature microRNA | Hairpin precursor processed by Drosha/Dicer | Animals and plants | Hairpin plus mature duplex, Argonaute loading | Distinguishing from other small reads |
| siRNA | Small interfering RNA | Dicer processing of double-stranded RNA | Many eukaryotes | Size, pathway signatures | Confusion with miRNA or degradation |
| piRNA | PIWI-interacting RNA | PIWI loading, ping-pong amplification | Animal germlines | PIWI association, cluster origin | Cluster annotation, somatic absence |
| Bacterial sRNA | Small regulatory RNA | Independent transcription or processing | Bacteria | Strand-specific reads, chaperone association | Antisense vs readthrough, strain variation |
| CRISPR RNA | crRNA guide | CRISPR array transcription and processing | Bacteria and archaea | Repeat-spacer arrays, effector context | Array vs spacer counting |
| lncRNA | Long noncoding transcript | Pol II transcription, often spliced/polyadenylated | Eukaryotes | Transcript boundaries, low coding potential | Functional vs readthrough or promoter byproduct |
| circRNA | Circular RNA | Back-splicing or circularization | Eukaryotes | Back-splice junction, exonuclease resistance | True circle vs artifact |
| Organellar RNA | Mitochondrial/plastid RNAs | Polycistronic transcription, processing, editing | Eukaryotic organelles | Organellar contigs, processing evidence | Nuclear-dependent maturation, assembly omission |
| Viral RNA | Genomic or subgenomic viral RNA | Replication by viral polymerase | RNA viruses, retroviruses | Viral assembly, sequencing | Genome vs mRNA vs replication intermediate |
| Viroid RNA | Infectious circular RNA | Host-polymerase rolling-circle replication | Plant viroids | Circular RNA, infectivity | Not a host gene, boundary case |
| Repeat-derived RNA | Repeat or transposon transcript | Transcription from repeats or mobile elements | Across life | Mapped reads, processing into small RNAs | Family vs locus assignment, mapping ambiguity |
Table 14.2. Repertoire Differences Across Genetic Systems. Comparison of the RNA repertoires carried by major genetic systems and the annotation hazards specific to each.
| System | Core RNA classes | Specialized RNA classes | Mobile or viral components | Main annotation hazards |
|---|---|---|---|---|
| Bacteria | mRNA, rRNA, tRNA | sRNAs, riboswitches, tmRNA, RNase P RNA | Prophage, plasmid, transposon RNAs | Short genes, antisense vs readthrough, strain variation |
| Archaea | mRNA, rRNA, tRNA | Modification guide RNAs, CRISPR RNAs, sRNAs | CRISPR arrays, mobile elements | Undersampling, lineage-specific sRNAs |
| Animals | mRNA, rRNA, tRNA | snRNA, snoRNA, miRNA, piRNA, lncRNA, circRNA | Retroelement and endogenous viral RNAs | Isoform complexity, repeat reads, cell-type specificity |
| Plants | mRNA, rRNA, tRNA | miRNA, siRNA, snoRNA, lncRNA | Transposon RNAs, viroids and viruses | Large repeat content, organelle mixing |
| Fungi | mRNA, rRNA, tRNA | snoRNA, some siRNA and lncRNA | Retrotransposon RNAs, mycoviruses | Lineage variation, RNAi loss in some species |
| Mitochondria | Reduced mRNA, rRNA, tRNA | Edited and processed transcripts | Rare | Nuclear-dependent processing, assembly omission |
| Plastids | mRNA, rRNA, tRNA | Edited, polycistronic transcripts | Rare | Editing sites, nuclear factor dependence |
| RNA viruses | Genomic RNA | Subgenomic RNAs, structured regulatory regions | The virus itself | Genome vs mRNA vs replication intermediate |
| Retroviruses | Genomic RNA (also template) | Spliced viral mRNAs, packaging signals | Endogenous retroviruses | Genome vs template distinction, ERV mapping |
| Viroids | Circular genomic RNA, no protein | Structured rod-like RNA | Infectious agent itself | Not a host gene, host-read contamination |
| Mobile elements | Element transcript | Source of small RNAs (piRNA, siRNA) | Transposition intermediates | Mapping ambiguity, active vs relic |
Table 14.3. Evidence Required for RNA Repertoire Claims. Evidence thresholds, common artifacts, and related chapters for the main kinds of RNA repertoire claims.
| Claim | Minimum evidence | Stronger evidence | Common artifact | Related chapter |
|---|---|---|---|---|
| Gene exists | Genomic locus with class features | Conserved sequence or structure, expression | Assembly collapse or contamination | Chapter 12 |
| RNA is expressed | Reproducible reads from appropriate library | Strand specificity, end mapping, independent methods | Mapping ambiguity, readthrough | Chapter 18 |
| Mature RNA is produced | Processed product detected | Defined ends, processing signatures, modification | Degradation products | Chapter 18 |
| Fragment is regulated | Reproducible fragment with cleavage pattern | Protein association, condition dependence, perturbation | Degradation or library bias | Chapter 16 |
| Class is absent | Sought but not found under stated methods | Method coverage appropriate for the class | Non-detection mistaken for absence | Chapter 5 |
| Pseudogene is functional | Locus-specific expression | Parent-mapping exclusion, perturbation and rescue | Paralogous mapping, readthrough | Chapter 16 |
| Lineage-specific RNA is conserved | Defined comparison set and search method | Covariance or structure, synteny conservation | Undersampling, sequence divergence | Chapter 62 |
A protein-coding gene is often introduced as a DNA segment that encodes a protein. For RNA biology, that definition is incomplete because the first stable product of the gene is an RNA. In bacteria, many mRNAs are transcribed from operons, where one promoter can produce a polycistronic transcript containing multiple protein-coding regions. In eukaryotes, many protein-coding genes produce precursor mRNAs that are capped, spliced, cleaved, polyadenylated, exported, translated, localized, and degraded. The mRNA carries regulatory information in its 5′ untranslated region, coding sequence, 3′ untranslated region, introns, splice junctions, RNA structure, chemical modifications, and RNA-binding protein sites.
This matters for repertoire work because excluding protein-coding genes from an RNA repertoire creates a false division. A catalog of all RNA-producing features must include mRNA genes while still distinguishing the protein-coding function from RNA-level regulation. For example, a stress-response mRNA may be counted as a protein-coding gene in a genome annotation, as a transcript in an RNA-seq experiment, as a translated molecule in ribosome profiling, and as a regulatory target in a small RNA study. Those are different layers of the same gene’s RNA biology.
Protein-coding categories also have boundary cases. Some annotated lncRNAs contain small open reading frames that encode microproteins. Some protein-coding loci produce noncoding isoforms. Some viral RNAs serve as mRNA and genome. Some upstream open reading frames regulate translation rather than define the main protein product. Chapter 77 treats noncanonical coding potential in detail. In Chapter 14, the key rule is that coding potential and RNA product identity should be stated separately.
A noncoding RNA gene is a gene whose primary functional product is an RNA rather than a protein. The term is useful but potentially misleading. “Noncoding” says what the RNA is not: it is not primarily translated into a protein. It does not say what the RNA does. Some noncoding RNAs are among the most ancient and essential molecules in cells. rRNAs form the catalytic and structural core of ribosomes. tRNAs decode codons by carrying amino acids to the ribosome. RNase P RNA participates in tRNA 5′ end maturation in many organisms. Spliceosomal snRNAs recognize splice sites and participate in pre-mRNA splicing chemistry.
Other noncoding RNAs are regulatory. Bacterial sRNAs can pair with mRNAs or bind proteins to affect translation and stability. miRNAs guide Argonaute proteins to target mRNAs in animals and plants. siRNAs guide RNA interference pathways. piRNAs help protect many animal germlines from transposable elements. CRISPR RNAs guide CRISPR-Cas immune complexes. Riboswitches fold into ligand-binding structures that control transcription, translation, splicing, or RNA stability. Some lncRNAs recruit, scaffold, titrate, or organize proteins and nucleic acids. Some circRNAs can bind proteins or miRNAs, and a minority may be translated.
The evidence standard differs by class. A conserved tRNA can be annotated from sequence and structure with high confidence, although mature expression and modification remain separate questions. A miRNA requires evidence for a hairpin precursor and characteristic processing, not merely a small read of the right length. A lncRNA function claim usually requires more than expression and differential abundance; mechanism claims need perturbation, localization, molecular partners, and appropriate controls. This is the main caution encoded by.
RNA classes are not all defined by the same property. Some are defined by function: tRNAs carry amino acids, rRNAs form ribosomes, CRISPR RNAs guide defense complexes. Some are defined by biogenesis: miRNAs are processed from hairpin precursors, piRNAs are loaded into PIWI proteins through specialized pathways, circRNAs are produced by back-splicing or other circularization routes. Some are defined by structure: riboswitch aptamer domains and many ribozymes have conserved folds. Some are defined by genomic or transcript features: lncRNAs are long transcripts with little or no protein-coding potential, and enhancer RNAs are associated with enhancer transcription.
Because class labels use different axes, a single RNA can sit near multiple categories. A viral RNA may be a genome, mRNA, regulatory structure, and packaging signal. A repeat-derived transcript may be a lncRNA by length and coding potential, a mobile-element RNA by origin, and a source of small RNAs by processing. A tRNA fragment may be a product of a housekeeping tRNA gene but behave as a stress-regulated small RNA in a particular context. A careful repertoire table states both origin and product class rather than forcing a single label.
Box 14.1. Noncoding Does Not Mean Nonfunctional
- A noncoding RNA is not translated as its primary product; the term describes coding potential, not function or importance.
- The definition includes essential structural RNAs, guide RNAs, catalytic RNAs, regulatory RNAs, and many transcripts whose functions are unknown.
- The caution has two sides: do not dismiss a transcript because it is noncoding, and do not infer function merely because it is noncoding.
A noncoding RNA is not translated as its primary product. That definition includes essential structural RNAs, guide RNAs, catalytic RNAs, regulatory RNAs, and many transcripts whose functions are unknown. The correct caution has two sides: do not dismiss a transcript because it is noncoding, and do not infer function merely because it is noncoding.
Bacterial genomes are often smaller and more compact than eukaryotic nuclear genomes, but bacterial RNA repertoires are not simple. A typical bacterial chromosome encodes mRNAs, rRNAs, tRNAs, transfer-messenger RNA in many lineages, RNase P RNA, signal recognition particle RNA, riboswitches, attenuation leaders, antisense RNAs, small regulatory RNAs, CRISPR arrays in many species, toxin-antitoxin RNAs, prophage RNAs, plasmid RNAs, and transposon-associated RNAs. Many bacterial transcripts are polycistronic, meaning that one RNA can contain multiple protein-coding regions. Many regulatory RNAs act by base pairing with mRNAs, changing access to ribosome-binding sites or recruiting RNA decay machinery.
A bacterial small RNA, or sRNA, is usually a short RNA that regulates translation, mRNA stability, transcription, protein activity, or stress response. Some bacterial sRNAs require RNA chaperones such as Hfq or ProQ; others act through different protein partners or direct RNA pairing. The repertoire of bacterial sRNAs is shaped by species, strain, growth condition, stress exposure, infection state, and library design. A study of Aggregatibacter actinomycetemcomitans showed that small RNA repertoires can vary within a species, supporting the general claim that strain-level context matters for bacterial repertoire annotation.
Bacterial annotation hazards include short genes missed by open reading frame filters, antisense RNAs confused with readthrough, small RNAs detected only under stress, phage-derived transcripts mistaken for host genes, and repetitive mobile elements that confound mapping. Strong bacterial RNA repertoire work therefore combines genome context, strand-specific transcript evidence, transcription start and end evidence where available, conservation or covariance for structured RNAs, and perturbation for regulatory claims.
Archaea require their own category because archaeal RNA biology is neither simply bacterial nor simply eukaryotic. Archaeal genomes often have compact organization and operons, yet archaeal transcription machinery resembles eukaryotic RNA polymerase II systems more than bacterial RNA polymerase systems. Archaeal repertoires include mRNAs, rRNAs, tRNAs, RNase P components, signal recognition particle RNA, guide RNAs for RNA modification, CRISPR RNAs, antisense RNAs, and lineage-specific small RNAs. Archaeal CRISPR-Cas systems have been especially important for understanding RNA-guided defense.
The evidence logic is familiar but the details differ. Archaeal tRNAs and rRNAs can be found with conserved structure and homology. Archaeal guide RNAs and CRISPR RNAs require attention to repeat-spacer arrays, processing sites, effector proteins, and genome defense context. Archaeal small RNAs may be condition-dependent or lineage-specific, and recent archaeal RNA-processing reviews emphasize that archaeal repertoires include guide RNAs, regulatory RNAs, CRISPR RNAs, and lineage-specific processing systems rather than a simple bacterial-like small-RNA catalog.
Eukaryotic nuclear genomes contain the RNA classes shared with other cellular life plus expanded repertoires created by introns, chromatin, repeats, and specialized RNA-processing pathways. Eukaryotic cells encode mRNAs, rRNAs, tRNAs, snRNAs, snoRNAs, scaRNAs, RNase P and MRP RNAs, SRP RNA, telomerase RNA, Y RNAs, vault RNAs, miRNAs, endogenous siRNAs in some lineages, piRNAs in many animals, lncRNAs, circRNAs, enhancer-associated RNAs, promoter-associated RNAs, antisense transcripts, and repeat-derived RNAs.
Eukaryotic repertoires are difficult to count because transcription and processing are extensive. A locus may produce multiple capped starts, alternative exons, retained introns, different 3′ ends, and cell-type-specific isoforms. A long transcript may overlap a protein-coding gene on the opposite strand. A lncRNA may be a stable functional molecule, a promoter-associated byproduct, a readthrough transcript, or a transcript whose act of transcription matters more than the RNA product. A circRNA may come from exonic back-splicing, intronic circularization, or technical artifacts. Chapter 18 treats transcript models and isoform versioning, while Chapter 90 treats lncRNA classes in depth.
Eukaryotes also illustrate why RNA repertoire is not identical to expression abundance. Some RNA genes are constitutively present in the genome but expressed only in specific cell types. piRNA clusters are central in many animal germlines but not in all somatic transcriptomes. Certain lncRNAs are highly lineage-specific. Many enhancer-associated transcripts are unstable. A repertoire claim must say whether it means “encoded in the genome,” “detected in a sample,” “active in a cell type,” or “functionally validated.”
Mitochondria, chloroplasts, plastids, and related organelles carry reduced genetic systems with RNA repertoires that are small in gene count but rich in processing logic. Many organellar genomes encode a subset of mRNAs, rRNAs, and tRNAs. Some organelles import tRNAs or RNA-processing factors from the nucleus. Organellar transcripts may be polycistronic, extensively processed, edited, stabilized by RNA-binding proteins, or regulated by polyadenylation and degradation pathways. Chapter 17 gives organellar RNA genes their detailed treatment.
For Chapter 14, organelles teach two general rules. First, an organism’s RNA repertoire is not limited to the nuclear or main chromosome. A human cell has nuclear and mitochondrial RNA-producing genetic systems; a plant cell has nuclear, mitochondrial, and plastid systems. Second, gene presence and mature product availability can diverge. An organelle may encode an RNA gene whose mature product depends on nuclear-encoded proteins, imported RNA factors, editing, or processing. Plant organelle genetics, human mitochondrial RNA-processing reviews, and mitochondrial RNA-maturation syntheses support this boundary rule while Chapter 17 gives the full organellar treatment.
Viruses and mobile elements challenge ordinary gene categories because RNA can be genetic material, message, regulatory element, structural element, enzyme substrate, packaging signal, or replication intermediate. Retroviral genomic RNA is a clear example: the same RNA molecule functions as the packaged genome, as a substrate for reverse transcription, and as part of the viral life cycle’s assembly and packaging logic. Alphavirus RNA life cycles similarly require the reader to distinguish genomic RNA, subgenomic RNA, replication intermediates, and structured RNA regions.
Viroids are even more extreme. A viroid is a small infectious circular RNA that does not encode proteins. It is not a host noncoding RNA gene in the ordinary sense, but it is an RNA genetic agent. Viroids show why a repertoire chapter must include RNA entities that are outside standard host gene annotations.
Mobile elements produce RNAs with several possible identities. A retrotransposon RNA can be an intermediate for copying the element. A transposon-derived transcript can be a source of small RNAs that silence transposons. A repeat-derived long transcript can influence chromatin or trigger innate immune sensing. A mobile-element fragment can also be a mapping artifact or an inactive relic. Chapter 16 expands this topic; the present chapter establishes that repertoire annotation must record origin, evidence, and activity separately.

Figure 14.2. RNA Classes Across Biological Systems. Different genetic systems produce overlapping but distinct RNA repertoires. The same terms, such as guide RNA or genomic RNA, must be interpreted in organism and element context. The figure compares bacterial, archaeal, eukaryotic, organellar, viral, viroid, and mobile-element repertoires to emphasize that shared labels can mean different things in different systems.
Figure 14.2. RNA Classes Across Biological Systems. The planned figure [Chapter 14](chapter1013.md).fig.02 should compare bacterial, archaeal, eukaryotic, organellar, viral, viroid, and mobile-element repertoires. The figure should emphasize that the same label, such as “guide RNA” or “genomic RNA,” means different things in different biological systems.
RNA gene repertoires differ not only by class presence but also by copy number. Gene-copy variation means that genomes, strains, species, or individuals differ in how many copies of a gene or RNA gene class they carry. Copy number can influence dosage, robustness, specialization, or evolutionary experimentation. It can also reflect recent duplication, incomplete genome assembly, collapsed repeats, pseudogenization, or annotation artifacts.

Figure 14.3. Copy Number and RNA Gene Repertoire Variation. RNA gene repertoires change by copy-number variation, duplication, deletion, and decay. Counting copies without classifying activity can mix functional genes, paralogs, pseudogenes, and fragments. The figure illustrates tRNA copy-number variation, miRNA paralog duplication, pseudogenes and processed copies, repeat-derived fragments, and annotation confidence labels.
tRNAs are a useful running example because they are ancient and structurally conserved, yet they vary in copy number and anticodon representation across cellular life. An anticodon is the three-nucleotide part of a tRNA that pairs with an mRNA codon during translation. A genome’s tRNA gene repertoire can therefore be described by the number of tRNA loci, the anticodons represented, the amino acids served, the predicted tRNA isotypes, the mature tRNA pool, and the tRNA modifications needed for decoding. Those are related but not identical inventories.
Other RNA classes also vary in copy number. Bacterial rRNA operon copy number can differ among organisms with different growth strategies. Eukaryotic rDNA repeats occur in large arrays. snRNA and snoRNA families may have multiple paralogs. miRNA families can expand and lose members. CRISPR arrays vary by spacer acquisition history. piRNA clusters differ among animal lineages. A repertoire statement should therefore avoid saying simply that an organism “has tRNAs” or “has miRNAs”; it should specify which copies, families, products, and evidence level are being discussed.
A paralog is a homologous gene produced by duplication. Paralogous RNA loci can follow several paths. One copy may retain the ancestral role while another decays. Both copies may remain redundant. The copies may divide ancestral functions, a process often called subfunctionalization. One copy may acquire a new expression pattern, target, processing route, or molecular partner. In RNA repertoires, paralogy can be subtle because the product may be defined by structure, modification pattern, target recognition, or genomic context rather than protein sequence.
For example, duplicated miRNA hairpins can produce related mature miRNAs with overlapping but not identical target repertoires. Duplicated snoRNAs may guide modification at similar or distinct sites. Duplicated tRNA genes may contribute to codon-specific decoding capacity or may be silent. Paralogs should not be collapsed automatically unless the biological question is explicitly about family-level presence rather than locus-level or product-level differences.
A pseudogene is a gene-derived sequence that has lost the ancestral coding or structural capacity. Pseudogenes can arise through duplication followed by mutation, retrotransposition of a processed mRNA, or decay of a once-functional gene. In RNA repertoire work, pseudogenes matter for three reasons. First, pseudogene sequences can be transcribed. Second, pseudogene similarity to functional genes can confuse read mapping. Third, some pseudogene-derived RNAs may have regulatory roles, although such claims require strong controls.
Controls are essential because pseudogene transcription can be overcalled. Reads may map equally well to a parent gene and pseudogene. Apparent pseudogene expression may come from readthrough transcription, antisense transcription, repetitive sequence, or incomplete annotation. A functional pseudogene RNA claim should show reproducible expression from the pseudogene locus, correct strand and boundaries, exclusion of parent-gene mapping, and mechanism-specific evidence. If the claim is that a pseudogene RNA regulates another RNA, perturbation and rescue are stronger than correlation alone. This chapter records the general caution; Chapter 16 treats pseudogene and repeat-derived RNA mechanisms in detail.
Many RNA molecules are processed into smaller RNAs. A pre-tRNA becomes a mature tRNA after 5′ leader removal, 3′ trailer processing, CCA addition in many systems, intron removal in some tRNAs, and modification. A primary miRNA transcript is processed into a precursor hairpin and then into a mature miRNA duplex. rRNA precursors are cleaved and modified during ribosome biogenesis. These processing products are part of the normal RNA repertoire.
Fragments are more difficult. A tRNA-derived fragment, rRNA fragment, Y RNA fragment, snoRNA-derived RNA, or mRNA cleavage product may be regulated and functional in some contexts. It may also be a degradation product from extraction, stress, nuclease activity, or library preparation. A fragment should not be counted as a distinct RNA gene merely because it appears in sequencing data. The evidence should identify the precursor, cleavage pattern, reproducibility, condition dependence, protein association, subcellular localization, and functional consequence where relevant.
Modification-aware methods are especially important for fragments. Standard small RNA workflows can under-detect heavily modified RNAs because modifications interfere with adapter ligation or reverse transcription. PANDORA-seq used enzymatic treatment to reduce modification-related barriers and expanded the observed small RNA repertoire, including tRNA-derived and other regulatory small RNAs. The biological lesson is not that every newly visible fragment is functional. The lesson is that methods determine which molecules are visible enough to enter the repertoire discussion.
Repeat-derived RNAs come from tandem repeats, transposable elements, retroelements, satellite sequences, endogenous viral sequences, simple repeats, or other repetitive regions. These RNAs can be important. They can participate in transposon silencing, heterochromatin formation, innate immune sensing, nuclear body organization, repeat expansion disease, and genome evolution. They can also be the most artifact-prone part of a repertoire because short reads often cannot be assigned uniquely to one repeat copy.
Repeat-derived RNA claims should specify whether the evidence is locus-specific or family-level. If reads map to many copies, the claim may support expression of a repeat family but not a specific locus. If a long-read method assigns a transcript to a unique locus, confidence improves but still depends on genome assembly and mapping quality. If a repeat-derived transcript overlaps a protein-coding gene or pseudogene, strand and boundary evidence matter. Repeats are an area where repertoire annotation should explicitly use confidence levels rather than binary presence-or-absence calls.
Box 14.2. Count Loci, Molecules, or Mature Products?
- A tRNA locus, a mature tRNA, a tRNA isodecoder family, a charged aminoacyl-tRNA, and a tRNA fragment each answer a different question.
- The same is true for miRNA genes, precursor miRNAs, mature miRNAs, lncRNA loci, lncRNA isoforms, and repeat-derived fragments.
- A repertoire count is meaningful only when the counted object is explicitly stated.
A tRNA locus, a mature tRNA, a tRNA isodecoder family, a charged aminoacyl-tRNA, and a tRNA fragment answer different questions. The same is true for miRNA genes, precursor miRNAs, mature miRNAs, lncRNA loci, lncRNA isoforms, and repeat-derived fragments. A repertoire count is meaningful only when the counted object is stated.
A rare RNA class is uncommon, narrowly distributed, condition-specific, technically difficult to detect, or newly recognized. Rare RNA classes can be biologically central in the right context. A defense RNA may matter only during infection. A developmental RNA may appear only in a transient cell state. A germline RNA may be absent from somatic samples. A stress-induced bacterial sRNA may be invisible in rich laboratory medium. A viroid may matter only in infected plant tissue.
Rare RNA classes challenge annotation pipelines because most automated systems are trained on common patterns. A pipeline optimized for protein-coding genes may miss structured RNAs. A small RNA pipeline optimized for miRNAs may misclassify tRNA fragments. A database model derived from well-sampled animals may not generalize to protists or archaea. A method that captures polyadenylated RNAs may miss many bacterial RNAs, histone mRNAs, non-polyadenylated lncRNAs, organellar RNAs, and degraded or processed small RNAs.
Lineage-specific RNA classes or loci are restricted to a clade, species group, strain, organelle lineage, virus family, or mobile-element family. Lineage specificity can be real. Genomes evolve new repeats, mobile elements, defense systems, and regulatory transcripts. However, apparent lineage specificity can also reflect undersampling, poor assemblies, weak homology search, rapid sequence divergence, or inconsistent annotation. A lineage-specific claim should therefore state the comparison set and the search method.
For structured RNAs, sequence divergence is a major problem. Homology may be detectable by conserved secondary structure even when primary sequence is weak. Chapter 62 and Chapter 140 treat covariance models and RNA homology search in detail. For lncRNAs, primary-sequence conservation can be weak even when promoter context, synteny, structure, or function is conserved. For bacterial sRNAs, strain-level variation can be real and biologically relevant, as the Aggregatibacter example illustrates.
Environment-specific RNAs are detected or functional only under particular conditions. Conditions can include nutrient limitation, temperature shift, oxidative stress, hypoxia, infection, developmental stage, differentiation state, immune activation, drug exposure, or cell-cycle phase. A condition-specific claim must report the condition and the sampling strategy. “Not detected” in one condition should not be generalized to absence from the organism.
Single-cell and spatial transcriptomics have made condition and cell-state specificity more visible, but they also introduce label dependence. A single-cell RNA-seq analysis may label a cell as a fibroblast, macrophage, oligodendrocyte, or disease-associated state. That label helps interpret expression, but it is not itself an RNA gene class. Reviews of single-cell annotation emphasize the dependence on reference data, integration method, and evaluation. Model-assisted cell-type annotation studies, including GPT-4-based annotation, illustrate that labels can be useful while still requiring validation. The caution for RNA repertoires is direct: biological labels organize samples, but RNA gene categories require molecule-level and locus-level evidence.
Missing categories are expected in every repertoire. A class may be truly absent. It may be present but not expressed in the sampled condition. It may be expressed but chemically modified, highly structured, too short, too long, non-polyadenylated, unstable, repetitive, or too divergent for the method used. It may be filtered out during quality control. It may be present in a database but under a different name. It may be part of a viral, organellar, plasmid, or symbiont sequence omitted from the assembly.

Figure 14.4. Why Some RNA Classes Are Missing. Missing RNA categories can reflect true absence, incomplete sampling, condition-specific expression, modification-blocked library construction, weak homology models, or database scope. The figure contrasts these non-detection failure modes with genuine biological absence to show why a missing class is not automatically an absent class.
PANDORA-seq provides a concrete example of method-dependent visibility. By addressing modification barriers, it expanded detection of small RNAs that standard workflows underrepresented. The discussion of “missing allosteric ribozymes” makes a related point for functional structured RNAs: if discovery strategies search the wrong sequence space, condition, or assay readout, the resulting catalog can be biased. Therefore, absence from a repertoire table should be annotated as “not detected under these methods” unless the evidence supports true absence.
Box 14.3. A Missing RNA Class May Be a Method Problem
- A standard RNA-seq dataset can miss non-polyadenylated RNAs, highly modified small RNAs, structured RNAs, unstable transcripts, repeat-derived RNAs, and condition-specific RNAs.
- Modification-aware small-RNA detection can expand the visible repertoire, and discovery of overlooked structured RNAs such as allosteric ribozymes shows how search strategy biases catalogs.
- A complete statement says which methods were used and which classes those methods are expected to miss.
A standard RNA-seq dataset can miss non-polyadenylated RNAs, highly modified small RNAs, structured RNAs, unstable transcripts, repeat-derived RNAs, and condition-specific RNAs. A complete statement says which methods were used and which classes those methods are expected to miss.
Annotation confidence is the degree to which an RNA-producing feature has been correctly located, bounded, classified, and interpreted. It is not a single property. A record may have a confident genomic locus but uncertain mature ends. Another record may have strong expression evidence but uncertain function. A third record may have conserved structure but no expression in the sampled dataset. Confidence should be attached to the specific claim.
For a gene-existence claim, minimum evidence may be a genomic locus with sequence or structure features expected for the class. Stronger evidence includes expression, correct processing, conserved features, and absence of assembly artifacts. For an expression claim, minimum evidence is reproducible reads from an appropriate library. Stronger evidence includes strand specificity, end mapping, independent methods, and controls for mapping ambiguity. For a function claim, minimum evidence depends on mechanism; stronger evidence includes perturbation, rescue, biochemical interaction, localization, and physiological consequence.
A practical confidence ladder begins with raw detection and ends with mechanism. The first rung is sequence compatibility: a locus resembles a known RNA class or produces reads. The second rung is boundary support: transcription start sites, processing sites, splice junctions, mature ends, or circular junctions support the product model. The third rung is biogenesis: the RNA shows the expected precursor, processing enzyme dependence, protein association, modification pattern, or subcellular localization. The fourth rung is comparative or structural support: homologs, covariance, conserved motifs, synteny, or conserved folding support the class. The fifth rung is functional evidence: perturbation changes a relevant phenotype or molecular output, ideally with rescue or orthogonal validation.

Figure 14.5. Annotation Confidence Ladder for RNA Genes. Confidence increases when independent evidence supports boundaries, strand, reproducible expression, biogenesis, context, conservation or lineage-specific plausibility, and perturbation. The figure shows confidence rising from raw read support to boundary and strand support, biogenesis signatures, comparative or structural support, and finally perturbation or rescue, and notes that different RNA classes can enter the ladder at different rungs.
Different RNA classes emphasize different rungs. tRNA and rRNA annotation can rely heavily on conserved structure, but expression and processing remain important. miRNA annotation emphasizes hairpin precursor, Drosha or Dicer processing where applicable, mature duplex pattern, Argonaute loading, and target repression. circRNA annotation emphasizes back-splice junctions, resistance to exonuclease treatment in some assays, independent validation, and exclusion of template switching or trans-splicing artifacts. lncRNA annotation emphasizes transcript boundaries, coding-potential assessment, chromatin and expression context, and mechanism-specific functional evidence.
Computational annotation is essential because no human curator can manually inspect every genome and transcriptome. Yet computational labels are hypotheses. A small RNA annotation tool can reduce label fragmentation by integrating classes, but its output depends on the libraries, genome assembly, mapping parameters, databases, and rule sets used. ITAS was designed as an integrated transcript annotation approach for small RNA, and it illustrates both the usefulness and dependence of such systems.
Metadata quality is part of annotation confidence. If the organism, strain, tissue, cell type, developmental stage, treatment, infection state, or library protocol is mislabeled, a repertoire interpretation can be wrong even when the sequence analysis is technically sound. Bias-invariant RNA-sequencing metadata annotation work supports the broader point that metadata and sample labels affect downstream interpretation.
Every repertoire table should distinguish true absence, non-detection, and not assessed. True absence means that a class is absent with evidence appropriate to the genome and methods. Non-detection means that the class was sought but not found under specified conditions and thresholds. Not assessed means that the methods or annotation system did not evaluate the class. This distinction is especially important for rare, modified, repetitive, organellar, viral, or environment-specific RNAs.
For example, a poly(A)-selected eukaryotic RNA-seq dataset should not be used to claim absence of many non-polyadenylated RNAs. A short-read dataset should not be used to make strong locus-specific repeat claims when reads map to many copies. A genome assembly lacking organellar contigs should not be used to infer absence of organellar RNA genes. A small RNA dataset without modification-aware treatment should not be treated as complete for tRNA-derived fragments or other modified small RNAs.
Figure 14.5. Annotation Confidence Ladder for RNA Genes. The planned figure [Chapter 14](chapter1013.md).fig.05 should show increasing confidence from raw reads to boundary support, biogenesis evidence, comparative or structural evidence, and perturbation or rescue evidence. It should also show that different RNA classes can enter the ladder at different rungs.
Box 14.4. Pseudogene Transcription Needs Mapping Controls
- Pseudogene sequences can be transcribed, but their similarity to parent genes makes read mapping ambiguous.
- Apparent pseudogene expression may come from readthrough, antisense transcription, repetitive sequence, or incomplete annotation.
- A functional pseudogene RNA claim should show reproducible locus-specific expression, correct strand and boundaries, exclusion of parent-gene mapping, and, for regulatory claims, perturbation and rescue.
Genome annotation provides the first layer of an RNA repertoire. Annotation can identify protein-coding genes, structural RNAs, repeats, organellar genes, mobile elements, and conserved RNA families. Comparative genomics strengthens annotation when homologous loci, conserved synteny, conserved sequence motifs, or conserved RNA structures support a class. Covariance models are especially important for structured RNAs because compensatory base-pair changes can preserve structure even when sequence changes.
The limitation is that genomes are assemblies, not direct organisms. Repetitive RNA genes can be collapsed. Small genes can be missed. Contamination can introduce foreign RNA genes. Organellar or plasmid sequences may be omitted or misassembled. Genome annotation also does not prove expression under a given condition.
Transcriptome sequencing shows which RNAs are captured by a protocol. Bulk RNA-seq can reveal abundant mRNAs and long transcripts. Long-read sequencing can improve isoform and end resolution. Small RNA-seq can reveal miRNAs, siRNAs, piRNAs, and fragments. Direct RNA sequencing and modification-aware workflows can address some biases, although each method has its own error profile. Strand specificity, size selection, rRNA depletion, poly(A) selection, cap capture, and end-mapping all change the visible repertoire.
Table 14.4. Detection Biases by RNA Class. Detection barriers, helpful methods, and remaining limitations for RNA classes that are easily missed or misassigned.
| RNA class | Detection barrier | Method that helps | Remaining limitation |
|---|---|---|---|
| Modified tRNAs | Modifications block reverse transcription and ligation | Modification-aware workflows | Residual modification bias |
| tRNA fragments | Modifications and length filters | Modification-aware small RNA-seq | Degradation vs functional fragment |
| miRNAs | Confusion with other small reads | Hairpin and processing criteria | Low-abundance or novel miRNAs |
| piRNAs | Germline-restricted, cluster origin | PIWI immunoprecipitation, germline sampling | Somatic absence, cluster annotation |
| circRNAs | Missed by poly(A) selection | Back-splice junction detection, exonuclease treatment | True circle vs artifact |
| lncRNAs | Low abundance, weak conservation | rRNA depletion, long-read sequencing | Function vs transcriptional byproduct |
| Organellar RNAs | Assembly omission, extensive processing | Organelle-aware assembly, targeted sequencing | Nuclear-dependent maturation |
| Repeat RNAs | Multi-mapping ambiguity | Long-read and repeat-aware mapping | Family vs locus assignment |
| Structured ribozymes | Weak homology, wrong search space | Covariance models, structure-based search | Undiscovered or missing classes |
The limitation is that transcript detection is not equivalent to function. A transcript may be unstable, low abundance, readthrough, degradation, or an artifact. Conversely, failure to detect a transcript does not prove absence. The evidence must be matched to the class and question.
Biogenesis evidence asks whether the RNA is produced by the pathway expected for its class. miRNAs should show precursor and mature products consistent with miRNA processing. piRNAs should show PIWI association and pathway-specific signatures. CRISPR RNAs should correspond to CRISPR arrays and processing. tRNA fragments should have cleavage patterns and modification-aware detection consistent with their proposed origin. circRNAs should show junction evidence and artifact controls.
Perturbation evidence asks what happens when the locus, RNA, processing enzyme, binding partner, or pathway is altered. Stronger evidence links the perturbation to a specific molecular or cellular output and, where possible, rescues the effect. Functional evidence is not required to list a transcript as detected, but it is required to claim that the transcript is a functional RNA gene product.
Across bacteria, RNA repertoires are shaped by genome compactness, operons, phage, plasmids, stress responses, and RNA chaperones. Across archaea, repertoires are shaped by CRISPR systems, guide RNAs, compact genomes, and distinctive transcription and processing. Across eukaryotes, repertoires are shaped by introns, chromatin, cell type, development, repeats, organelles, and specialized small RNA pathways.
Across organelles, repertoires are shaped by genome reduction and dependence on nuclear factors. Across viruses, repertoires are shaped by genome polarity, segmentation, replication strategy, packaging, host range, and immune evasion. Across mobile elements, repertoires are shaped by transposition mechanism, host silencing pathways, reverse transcription, and genomic copy number. Across environmental samples, repertoires are shaped by community composition, incomplete genomes, contamination, and uneven sampling.
The same RNA class can therefore have different meanings by context. A guide RNA in a CRISPR system is not the same as a guide RNA for organellar editing. A genomic RNA in an RNA virus is not the same as an mRNA transcribed from a nuclear locus. A repeat-derived RNA in a germline defense pathway is not the same evidential object as a repeat-derived read in a noisy short-read dataset.
RNA repertoire discovery depends on computational and experimental choices. Homology search, covariance modeling, transcript assembly, small RNA classification, repeat-aware mapping, long-read isoform analysis, direct RNA sequencing, and metadata curation all shape the final catalog. Chapters 128, 132, 140, and 141 treat these methods more deeply.
Computational repertoire discovery usually follows a causal workflow. First, define the biological system and assembly. Second, choose which RNA classes are in scope. Third, collect evidence from sequence, structure, expression, and processing. Fourth, map reads with rules appropriate for repeats, paralogs, organelles, and viruses. Fifth, classify candidate RNAs using curated databases and class-specific criteria. Sixth, assign confidence levels. Seventh, record versioned identifiers, sample metadata, and unresolved categories.
Model-assisted annotation can help organize large datasets, especially in single-cell contexts, but labels should be checked against reference data and evaluation sets. A cell-type label, disease-state label, or cluster label is not a new RNA gene class. It is context for interpreting which RNA genes and products are active in a sample.
Box 14.5. Cell-Type Labels Are Not RNA Gene Classes
- Single-cell and spatial annotation identifies biological states and cell types, which help interpret which RNA genes are active in a sample.
- A cell-type, disease-state, or cluster label is context, not an RNA gene class.
- RNA gene repertoire annotation identifies RNA-producing genomic features and products through molecule-level and locus-level evidence.
Open questions:
Common misconceptions: