Chapter 16. Repeat Classes, Annotation, and Repeat-Derived RNA Outputs

Scope Note

This chapter explains how repetitive genomic sequence and repeat-derived RNA outputs are classified, annotated, and measured. The phrase “repeat-derived RNA” covers several different record types: an RNA intermediate that defines a mobile-element class, a tandem-repeat transcript, a pseudogene transcript related to a parent gene, a small RNA mapped to transposon sequence, a long noncoding RNA containing repeat modules, a double-stranded RNA formed from inverted repeats, or a sequencing signal that cannot be assigned confidently to one genomic source. These cases share one practical problem: repeated sequence makes molecule identity, locus origin, abundance, and record stability harder to establish than for a single-copy mRNA.

Chapter 13 introduces ancient mobile elements and RNA-linked genome innovation. Chapter 14 treats RNA gene repertoires and warns that repeats, pseudogenes, fragments, and paralogs complicate gene counts. Chapter 15 treats genome architecture as an RNA-production system. This chapter owns repeat classes, loci, family and subfamily labels, pseudogene objects, RNA-output types, multimapping, annotation artifacts, and evolutionary turnover of records. Host regulation, immunity, development, aging, cancer, neurological disease, and causal function evidence hand off to Chapter 99. Reverse transcription, target-primed reverse transcription, packaging, recombination, integration, and other replication-cycle mechanisms hand off to Chapter 120.

Executive Summary

Repeats are genomic sequences present in multiple copies. Major categories include tandem repeats, interspersed repeats, DNA transposons, non-LTR retrotransposons such as LINEs and SINEs, LTR retrotransposons, endogenous retroviruses, pseudogenes, low-complexity sequence, and repeat fragments embedded in host genes. A repeat can be transcribed from its own promoter, by readthrough, or as part of an exon, intron, untranslated region, or long noncoding RNA. Corresponding records include full element RNAs, mobile intermediates, embedded modules, small RNAs, duplexes, pseudogene RNAs, translated outputs, fragments, and ambiguous mapping signals.

LINEs, SINEs, LTR retrotransposons, and endogenous retroviruses are central annotation classes because their historical or current mobility involves RNA. A full-length LINE record can encode an RNA intermediate and the proteins needed for autonomous mobilization; a SINE record is usually nonautonomous; an LTR element has terminal repeats and internal regions; and an endogenous retrovirus is inherited germline sequence derived from a retroviral integration. Naming the class-defining RNA intermediate is necessary here, but the reverse-transcription and integration mechanisms belong to Chapter 120.

Pseudogene RNAs require the same caution. A pseudogene is a gene-derived locus that has lost the ancestral coding or structural role, but the locus may still be transcribed. Processed, duplicated, and unitary pseudogenes have different origins and diagnostic features. Parent genes and pseudogenes can remain so similar that short reads are assigned incorrectly. A strong RNA-output annotation must distinguish the pseudogene locus from its parent, identify strand and transcript boundaries, and state whether evidence is unique, probabilistic, or family-level.

Repeat-derived small RNAs, long RNAs, and double-stranded RNAs are distinct annotated outputs. A small-RNA record needs length, ends, strand, pathway association, and family-versus-locus resolution. A repeat-containing long RNA needs transcript context and unique flanking evidence. A double-stranded repeat-RNA record needs evidence for complementary strands or an inverted-repeat duplex rather than sequence repetition alone. The host consequences of these outputs are synthesized in Chapter 99.

Expanded-repeat loci illustrate output plurality without requiring disease-mechanism adjudication here. One locus can yield sense and antisense repeat RNAs, processed fragments, structured or double-stranded species, and translated products. Annotation should record repeat sequence, length, interruptions, strand, transcript context, and product class. The consequences of those products belong to Chapter 99.

The central annotation rule is simple but demanding: a repeat-derived RNA record should not claim finer resolution than its sequence and mapping evidence support. Records should state repeat family or subfamily, locus or copy class, strand, RNA length, transcript context, library type, mapping policy, genome build, repeat annotation version, expression context, and validation method. Without that information, apparent expression may reflect multimapping, collapsed assembly, hidden polymorphism, readthrough transcription, fragments, or annotation drift.

Concept Inventory

  • Repeat: genomic sequence present in multiple copies or composed of recurring units; subclasses include tandem, interspersed, low-complexity, gene-derived, and mobile-element records.
  • Transposable element: a mobile genetic element or descendant record; annotations must distinguish intact, partial, inactive, polymorphic, and embedded copies.
  • LINE and SINE: long and short interspersed retrotransposon classes. Intact LINEs are generally autonomous, whereas SINEs are generally nonautonomous; most genomic copies are not equivalent to active elements.
  • LTR retrotransposon and endogenous retrovirus: terminal-repeat-flanked element classes whose records range from intact loci to solo LTRs and inherited fragments.
  • Pseudogene: a processed, duplicated, or unitary gene-derived locus that lost an ancestral role but can retain transcriptional output and sequence similarity to a parent gene.
  • Repeat-derived RNA: an umbrella record that must be refined to full element, mobile intermediate, embedded module, small RNA, lncRNA, duplex, pseudogene RNA, translated output, fragment, or ambiguous signal.
  • Mapping ambiguity and turnover: repeated sequence limits locus assignment, while assemblies and annotation libraries can split, merge, reclassify, or deprecate repeat records over time.

What to Know Before Reading This Chapter

The main prerequisite is the difference between a genomic locus and a sequence family. A locus is a physical position in a genome assembly, such as one LINE-1 insertion on a particular chromosome. A family is a group of related repeat copies, such as AluY or L1HS. A short sequencing read may identify the family but not the locus. Therefore, a statement such as “Alu is upregulated” may mean that reads from many Alu copies increased, that a few young copies increased, that a host gene containing Alu sequence increased, or that an analysis pipeline changed how multimapping reads were counted.

A second prerequisite is the difference between a repeat locus, a transcript context, and an RNA output. The same family sequence can occur in an independently initiated element RNA, inside a host pre-mRNA, in a processed small RNA, or in an ambiguous multimapping read. Annotation should name the object actually supported by the assay.

A third prerequisite is that repeat annotations are versioned models. Genome assemblies, consensus libraries, family hierarchies, polymorphic-insertion catalogs, and transcript annotations change. A coordinate or family name is interpretable only with the assembly and annotation release that defined it.

The chapter uses five running examples. LINE-1 illustrates an autonomous class and a mobile RNA intermediate. Alu illustrates independent SINE RNA, embedded repeat modules, inverted-repeat duplexes, and mapping ambiguity. Telomeric repeat-containing RNA illustrates tandem-repeat transcription. Pseudogenes illustrate parent-locus ambiguity. piRNA records illustrate why repeat-derived small RNAs require size, strand, cluster, and pathway-association fields.

16.1. Repeat classes and their RNA outputs

Figure 16.1. Repeat Classes and RNA-Output Types

Figure 16.1. Repeat Classes and RNA-Output Types. A repeat family is a DNA classification, whereas an RNA-output record describes the molecule and supported source resolution.

Table 16.1. Repeat Classes and RNA Products. Repeat-class annotation should connect the genomic object to its observed RNA outputs and required qualifiers; shared sequence, overlap, and multimapping can make one RNA record compatible with several repeat origins.

Repeat class Representative object RNA-output records Required qualifiers Main ambiguity
Tandem repeat Telomeric, satellite, microsatellite, expanded locus Long repeat RNA, antisense RNA, fragment, duplex candidate Motif, copy number, interruption, locus, strand Repeat length and assembly representation
LINE Full-length or truncated non-LTR element Full element RNA, antisense RNA, embedded fragment Family, subfamily, completeness, coding capacity, locus Rare intact copies among abundant fragments
SINE Alu or other nonautonomous element Independent SINE RNA, embedded module, inverted-repeat duplex Family, orientation, transcript context, locus resolution Extreme multimapping
LTR or ERV Intact, partial, solo-LTR, or inherited retroviral record Full or partial element RNA, LTR-initiated host RNA, embedded fragment Terminal/internal region, completeness, family, boundaries Inherited record versus current infection or active replication
Pseudogene Processed, duplicated, or unitary locus Sense, antisense, spliced, readthrough, small-RNA, ambiguous signal Parent gene, unique variants, strand, boundaries Parent-gene similarity

16.1.01. From Repeated DNA to RNA Molecules

The phrase “repeat class” names the type of repeated DNA, but the phrase “RNA output” names the molecule or RNA signal produced from that DNA. The distinction matters because one repeat class can produce several RNA outputs. A LINE-1 locus can produce a full-length retrotransposon RNA, a truncated transcript, an antisense transcript, a host gene exonized fragment, or small RNA fragments. A tandem repeat locus can produce a long repeat-containing RNA, short processed fragments, or a double-stranded structure if antisense or inverted-repeat transcripts are present.

Repeats enter RNA in four common ways. First, some repeats have promoters and are transcribed as repeat loci. Full-length LINE-1 RNA and some ERV transcripts belong in this category. Second, repeats can be embedded inside host genes, especially in introns and untranslated regions; the repeat sequence then appears in pre-mRNA or mature RNA without the repeat itself initiating transcription. Third, repeats can donate regulatory sequences such as promoters, enhancers, splice sites, or polyadenylation sites that change host RNA production. Fourth, RNA processing or decay can release repeat-containing fragments that are detected by sequencing.

These routes can overlap. A primate Alu insertion in an intron may be transcribed as part of a pre-mRNA. If another Alu copy in the opposite orientation lies nearby, the two Alu sequences can base-pair and create a double-stranded region. That duplex may be edited by ADAR enzymes, retained in the nucleus, processed, or sensed under some conditions. The same Alu family can also produce Pol III-transcribed SINE RNAs in stress or disease contexts. Therefore, a family name alone is not sufficient; the reader needs the locus, strand, RNA form, and cellular context.

Table 16.1 should be read as a classification aid rather than a claim that all members of a class behave the same way. Tandem repeats often raise questions about RNA structure and repeat length. LINEs raise questions about autonomous retrotransposition and host gene effects. SINEs raise questions about nonautonomous mobilization, editing, and multimapping. LTR retrotransposons and ERVs raise questions about promoter activity, viral-like RNAs, and immune recognition. Pseudogenes raise questions about parent-gene homology and regulatory evidence.

16.1.02. Tandem Repeat RNAs and Repeat Expansion Logic

A tandem repeat contains repeat units arranged head to tail. The repeated unit may be a few nucleotides, as in microsatellites, or longer units in satellite DNA. Telomeric repeats are tandem repeats at chromosome ends. Disease-associated nucleotide repeat expansions are tandem repeats whose length can exceed a pathogenic threshold at a specific locus. The RNA consequence depends on whether the repeat lies in a transcribed region, which strand is transcribed, and whether the repeat length changes RNA structure, RNA processing, translation, or chromatin.

Telomeric repeat-containing RNA provides a non-disease example. Telomeres contain repetitive DNA at chromosome ends, and telomeric repeat-containing RNAs can be detected in cells. A recent study reported increased telomeric repeat-containing RNA in aged human cells, illustrating that tandem repeat transcripts can vary with cellular state rather than being only passive genomic sequence. Telomere RNA biology is treated in greater depth in the telomere chapter, but the example is useful here because it shows that a repeat class traditionally discussed as DNA can have RNA outputs.

Expanded-repeat loci provide a demanding annotation example. A record should state the repeat motif, genomic locus, strand, repeat length, interruption pattern, transcript boundaries, full-length versus fragmentary status, and whether a reported product is RNA, small RNA, duplex RNA, or peptide. Short-read alignments often estimate repeat-containing transcript abundance without measuring the expanded tract itself, while long-read and targeted assays can connect repeat length to a specific RNA molecule. Disease mechanisms and host consequences hand off to Chapter 99.

16.1.03. Interspersed Repeats as RNA Modules

Interspersed repeats are distributed across the genome and often occur as transposable-element copies or fragments. Because they are numerous, host transcripts frequently contain repeat-derived sequence in introns, untranslated regions, exons, or long noncoding RNAs. Annotation should distinguish an independently initiated element RNA from repeat sequence embedded in a larger host transcript and from fragments released during processing or decay.

The word “module” is useful if used carefully. A repeat-derived module is a sequence block embedded in a larger transcript or regulatory record. For example, Alu sequence can occur in one or both orientations, LTR sequence can coincide with a transcript start, and ERV fragments can contribute internal transcript sequence. The record should describe coordinates and transcript context without assuming a host function from sequence overlap alone.

Repeat expression can change with tissue, cell state, stress, development, or disease. Those associations are useful metadata but do not by themselves identify the producing locus or establish a host consequence. A repeat-expression record should preserve its condition labels and supported resolution, while causal interpretation hands off to Chapter 99.

Box 16.1. Repeat Expression Is Not Repeat Function

  • Checklist: identify molecule and source resolution; distinguish independent transcript from embedded sequence; record abundance and condition; hand causal host interpretation to Chapter 99.

A repeat-matching signal can represent a mobile intermediate, an embedded transcript module, a small RNA, a duplex, a fragment, or an ambiguous family-level output. Annotation begins with molecule identity and stops at the resolution the data support; host consequences and function evidence belong to Chapter 99.

16.2. Annotation of LINEs, SINEs, LTR elements, endogenous retroviruses, and mobile RNA intermediates

Table 16.2. Annotation Fields for Mobile RNA Intermediates. Mobile-element records should state autonomy, the class-defining RNA intermediate, transcript boundaries, coding capacity, and processing features; these fields identify the RNA object without substituting for the replication mechanism.

Element class Autonomy label Class-defining RNA object Essential annotation fields Replication handoff
LINE Autonomous when intact Full-length element RNA Family, subfamily, completeness, coding status, strand, locus Target-primed reverse transcription in Chapter 120
SINE Nonautonomous Independent SINE RNA or embedded repeat output Family, orientation, transcript context, borrowed-machinery availability when tested Mobilization mechanism in Chapter 120
LTR retrotransposon Autonomous or nonautonomous by family Terminal-repeat-flanked element RNA LTR/internal regions, completeness, family, transcript boundaries Reverse transcription and integration in Chapter 120
Endogenous retrovirus Inherited retroviral-derived record Full or partial ERV RNA Locus, provirus completeness, orientation, family, transcript form Retroviral replication ancestry and mechanisms in Chapter 120
DNA transposon DNA-level mobility class Transposase mRNA, guide RNA, small RNA, or embedded output Class, terminal repeats, coding status, RNA-output type Excision/integration mechanism in Chapter 120

16.2.01. LINE Records and Autonomous RNA Intermediates

LINEs are long interspersed nuclear elements. A full-length autonomous LINE record contains the sequence information needed to generate a class-defining RNA intermediate and element-encoded proteins. In humans, LINE-1 is the major currently active autonomous family, but most genomic LINE-1-like records are truncated, mutated, rearranged, or old fragments. Annotation should therefore distinguish family and subfamily, full-length versus truncated structure, intact coding capacity, strand, locus, polymorphic insertion status, and whether an RNA observation spans the element’s diagnostic boundaries. Target-primed reverse transcription defines the replication class, but its molecular cycle belongs to Chapter 120.

LINE-1 biology matters for RNA annotation because the RNA has several identities at once. It is a transcript, a template for reverse transcription, a translation substrate, a component of an RNP, and a possible source of host regulatory sequence after insertion. A full-length active LINE-1 RNA differs from a short read mapping to LINE-1 sequence inside a host intron. A chapter or dataset that says “LINE-1 expression increased” should specify whether the evidence supports full-length LINE-1 RNA, family-level repeat signal, locus-specific expression, antisense promoter activity, or host transcripts containing LINE-derived fragments.

Boundary cases are common. Many LINE-1 copies are 5′ truncated, some contain disabling mutations, and some reads derive from old LINE fragments embedded in host genes. A rare intact copy and a widespread fragment family can therefore generate similar short-read signals with different object meanings. Repeat-aware annotation needs family-, subfamily-, locus-, and transcript-context views rather than a single undifferentiated “LINE expression” label.

16.2.02. SINE and Alu RNAs as Nonautonomous Repeats

SINEs are short interspersed nuclear elements. Many SINEs derive from small structured RNAs, such as 7SL RNA or tRNAs, and are transcribed by RNA polymerase III when active as independent SINE transcripts. Primate Alu elements are derived from 7SL-related sequence and are present in very high copy number. An Alu RNA can exist as a SINE transcript, an embedded segment of a host pre-mRNA, an element inside a lncRNA, or one half of an inverted-repeat pair.

SINEs are annotated as nonautonomous because they usually lack the proteins required for their own mobilization and historically or currently use machinery supplied by autonomous elements. This class distinction does not mean that each SINE copy is active. Records should distinguish an independent SINE transcript, repeat sequence embedded in another transcript, a locus-specific insertion, and a family-level multimapping signal. The borrowed replication mechanism is treated in Chapter 120.

Alu sequence illustrates why RNA-output type matters. A 2025 study of an Alu RNA pseudoknot and SRP9/SRP14 association used a defined Alu RNA molecule rather than an undifferentiated family-level count. The annotation lesson is to identify whether the object is an independent SINE RNA, an embedded repeat module, an edited inverted-repeat duplex, or a fragment. Host functional interpretation belongs to Chapter 99.

Alu elements are a major source of mapping ambiguity. Young Alu copies can be highly similar. Inverted Alu pairs inside transcripts can form double-stranded RNA and become substrates for adenosine-to-inosine editing. Short reads that overlap edited Alu sequence may align poorly or ambiguously. Long reads help by connecting repeat sequence to unique flanking sequence, but long-read error profiles, incomplete transcript coverage, and collapsed genome assemblies remain relevant. A credible Alu RNA claim should state whether it is family-level, subfamily-level, locus-level, or transcript-context-level.

16.2.03. LTR Retrotransposons and Endogenous Retroviruses

LTR retrotransposons are retroelements flanked by long terminal repeats. The LTRs contain regulatory sequence that can initiate transcription, terminate transcription, and influence nearby host genes. Endogenous retroviruses are retrovirus-derived sequences that entered the germline and became inherited as host genomic sequence. Their original retroviral architecture often included gag, pol, and sometimes env-like regions, but many endogenous copies are now fragmented or mutated.

LTR and ERV annotations can yield several RNA-output records. A record may represent a full or partial element transcript, an LTR-initiated host transcript, an internal fragment embedded in an mRNA or lncRNA, a bidirectional output, or a double-stranded repeat-RNA candidate. The annotation should identify terminal and internal regions, orientation, completeness, family or subfamily, transcript boundaries, and unique flanking sequence.

Plant LTR retrotransposons provide a comparative annotation example. An eccDNA-based study identified low-copy-number LTR records active in carrot callus cultures. The study illustrates that genomic copy number, RNA detection, and evidence of recent mobility are different record fields. Steady-state RNA abundance alone should not be used to label an element replication-competent.

An endogenous-retrovirus record denotes inherited germline sequence derived from retroviral integration; it does not denote current infection. Expression status, locus completeness, coding potential, family assignment, and transcript form should be represented separately. Host regulatory, immune, developmental, or disease consequences of an ERV-derived output belong to Chapter 99, and retroviral replication mechanisms belong to Chapter 120.

16.2.04. DNA Transposon and Mobile-RNA Boundary Records

DNA transposons are classified separately from retrotransposons because their mobility does not require an RNA-to-DNA intermediate, although their loci can produce transposase mRNAs, guide RNAs, small RNAs, or host transcripts containing transposon sequence. Some RNA-guided mobile systems blur simple category labels. This chapter records the element class and RNA-output type; the copying, excision, integration, and RNA-guided replication mechanisms belong to Chapter 120.

16.3. Pseudogene loci, transcription, and RNA-output annotation

16.3.01. What Pseudogenes Are and Why They Are Transcribed

A pseudogene is a sequence related to a gene but disabled for the ancestral gene function. This definition is historical and comparative. It does not say whether the locus is transcribed today. A processed pseudogene forms when a mature mRNA is reverse-transcribed and inserted back into the genome; it often lacks introns and may contain a poly(A)-derived tract. A duplicated pseudogene forms when a gene duplication is followed by mutation or regulatory loss in one copy. A unitary pseudogene is a disabled ancestral gene locus without a functional paralog in the same genome.

Pseudogene transcription has several possible causes. A processed pseudogene may land near an active promoter or acquire regulatory elements. A duplicated pseudogene may retain some ancestral promoter activity. A nearby gene may read through the pseudogene. A transposable element inside or near the pseudogene may provide a promoter. Chromatin derepression may increase transcription across many normally silent loci. Therefore, detecting a pseudogene transcript is the beginning of interpretation, not the conclusion.

The RNA product can also differ from the annotated pseudogene. The transcript may cover only part of the pseudogene. It may be antisense. It may be spliced with nearby exons. It may be an unstable nuclear RNA. It may be a small-RNA precursor. It may be indistinguishable from the parent gene by short-read sequencing. For these reasons, pseudogene expression should be described with transcript boundaries, strand, library method, and mapping controls.

16.3.02. Pseudogene RNA-Output Records and Parent-Gene Ambiguity

Pseudogene RNA-output records should first describe what was measured. Possible records include a sense or antisense primary transcript, a spliced RNA, a readthrough product from a neighboring locus, a small-RNA precursor, a repeat-containing fragment, or a signal that cannot be resolved from the parent gene. These output labels do not assert regulatory function. Proposed host consequences and causal evidence belong to Chapter 99.

Parent-gene ambiguity is the central annotation problem. Processed pseudogenes can resemble mature parent-gene mRNA, while duplicated pseudogenes can preserve intron-exon organization and long stretches of near identity. Unique substitutions, pseudogene-specific junctions, unique flanking sequence, long reads, phased variants, and locus-specific assays can support assignment. When none is available, a gene-family or parent-pseudogene-group label is more accurate than a false locus call.

Pseudogene-derived sequence can also appear in small-RNA libraries. Such a record should state length, strand, end pattern, parent-gene ambiguity, pathway association when measured, and whether the source is one locus or a sequence family. A small fragment matching a pseudogene is not automatically a processed regulatory small RNA.

16.3.03. Evidence Standards for Pseudogene RNA Annotation

The first evidence standard is molecular identity. The study must show that the RNA comes from the pseudogene locus rather than the parent gene. Unique sequence differences, splice junctions, long reads spanning pseudogene-specific variants, or targeted assays can help. Short reads that map equally to the pseudogene and parent gene are not sufficient unless the quantification model explicitly handles ambiguity and the conclusion is framed at the correct resolution.

Box 16.2. Multimapping Reads Are Data, Not Trash

  • Guidance: retain family-level information; declare whether reads were discarded, fractionally assigned, probabilistically assigned, or anchored by unique sequence; report uncertainty; do not name a locus without locus-specific evidence.

The second evidence standard is transcript context. The annotation should state whether signal is nascent or steady-state, sense or antisense, nuclear or cytoplasmic when known, full length or fragmentary, and constitutive or condition-associated. A signal detected only after strong stress or epigenetic treatment should not be generalized to an ordinary state.

The third evidence standard is orthogonal assignment. Long reads can connect pseudogene-specific variants to transcript boundaries; targeted amplification can test strand and splice structure; unique k-mers can support locus identity; and genomic sequencing can confirm that the source allele exists in the sample. Each method has biases, so agreement among independent sequence features is stronger than repeated use of the same ambiguous short reads.

The fourth evidence standard is an explicit record status. A locus can be annotated as transcription-supported, transcript-structure-supported, family-level-only, ambiguous with parent gene, or unobserved in the tested condition. These states are more informative than a binary expressed/not-expressed label and can be revised as assemblies or long-read evidence improve.

16.4. Repeat-derived small RNAs, lncRNAs, and dsRNAs as annotated outputs

16.4.01. Repeat-Derived Small-RNA Records

Small RNAs are short regulatory RNAs that guide protein complexes to targets by base pairing or sequence-related recognition. Repeat-derived small RNAs are produced from transposons, tandem repeats, pseudogene fragments, inverted repeats, or other repetitive sequence. They can be functional products of genome defense pathways, but they can also be RNA degradation fragments misannotated as regulatory RNAs.

piRNAs are a major annotated repeat-derived small-RNA class. A piRNA record commonly includes a characteristic length range, strand distribution, terminal chemistry or sequence bias, association with Piwi-clade Argonaute proteins, and mapping to repeat families or genomic clusters. Those fields distinguish a curated piRNA output from an arbitrary short repeat-matching fragment. Biogenesis and host consequences are treated in Chapter 88 and Chapter 99.

Repeat-derived siRNA records are common in plants, fungi, and several animal systems. Annotation should record size class, strand, precursor evidence, Dicer or Argonaute association when measured, repeat-family assignment, and whether multimapping is resolved at family or cluster level. Similar lengths alone do not make every repeat fragment an siRNA.

The evidence basis for repeat-derived small RNAs includes size distribution, strand bias, end chemistry, dependence on pathway enzymes, loading into Argonaute or Piwi proteins, mapping to repeat families, and target effects. Mapping remains difficult. A 26-nucleotide small RNA matching a transposon family may map to hundreds of loci. Sometimes the family-level assignment is the correct biological level because the small RNA regulates a family. Sometimes locus-level origin matters because a specific cluster produces the precursor. The analysis should state which level is supported.

16.4.02. Repeat-Containing lncRNA Records

Long noncoding RNAs, or lncRNAs, are transcripts longer than small RNAs that do not have a primary protein-coding role. Many lncRNAs contain repeat sequence because genomes are repeat-rich. Annotation should distinguish an lncRNA initiated by a repeat-derived promoter, an lncRNA with embedded repeat modules, a transcript composed largely of repeats, and a locus-level repeat transcript that merely overlaps an lncRNA model.

Repeat modules can make lncRNAs hard to analyze. If a lncRNA contains several Alu elements, reads from the repeat segments may map to many loci, while reads from unique exons identify the lncRNA. If the repeat modules drive function, deleting only unique exons may not test the relevant motif. If the lncRNA is regulated by an LTR promoter, perturbing the LTR may affect both promoter DNA and RNA output. If the lncRNA contains inverted repeats, it may form double-stranded structures that change localization or immune sensing.

The phrase “repeat-derived lncRNA” is too broad unless the repeat contribution is specified. A useful record names the transcript model, repeat family and coordinates, orientation, exon or intron context, unique sequence anchors, and whether the repeat supplies initiation, an embedded module, an inverted pair, or most of the transcript body. General lncRNA biology and host consequences hand off to Chapter 90, Chapter 91, and Chapter 99.

16.4.03. Double-Stranded Repeat-RNA Records

Double-stranded RNA, or dsRNA, is RNA in which complementary strands form an extended duplex. Viral replication can generate dsRNA, but cells can also produce self-derived dsRNA from inverted repeats, antisense transcription, mitochondrial transcripts, structured RNAs, or repeat-containing transcripts. Repeat-derived dsRNA is especially common when similar copies occur in opposite orientations within the same transcript or when sense and antisense transcription overlap.

A double-stranded repeat-RNA record should state how duplex formation was inferred: inverted repeats within one transcript, overlapping sense and antisense products, intermolecular pairing, editing patterns, duplex-enrichment assays, or direct structural evidence. Length, orientation, compartment, end chemistry, editing state, and protein association may be relevant record fields. Host sensing and immune consequences belong to Chapter 99.

Repeat-derived dsRNA can become more detectable when transcription, processing, editing, or compartmentalization changes. Those contexts should be recorded without converting an output annotation into an immune-mechanism claim. Sensor biochemistry and disease synthesis are treated in Chapter 99, Chapter 108, and Chapter 109.

Detection of dsRNA has artifacts. Antibodies recognizing dsRNA can vary in specificity. RNA extraction and library preparation can create or lose duplex information. Short reads may not prove that two complementary repeat segments existed as a duplex in the cell. ADAR editing can imply dsRNA formation, but editing site detection depends on mapping quality and RNA abundance. Strong claims combine strand-specific transcript evidence, editing patterns, dsRNA enrichment, sensor dependence, and perturbation of the repeat-containing source.

16.4.04. Expanded-Repeat and Translated-Output Records

Expanded-repeat loci can yield multiple annotated products: sense and antisense RNAs, alternative transcript boundaries, structured or duplex RNA, processed fragments, RNA foci detected by imaging, and repeat-associated translated products. A record should state the repeat motif, repeat length and interruptions, strand, transcript context, cell type, assay, and whether the observation refers to RNA or peptide.

Repeat-associated non-AUG translation is named here only to distinguish a translated output from the repeat RNA that served as template. The molecular translation mechanism and disease consequences are not repeat-annotation questions; they hand off to Chapter 99.

The annotation boundary is therefore object-specific. Evidence for repeat RNA abundance should not be entered as evidence for peptide abundance, and peptide detection should not be used to infer the amount or structure of every repeat RNA isoform. Chapter 99 compares their host consequences and causal evidence.

16.5. Multimapping, annotation artifacts, and evolutionary turnover of repeat records

16.5.01. What a Repeat-Derived RNA Record Must Specify

Figure 16.5. Artifact Triage for Repeat-Derived RNA Records

Figure 16.5. Artifact Triage for Repeat-Derived RNA Records. Every repeat-derived RNA record should stop at the finest resolution supported by its sequence, mapping, and orthogonal evidence.

Table 16.5. Repeat RNA Annotation Failure Modes. Apparent repeat-derived RNAs can result from multimapping, readthrough, fragment accumulation, assembly error, or incomplete annotation; targeted controls determine whether a specific repeat-RNA label is supported or must remain unresolved.

Failure mode Misleading appearance Control Supported resolution after control
Random assignment of multimappers False locus-specific expression Probabilistic or family-aware quantification Family, subfamily, or confidence-weighted locus
Discarding all multimappers Absent repeat-family expression Retain ambiguous counts with declared model Family-level abundance
Collapsed or missing assembly sequence Merged copies or absent polymorphic insertion Long-read or pangenome assembly and genomic validation Sample-specific locus or unresolved group
Parent-pseudogene similarity False pseudogene or parent expression Unique variants, junctions, flanks, or long reads Locus, parent-pseudogene group, or unresolved
Intronic or readthrough signal False independent repeat transcript Strand-aware start/end and transcript-context evidence Embedded module, nascent overlap, or independent RNA
Annotation-library drift Apparent expression change between analyses Preserve library version and remap consistently Versioned comparison with split/merge history

Repeat-derived RNA records are unusually sensitive to missing metadata. A minimal expression record should specify the repeat family or subfamily, locus or copy class if known, strand, RNA length range, library method, sequencing read length, genome build, repeat annotation source and version, mapping strategy, multimapping policy, and expression context. Host-function evidence is a separate layer owned by Chapter 99.

The level of resolution should match the data. If reads cannot be assigned to a specific Alu copy, the claim should not name a specific Alu locus. If only total family-level signal was measured, the claim should be family-level. If long reads connect a repeat to unique flanking exons, the claim can describe a transcript isoform. If nascent RNA data show transcription through a repeat but steady-state RNA is absent, the claim should distinguish transcription from stable RNA accumulation. This discipline prevents overinterpretation without discarding real repeat biology.

Figure 16.6. Repeat-Locus Anatomy and the RNA Evidence Visible from Each Locus State

Figure 16.6. Repeat-Locus Anatomy and the RNA Evidence Visible from Each Locus State. Repeat-family signal becomes a locus-level or full-length RNA claim only when reads or orthogonal evidence resolve the boundaries and unique features required by that locus class.

Repeat-aware quantification can use several strategies. One strategy discards multimapping reads, which reduces false locus assignment but undercounts repetitive families. Another strategy assigns multimapping reads probabilistically, which estimates family or locus abundance but depends on model assumptions. A third strategy uses repeat-family consensus annotation. A fourth uses uniquely mappable flanking sequence. TEtranscripts, Telescope, and SQuIRE illustrate complementary approaches for including transposable elements, estimating retrotranscriptome expression, and resolving locus-specific interspersed-repeat expression from RNA-seq data. Long-read sequencing can connect repeats to larger transcript structures. No strategy is universally best; the correct method depends on whether the biological question is family activity, locus-specific transcription, transcript isoform structure, or functional molecule identity.

16.5.02. Common Artifact Classes

Multimapping is the most common artifact. A read from a young repeat family may align equally well to many locations. If the pipeline randomly assigns the read to one locus, apparent locus-specific expression can be false. If the pipeline discards the read, repeat-family expression can be underestimated. If the pipeline assigns reads only to annotated genes, repeat-derived transcripts may be hidden inside host gene counts.

Genome assembly and annotation artifacts are also common. Repeats are hard to assemble because similar copies collapse or misjoin. A reference genome may lack polymorphic insertions present in the sample. RepeatMasker or related annotations may classify an element at one family resolution while the biological question needs another. Genome build changes can move coordinates, split loci, or change mappability. Metadata and annotation-context bias can also affect interpretation across public RNA-seq datasets.

RNA processing artifacts can mislead. Intronic reads can reflect nascent transcription, unspliced pre-mRNA, retained introns, or RNA decay. Small RNA reads can be functional small RNAs or degradation fragments. Circular RNA detection can be confounded by repeats that promote misalignment. Antisense signals can arise from library artifacts if strand specificity is weak. Apparent dsRNA may reflect extracted RNA annealing or antibody bias rather than an in-cell duplex. Each artifact has an experimental control; the control should be chosen before a functional story is built.

Perturbation artifacts are especially important. CRISPR deletion of a repeat removes DNA sequence as well as RNA output. CRISPR interference can repress nearby promoters. RNA interference against a repeat family can hit many transcripts because the sequence is shared. Antisense oligonucleotides can change splicing or nuclear retention. Overexpression of a repeat RNA can create nonphysiological abundance and localization. Rescue experiments, orthogonal perturbations, and endogenous-level assays help separate mechanism from method.

16.5.03. Evolutionary Turnover of Repeat Records

Repeat records turn over rapidly. New insertions arise, old copies accumulate mutations, recombination deletes or rearranges loci, and lineage-specific expansions create family and subfamily structures that differ among assemblies and populations. A repeat family present in one reference may be absent, collapsed, or differently named in another.

Turnover changes both coordinates and classifications. A locus can move between family assignments as consensus libraries improve; fragments can be merged into one record or split into several; a polymorphic insertion can become represented in a pangenome; and a young subfamily can be separated from an older broad family. RNA-output records must therefore preserve the genome assembly, annotation library, family hierarchy, and release used for quantification.

Comparative annotation should distinguish biological absence from record absence. Failure to detect a repeat-derived RNA in another species can reflect true lineage restriction, incomplete assemblies, missing repeat libraries, low mappability, untested conditions, or incompatible naming schemes. Host evolutionary consequences and exaptation claims belong to Chapter 99; this chapter records the object and its turnover.

16.5.04. Record Confidence and Revision

A practical confidence ladder begins with detection of repeat-matching signal. The next levels establish family or subfamily, supported locus resolution, strand, transcript boundaries, full-length versus fragmentary status, and orthogonal validation. Each field can carry its own confidence rather than forcing one binary score on the whole record.

Records should be revised rather than silently replaced when a new assembly, repeat library, long-read transcript, or mapping model changes the assignment. A revision should preserve the earlier identifier when possible, state whether the record was split, merged, reclassified, or deprecated, and document the evidence that changed. Functional consequence layers maintained elsewhere should point to the versioned repeat and RNA-output object.

Recent Consensus

Repeats are major sources of heterogeneous RNA outputs. A useful annotation connects the DNA object and family hierarchy to an RNA molecule type while preserving locus, strand, transcript context, and supported mapping resolution.

LINEs, SINEs, LTR elements, endogenous retroviruses, DNA transposons, tandem repeats, and pseudogenes require different object schemas. A class-defining RNA intermediate may be named, but detailed replication cycles hand off to Chapter 120.

Pseudogene transcription is real but difficult to assign. Parent-gene homology, readthrough, duplicated sequence, and incomplete transcript models make locus-level RNA annotation especially demanding.

Repeat-derived small RNAs, long RNAs, duplex RNAs, embedded modules, and translated products are different output classes. Each needs molecule-specific fields rather than one umbrella “repeat expression” record.

Multimapping is information about resolution, not automatically unusable noise. Family-level quantification can be valid when locus assignment is impossible, provided the model, denominator, annotation release, and uncertainty are explicit.

Repeat records turn over with assemblies, populations, consensus libraries, family hierarchies, and transcript evidence. Versioned records should preserve split, merge, reclassification, and deprecation history. Host consequences and causal function evidence hand off to Chapter 99.

Open Questions, Controversies, Deprecated Models, and Common Misconceptions

Open questions:

  • How should repeat and RNA-output schemas represent pangenome insertions, sample-specific repeats, and family hierarchies that differ among assemblies?
  • How should repeat RNA quantification represent multimapping reads when the correct biological resolution may be family-level rather than locus-level?
  • Which pseudogene RNA records can be resolved confidently from parent genes using long reads, phased variants, or unique flanking sequence?
  • How should small-RNA and dsRNA records represent family-level origin, duplex evidence, processing state, and multimapping uncertainty?
  • How should long-read and direct RNA sequencing be integrated with short-read repeat quantification for highly similar or polymorphic repeats?

Common misconceptions:

  • “Repeat-derived RNA is one molecule class.” Repeat sequence can occur in full-length element RNAs, embedded transcript modules, small RNAs, duplexes, fragments, or ambiguous family-level signals.
  • “Repeat expression proves repeat function.” Expression establishes an RNA signal at a supported resolution; host consequence and function require the evidence synthesized in Chapter 99.
  • “Every pseudogene read comes from the pseudogene locus.” Parent-gene similarity and readthrough can make locus assignment ambiguous.
  • “Multimapping reads are unusable.” Multimapping reads are informative at family or model-based levels if the analysis states its assumptions.
  • “Endogenous retrovirus expression means current infection.” ERV RNA can come from inherited genomic copies and should be annotated by locus and transcript form.
  • “The latest repeat library can be compared directly with every older result.” Family definitions, coordinates, and copy sets can change between releases.

Deprecated or weakened claims:

  • Broad “junk DNA” language is obsolete as a scientific explanation because it hides mechanistic diversity. It is still true that many repeat copies lack assigned function, but absence of assigned function should be stated directly rather than used as a category label.
  • Broad “everything is regulatory” interpretations are unsupported; this chapter catalogs objects and outputs without converting detection into function.
  • Randomly assigning every multimapping read to one locus is a weakened annotation practice when the data support only family-level abundance.