# Chapter 77. smORFs, Microproteins, Noncanonical ORFs, and Coding Potential of Annotated ncRNAs

## Scope Note

This chapter explains how short open reading frames (smORFs), upstream open reading frames (uORFs), alternative reading frames, and open reading frames embedded in annotated long noncoding RNAs (lncRNAs) or circular RNAs (circRNAs) are discovered, interpreted, and tested. The chapter treats translation evidence and biological function as related but separate questions: a ribosome can occupy an RNA without yielding a stable peptide, a peptide can be detected without being functional, and a functional RNA can also contain a translated region. The main goal is to give readers an evidence framework for deciding when a small coding region should be annotated as a protein-coding feature, a regulatory translation event, a probable artifact, or an unresolved candidate.

## Executive Summary

Small open reading frames are short nucleotide segments that can be translated into peptides. In many annotation systems, an open reading frame shorter than about 100 codons, sometimes 150 codons, has historically been treated with suspicion because random sequence produces many short start-to-stop intervals. That length filter protected genome annotations from false positives, but it also delayed recognition of genuine microproteins, which are small proteins or peptides with biological activity. Some microproteins act as membrane-associated regulators, some tune large protein complexes, some participate in developmental signaling, and some are presented to immune cells as short antigens. The same translation machinery that makes canonical proteins can also initiate on uORFs in 5′ untranslated regions, alternative ORFs overlapping annotated coding sequences, downstream ORFs, lncRNA-associated ORFs, and circRNA-associated ORFs.

The central interpretive difficulty is that "coding potential" has several meanings. A transcript may contain an ORF by chance. An ORF may be bound by ribosomes. Ribosomes may initiate and elongate through the ORF. Elongation may produce a stable peptide. The peptide may accumulate, localize, interact, and perform a function. A genetic perturbation may alter phenotype because of the peptide, because of the act of translation, because of RNA structure or RNA-binding protein occupancy, or because the transcript annotation is incomplete. Strong claims require evidence that separates these possibilities.

Ribosome profiling transformed this field by showing where translating ribosomes protect short RNA fragments from nuclease digestion. With sufficient read depth, triplet periodicity, initiation-site enrichment, stop-codon behavior, and reproducible ORF boundaries, ribosome profiling can support active translation of smORFs and uORFs. However, ribosome profiling alone usually cannot prove that a stable microprotein exists or that the peptide has a biological function. Proteomics and immunopeptidomics can detect peptide products, but short peptides are difficult to identify because they produce few unique mass spectra, are often low abundance, may be hydrophobic or unstable, and require carefully controlled search databases. Functional validation requires perturbations that preserve RNA-level effects while changing peptide production, such as start-codon mutation, stop-codon insertion, frame disruption, synonymous rescue, endogenous tagging, or rescue by peptide expression.

The current consensus is neither that most annotated ncRNAs secretly encode proteins nor that all noncanonical translation is noise. Annotated lncRNAs and circRNAs can contain translated ORFs, but the fraction that produce stable functional peptides is much smaller than the fraction with some ribosome association. uORFs are widespread regulators of downstream coding sequence translation, and many uORFs function through the act of translation rather than through the peptide product. Microproteins are real, biologically important, and still under-annotated, especially when they are short, lineage-specific, membrane-associated, or expressed only in specialized cell states. The field is moving toward tiered annotation in which ORF prediction, translation evidence, peptide evidence, conservation, and function are recorded separately instead of being collapsed into a binary coding or noncoding label.

## Concept Inventory

- **Small open reading frame (smORF):** A short start-to-stop coding interval, often operationally defined as fewer than 100 or 150 amino acids. The exact cutoff differs by database and study. The term describes sequence architecture, not proven function.
- **Microprotein or micropeptide:** A small translated product, commonly under 100 amino acids, that has evidence of accumulation or biological activity. "Microprotein" is often used when the product folds, localizes, or interacts like a protein; "micropeptide" is often used more broadly for short translated products.
- **Upstream open reading frame (uORF):** An ORF in the 5′ untranslated region of an mRNA, upstream of the main coding sequence. uORFs often regulate translation of the downstream coding sequence by capturing scanning ribosomes, altering reinitiation, or triggering decay.
- **Alternative ORF:** An ORF that differs from the annotated main coding sequence. It may be in a different frame, overlap the canonical coding region, start at a downstream or upstream initiation codon, or reside in a transcript previously classified as noncoding.
- **Annotated ncRNA with coding potential:** A transcript classified as noncoding by a particular annotation release but containing sequence or experimental evidence compatible with translation. The label does not prove that the transcript's main function is peptide-mediated.
- **Ribosome occupancy:** Enrichment of ribosome-protected RNA fragments on a transcript or ORF. Occupancy supports ribosome association but must be interpreted with fragment periodicity, drug treatment, nuclease bias, read mapping, and transcript isoform structure.
- **Peptide evidence:** Detection of a peptide product by mass spectrometry, immunoblotting, epitope tagging, immunoprecipitation, imaging, or other protein-level methods. Peptide evidence is strongest when the detected sequence is unique and endogenous.
- **Functional evidence:** Experiments showing that altering the ORF or peptide changes a biological phenotype through the peptide product or the act of translation, with RNA abundance, RNA localization, and neighboring regulatory elements controlled.

## What to Know Before Reading This Chapter

Readers should understand the basic organization of a eukaryotic mRNA: a 5′ cap, a 5′ untranslated region, a coding sequence, a stop codon, a 3′ untranslated region, and a poly(A) tail. [Chapter 67](chapter1062.md) explains how scanning ribosomes usually initiate at an AUG codon in a favorable sequence context, while [Chapter 70](chapter1065.md) explains how uORFs, internal ribosome entry sites, and RNA features regulate translation initiation. Those rules matter here because most noncanonical ORFs are discovered by asking whether ribosomes initiate at places that annotation pipelines did not designate as main coding sequences.

Readers should also distinguish a transcript from a gene model. A gene may produce several isoforms with different first exons, 5′ untranslated regions, splicing patterns, polyadenylation sites, and coding intervals. A transcript annotated as an lncRNA in one release may later be split, merged, reclassified, or found to share sequence with an unrecognized protein-coding isoform. [Chapter 18](chapter1017.md) treats transcript annotation as a versioned interpretation of evidence, not a permanent property of a genomic locus.

Finally, this chapter uses "translation" in two senses that must be kept separate. Biochemically, translation is the ribosome-catalyzed synthesis of a peptide from codons. Functionally, translation can regulate an RNA even when the peptide product is unstable or irrelevant. For example, a uORF can repress a downstream coding sequence because scanning ribosomes initiate upstream and then fail to reinitiate efficiently. In that case the act of uORF translation is functional even if the uORF peptide is rapidly degraded and never acts as an independent molecule.

## Core Mechanisms and Molecular Players

## 77.1. smORF and uORF discovery

Small ORFs are abundant in nucleotide sequences because stop codons occur frequently by chance. In a random sequence with the standard genetic code, a short start-to-stop interval is expected far more often than a long uninterrupted coding sequence. Early gene-finding systems therefore used length thresholds, codon composition, splice structure, transcript support, and cross-species conservation to avoid annotating thousands of accidental ORFs. This conservative strategy worked reasonably well for large canonical proteins, but it made short coding regions difficult to distinguish from background sequence. The discovery problem is therefore not simply "find all start codons." It is to identify which short ORFs are translated, which translation events matter, and which short ORFs should be represented in gene annotations.

The term smORF is intentionally broad. A smORF may lie in the 5′ untranslated region of an mRNA, in the main body of an annotated lncRNA, in an alternatively spliced exon, in a circRNA, in a pseudogene transcript, or overlapping a canonical coding sequence in another frame. The same sequence feature can receive different labels depending on transcript context. A short ORF upstream of a known coding sequence is a uORF. A short ORF that produces a functional peptide is often called a microprotein or micropeptide. A short ORF detected only by computational prediction remains a candidate smORF until translation and function are tested.

![Figure 77.1. Evidence Ladder for a Candidate smORF](../assets/figures/chapter1072_figure1.png)

**Figure 77.1. Evidence Ladder for a Candidate smORF.** A candidate small open reading frame should move through separate evidence levels: sequence plausibility, expression in a defined transcript isoform, ribosome profiling evidence, peptide detection, conservation, perturbation, and rescue. The figure should show that these levels are cumulative but not interchangeable. Ribosome occupancy supports translation, peptide detection supports product formation, and rescue experiments support function.

uORFs were recognized before genome-wide ribosome profiling because individual mRNAs showed leader sequences that altered downstream translation. A scanning ribosome begins at the 5′ end of a capped eukaryotic mRNA and inspects the leader sequence for an initiation codon in a suitable context. When a ribosome initiates at a uORF, it can terminate before reaching the main coding sequence. After termination, the 40S ribosomal subunit may dissociate, remain bound and reinitiate downstream, or stall in a way that affects mRNA decay. The outcome depends on uORF length, intercistronic distance, initiation context, availability of initiation factors, peptide-mediated stalling, and stress-regulated phosphorylation of initiation factors. [Chapter 70](chapter1065.md) treats these regulatory mechanisms in detail; this chapter focuses on how uORFs enter the broader landscape of noncanonical coding.

![Figure 77.2. Outcomes of uORF Translation](../assets/figures/chapter1072_figure2.png)

**Figure 77.2. Outcomes of uORF Translation.** A ribosome scanning from the 5′ cap can bypass a weak uORF start codon, initiate at the uORF and dissociate, terminate and reinitiate downstream, stall during uORF translation, or trigger decay pathways. The same uORF can therefore regulate a downstream coding sequence through ribosome behavior even when the uORF peptide is not a stable effector.

Genome-wide discovery uses several complementary signals. Sequence-based screens search for start codons, in-frame stop codons, codon bias, Kozak-like initiation contexts, predicted signal peptides or transmembrane helices, and conservation of coding frame. Transcriptomic data ask whether the candidate ORF is present in an expressed isoform. Ribosome profiling asks whether ribosome-protected fragments align to the ORF with three-nucleotide periodicity. Translation initiation profiling, commonly using drugs or conditions that enrich initiating ribosomes, can nominate start sites. Proteomics asks whether the peptide product can be observed. Genetic experiments ask whether disrupting the ORF changes a phenotype.

Each discovery route has a different false-positive profile. A sequence-only screen finds many accidental ORFs. Transcript expression proves only that RNA exists. Ribosome occupancy may mark elongating ribosomes, scanning ribosomes, stalled ribosomes, RNA-protein complexes that protect similar fragments, or mapping artifacts. Proteomic database searches can produce spurious matches when the search space is inflated with many predicted smORFs. Functional screens may detect RNA-level effects rather than peptide-level effects. Good discovery pipelines therefore keep evidence classes separate and upgrade confidence only when independent methods converge.

**Table 77.1. Evidence Thresholds for Noncanonical ORF Claims.** Different evidence classes support different claims. Sequence evidence supports a candidate ORF; transcript evidence supports an expressed substrate; ribosome profiling supports probable translation; proteomics supports peptide production; conservation supports evolutionary constraint; perturbation and rescue support function.

| Evidence class | What the evidence can support | What the evidence cannot support alone | Typical false positives | Useful controls |
| --- | --- | --- | --- | --- |
| **Sequence prediction** | A start-to-stop interval, initiation context, codon composition, or predicted motif makes an ORF plausible. | Expression, ribosome use, peptide production, or function. | Random short ORFs, repeat-derived intervals, near-cognate start noise, unrecognized overlap with another feature. | State genome assembly, transcript ID, ORF coordinates, start codon, frame, length cutoff, and annotation release. |
| **Transcript isoform evidence** | The candidate ORF lies in an expressed, correctly spliced RNA substrate in a defined sample. | Translation of that ORF or peptide accumulation. | Reads from paralogs, pseudogenes, host-gene isoforms, incomplete transcript models, or merged gene annotations. | Use isoform-resolving RNA evidence, junction support, strand specificity, long-read data where available, and annotation-version checks. |
| **Ribosome profiling** | Frame-specific ribosome-protected fragments support probable translation of the ORF in that condition. | Stable peptide existence, autonomous peptide function, or correct start-codon assignment by itself. | Nonribosomal RNP protection, nuclease bias, drug or lysis artifacts, ambiguous mapping, too few reads over a short ORF. | Require triplet periodicity, reproducible ORF boundaries, calibrated offsets, appropriate controls, and comparison with surrounding untranslated sequence. |
| **Translation initiation profiling** | Enriched initiating ribosomes nominate a candidate start codon, including possible non-AUG starts. | Complete elongation through the ORF, peptide stability, or function. | Drug-induced stalls, secondary peaks, weak near-cognate initiation events, start sites from an unmodeled isoform. | Pair initiation peaks with elongating Ribo-seq signal, start-codon mutation, replicate support, and transcript-isoform validation. |
| **Mass spectrometry proteomics** | A unique endogenous peptide or targeted spectrum supports peptide product formation. | The exact RNA template, translation mechanism, or biological function. | Inflated candidate databases, shared peptides, poor spectra, missed modifications, contaminant assignments. | Use strict target-decoy false-discovery control, unique peptide evidence, spectrum validation, targeted confirmation, and a controlled smORF database. |
| **Immunopeptidomics** | A peptide from a noncanonical ORF is produced, processed, and presented by major histocompatibility complex molecules. | Autonomous intracellular function or T-cell recognition without further assays. | Ambiguous source proteins, contaminant peptides, allele-assignment uncertainty, expanded-search false discovery. | Validate peptide uniqueness, HLA context, matched expression or translation evidence, and targeted peptide confirmation. |
| **Endogenous tagging** | The product can be localized or measured at the native locus when the tag does not disrupt the microprotein. | Native behavior if the tag changes stability, localization, or interactions. | Tag-driven stabilization, altered membrane topology, overlarge epitope effects, edited clone artifacts. | Use small or orthogonal tags, untagged peptide evidence where feasible, ORF-disrupting negative controls, and multiple edited clones. |
| **Conservation and comparative genomics** | Preserved reading frame, codon-level constraint, or syntenic conservation supports evolutionary selection. | Absence of function when conservation is weak, especially for young or lineage-specific ORFs. | Conservation of RNA structure, splicing elements, RBP motifs, or a canonical overlapping coding sequence. | Use codon-aware alignments, synteny, frame-specific tests, and checks for overlapping RNA or protein constraints. |
| **Perturbation and rescue** | ORF disruption plus peptide or ORF-preserving rescue can support causal peptide-mediated function. | Peptide mechanism if the edit also alters RNA abundance, splicing, localization, or regulatory motifs. | Nonsense-mediated decay, disrupted RNA structure, promoter or splice effects, off-target edits, overlapping ORF disruption. | Measure RNA abundance, splicing, localization, and peptide loss; include synonymous controls, alternative-start rescue, and peptide rescue. |

The discovery of smORFs also depends on cell state. A small ORF may be translated only during stress, differentiation, meiosis, embryogenesis, infection, immune activation, or in a specialized subcellular compartment. Standard bulk RNA sequencing or proteomics in proliferating cell lines can miss such events. Conversely, high-depth experiments in stressed cells can reveal many transient translation events that may not represent stable gene products. This cell-state dependence is one reason recent reviews of lncRNA-encoded and circRNA-encoded peptides emphasize validation of expression context, translational mechanism, and function rather than relying on a single genomic screen.

## 77.2. Microprotein biogenesis and function

A microprotein is made by the same basic chemistry as any other protein: a ribosome selects a start codon, charged transfer RNAs deliver amino acids, peptide bonds form in the ribosomal peptidyl transferase center, and translation terminates at a stop codon. What makes microproteins special is not a different decoding chemistry but a different set of constraints. A microprotein has little sequence length in which to encode multiple folded domains, regulatory motifs, and targeting signals. Many microproteins therefore act through compact surfaces, short transmembrane helices, amphipathic helices, intrinsically disordered motifs, or simple binding interfaces that modulate larger complexes.

The first causal step in microprotein biogenesis is initiation. Canonical AUG codons in favorable sequence context are easiest to interpret, but non-AUG initiation can occur at near-cognate codons such as CUG, GUG, UUG, ACG, and others, depending on organism, transcript context, and initiation factor behavior. Non-AUG initiation is harder to annotate because the background frequency of near-cognate codons is high and initiation efficiency is often low. For a candidate microprotein, the start site should ideally be supported by initiation profiling, frame-specific ribosome footprints, mutational loss of translation, and rescue with an ORF-preserving construct.

The second step is elongation through the ORF. Short ORFs have few codons over which to observe triplet periodicity, so statistical confidence is weaker than for long proteins. A 30-codon ORF can yield only a small number of distinct ribosome positions, and many reads may pile up near start or stop sites. Specialized analysis methods use ribosome release scores, frame periodicity, initiation peaks, and comparison with surrounding untranslated sequence to evaluate whether ribosomes behave like elongating ribosomes. Even then, elongation evidence should be described as translation evidence, not as proof of a stable functional product.

The third step is peptide fate. Some microproteins are co-translationally inserted into membranes, especially when they contain a hydrophobic helix. Others remain cytosolic, enter mitochondria, associate with endoplasmic reticulum membranes, localize to the nucleus, or are rapidly degraded by the proteasome. A short peptide can have a half-life measured in minutes, and a low steady-state abundance can make detection difficult. A peptide may also be produced only in a narrow developmental or stress window. Absence from a standard proteomic dataset is therefore weak negative evidence, especially for hydrophobic, basic, low-abundance, or condition-specific microproteins.

Functionally characterized microproteins often regulate larger protein complexes rather than acting as enzymes on their own. A small membrane microprotein can bind a transporter or pump and alter its conformation, localization, or turnover. A short mitochondrial peptide can tune respiratory function or stress responses. A small secreted peptide can act as a signal in development. A nuclear microprotein can modulate transcriptional regulators or RNA-binding proteins. These examples illustrate a general rule: a microprotein does not need a large catalytic domain to matter. It can change biology by occupying a binding surface, altering complex assembly, recruiting degradation machinery, or changing localization.

Microprotein function is usually tested by separating peptide effects from RNA effects. A start-codon mutation that abolishes translation but leaves RNA expression unchanged is informative, but only if the mutation does not disrupt an RNA structure, splice signal, RNA-binding protein motif, microRNA site, or overlapping ORF. A frameshift or premature stop codon can eliminate the peptide, but it may also trigger nonsense-mediated decay if placed in a transcript context recognized by that pathway. A peptide rescue experiment is stronger when the peptide is expressed from a separate transcript and restores the phenotype of an ORF-disrupting mutation. Endogenous tagging can show localization and abundance, but a tag may be larger than the microprotein and may alter its behavior. Strong functional evidence often requires several imperfect experiments that point to the same mechanism.

Reviews of lncRNA-encoded micropeptides emphasize that biogenesis and function must both be documented. A lncRNA-associated smORF may be translated, but the transcript may also function through chromatin binding, RNA-RNA pairing, RNA-protein scaffolding, or transcriptional interference. The presence of a microprotein does not erase RNA-level mechanisms. A single locus can produce both an RNA molecule with regulatory activity and a small peptide with a distinct function. Such dual-function loci are biologically plausible but require careful experimental separation.

## 77.3. Noncanonical ORFs in lncRNAs and circRNAs

![Figure 77.3. Locations of Noncanonical ORFs](../assets/figures/chapter1072_figure3.png)

**Figure 77.3. Locations of Noncanonical ORFs.** Noncanonical ORFs can occur in 5′ untranslated regions, 3′ untranslated regions, overlapping reading frames, downstream ORFs, annotated lncRNAs, circRNAs, pseudogene transcripts, and lineage-specific transcripts. The figure should distinguish transcript class from ORF class and should avoid implying that every depicted ORF is functional.

Long noncoding RNAs are operationally defined transcripts, usually longer than 200 nucleotides, that lack an annotated conventional protein-coding sequence. This definition is not a biochemical guarantee that no ribosome ever translates the transcript. It is a statement about annotation, evidence, and convention. Many lncRNA annotations were built by excluding transcripts with long conserved ORFs, but shorter ORFs remained abundant. As ribosome profiling became common, many annotated lncRNAs were found to associate with ribosomes. The interpretive question is whether that association reflects productive translation, regulatory scanning, experimental contamination, or incomplete transcript models.

One boundary case is a transcript that was annotated as lncRNA because its ORF was too short for historical gene prediction. If ribosome profiling, mass spectrometry, conservation, and genetic evidence all support peptide production and function, reclassification as a protein-coding transcript or dual-function transcript may be appropriate. Another boundary case is a transcript that produces a peptide in one species but lacks a conserved ORF in related species. Such a locus may represent de novo gene birth, lineage-specific adaptation, or weak translation without conserved function. Conservation is powerful when present, but lack of deep conservation is not definitive for short, recently evolved ORFs.

CircRNAs are covalently closed RNA molecules generated when a downstream splice donor is joined to an upstream splice acceptor by back-splicing. Because circRNAs lack a free 5′ cap and a poly(A) tail, they do not fit the standard cap-dependent scanning model. Nevertheless, some circRNAs can be translated when they contain internal ribosome entry elements, chemical modifications or structures that recruit initiation machinery, or engineered translation modules. A circRNA ORF can cross the back-splice junction, which creates a peptide sequence not present in the corresponding linear RNA. In principle, repeated translation around a circular template could yield repeating peptide products, although endogenous examples require stringent validation.

CircRNA translation has specific artifact risks. Linear RNAs from the same locus can contaminate circRNA preparations. Reverse transcription and read alignment can create apparent junction support. Ribosome footprints that cross a back-splice junction are stronger than generic ribosome occupancy on exonic sequence shared with linear transcripts, but junction-spanning reads can still be rare. Proteomic evidence is most convincing when the detected peptide uniquely spans a circRNA-specific junction or otherwise cannot be produced from the linear transcript. Functional experiments should distinguish the circRNA molecule, the back-splice junction, and the peptide product.

Alternative ORFs can also overlap canonical coding sequences. A ribosome may initiate upstream of the main start codon, downstream of it, or in a different frame within the coding region. Overlapping ORFs create annotation challenges because the same nucleotide sequence is constrained by more than one reading frame. If both products are functional, sequence evolution must preserve two protein-coding messages. In other cases, an overlapping ORF may be translated at low levels and serve as a source of antigenic peptides, a regulatory translation event, or an experimental artifact. Standard gene models often represent only the dominant coding sequence, so alternative ORFs can be hidden in plain sight.

The phrase "coding potential of annotated ncRNAs" should therefore be read as a hypothesis about a specific ORF in a specific transcript isoform and biological context. It should not be generalized to all molecules assigned to a gene symbol. A locus may have several transcript isoforms, one protein-coding and one noncoding. A lncRNA name may persist after reannotation. A circRNA may share exons with a host gene that encodes a canonical protein. Careful writing names the exact RNA species, the ORF coordinates, the start codon, the reading frame, the evidence class, and the cell or tissue in which translation was observed.

## Experimental Foundations and Evidence

## 77.4. Ribosome profiling and proteomics evidence

Ribosome profiling, first described as a genome-wide method by Ingolia and colleagues in 2009, relies on a simple physical idea. Translating ribosomes protect a short segment of mRNA from nuclease digestion. If protected fragments are sequenced and mapped to a transcriptome or genome, their positions reveal where ribosomes were located at the time of lysis. Because ribosomes advance codon by codon, high-quality ribosome profiling data often show three-nucleotide periodicity. The protected fragment length and offset between read end and ribosomal A site must be calibrated for each experiment.

For canonical coding sequences, ribosome profiling provides a strong view of translation dynamics. For smORFs, interpretation is more delicate. A short ORF has fewer codons, fewer possible protected fragments, and often lower expression. Initiation and termination peaks can dominate the signal. Reads may map ambiguously if the ORF lies in repetitive sequence, paralogous genes, pseudogenes, or shared exons. Transcript isoforms can make an apparent lncRNA footprint actually arise from an unannotated coding isoform. Nuclease digestion can generate protected fragments from RNA-binding proteins or structured RNA segments that resemble ribosome footprints. These issues do not invalidate ribosome profiling; they define the controls needed for noncanonical ORF annotation.

Strong ribosome profiling evidence for a smORF includes a discrete ORF boundary, reads in the expected frame, reduced signal after translation inhibition controls when appropriate, initiation-site enrichment at a plausible start codon, depletion after start-codon mutation, and reproducibility across biological replicates. Translation initiation profiling can enrich ribosomes at start sites and help distinguish the initiating codon from downstream ribosome positions. Active ribosome capture methods, such as RiboLace, attempt to enrich actively translating ribosomes rather than all ribosome-associated RNA. Streamlined mono- and di-ribosome profiling methods can improve sensitivity in yeast and human cells, while single-cell methods begin to ask how ribosome association varies across cell states. Each method adds evidence but also introduces its own selection biases.

Proteomics asks a different question: can the peptide product be observed? Liquid chromatography tandem mass spectrometry identifies peptides by matching spectra to a database of expected peptide sequences. For microproteins, database construction is critical. If the database includes millions of predicted short ORFs, the false discovery burden increases. If the database excludes noncanonical ORFs, real peptides cannot be found. A balanced search uses controlled candidate sets, target-decoy strategies, unique peptide requirements, retention-time or fragmentation support, and manual inspection for key claims. Targeted proteomics can then test candidate peptides with greater sensitivity.

Short peptides are intrinsically difficult for mass spectrometry. A 35-amino-acid microprotein may yield only a few tryptic peptides, and some may be too short, too long, too hydrophobic, too basic, or shared with other proteins. Membrane microproteins can be lost during extraction. Rapidly degraded peptides may be better detected by immunopeptidomics, where peptides bound to major histocompatibility complex molecules are purified and sequenced, than by whole-cell proteomics. Conversely, immunopeptidomics can prove that a peptide was produced and presented but not that it has an autonomous cellular function.

Protein-level evidence can also come from antibodies, epitope tags, and imaging. Antibodies against microproteins are hard to validate because the antigen is short and may share epitopes. Epitope tags can change localization, stability, and interaction surfaces. Endogenous genome editing is preferable to overexpression because overexpression can force weak ORFs into artificial abundance, bypass native initiation rules, and mislocalize the peptide. The strongest localization claims use an endogenous tag, an ORF-disrupting negative control, peptide rescue, and independent detection where feasible.

The key lesson is that ribosome profiling and proteomics answer complementary questions. Ribosome profiling asks whether ribosomes occupy and likely translate an ORF in a particular context. Proteomics asks whether peptide molecules accumulate or are presented. Neither method alone proves biological function. A claim that "this lncRNA encodes a functional microprotein" should rest on translation evidence, peptide evidence, and perturbation evidence that links the peptide to a phenotype.

**Table 77.2. Method Strengths and Failure Modes.** Ribosome profiling, initiation profiling, active ribosome capture, proteomics, immunopeptidomics, endogenous tagging, reporter assays, and CRISPR start-codon editing each contribute different information and have different failure modes.

| Method | Input material | Main output | Best use | Common artifact | Best companion assay |
| --- | --- | --- | --- | --- | --- |
| **Ribosome profiling** | Nuclease-digested ribosome-protected RNA fragments from cells or tissues. | Mapped footprints, frame periodicity, ORF boundaries, and relative ribosome occupancy. | Nominate translated smORFs and distinguish ORF-specific signal from surrounding untranslated sequence. | Nonribosomal RNP protection, nuclease bias, ambiguous mapping, initiation or termination pileups. | Translation initiation profiling, peptide detection, and ORF-disrupting edits. |
| **Translation initiation profiling** | Lysates enriched for initiating ribosomes, often after initiation-stalling treatments. | Candidate start-site peaks and initiation-codon assignments. | Resolve AUG and near-cognate start codons for short ORFs. | Drug-induced redistribution, secondary peaks, stress-induced initiation changes. | Standard Ribo-seq elongation signal and start-codon mutation. |
| **Active ribosome capture** | Affinity-captured active ribosomes and associated RNA. | Enriched set of actively translated RNA regions. | Reduce passive ribosome association when prioritizing candidate ORFs. | Capture bias, incomplete separation of active and stalled complexes, retained RNP contaminants. | Conventional Ribo-seq, proteomics, and translation-inhibitor controls. |
| **LC-MS/MS proteomics** | Extracted proteins or digested peptides searched against canonical and candidate ORF databases. | Spectra assigned to candidate microprotein peptides. | Demonstrate peptide product formation when unique spectra are available. | Search-space inflation, shared peptides, hydrophobic peptide loss, low-abundance false negatives. | Targeted proteomics, synthetic peptide spectra, and endogenous tagging. |
| **Immunopeptidomics** | Major histocompatibility complex-bound peptides purified from cells or tissues. | Presented peptide sequences and possible noncanonical antigen candidates. | Identify smORF-derived peptides relevant to immune recognition. | Presentation does not prove autonomous function; peptide source and HLA assignment can be ambiguous. | Transcript and Ribo-seq evidence, targeted MS, HLA binding analysis, and T-cell assays when immunogenicity is claimed. |
| **Endogenous tagging and imaging** | Genome-edited cells or organisms with a tag inserted at the native locus. | Product localization, abundance, stability, and sometimes interaction context. | Test whether a candidate microprotein accumulates in the expected compartment. | Tag size or placement can alter localization, turnover, membrane insertion, or binding surfaces. | Untagged peptide detection, ORF-disrupting controls, and rescue assays. |
| **Reporter assays** | Cloned UTRs, ORFs, start codons, or circRNA junctions linked to measurable reporters. | Sequence sufficiency for initiation, repression, reinitiation, or junction-dependent translation. | Dissect uORF or start-codon logic under controlled sequence changes. | Overexpression, non-native RNA structure, altered isoform context, artificial abundance. | Endogenous editing, Ribo-seq, and RNA abundance measurements. |
| **CRISPR start-codon or frame editing** | Endogenous genomic loci edited to disrupt initiation, frame, or stop-codon context. | Phenotype, RNA state, and peptide loss after locus-level perturbation. | Test whether an ORF is required in its native regulatory context. | RNA structural disruption, altered splicing, nonsense-mediated decay, off-target editing. | Peptide rescue, synonymous controls, transcript measurements, and independent edited alleles. |

## 77.5. Evolution and annotation of small coding regions

Evolutionary evidence has always been central to protein-coding annotation. A coding sequence accumulates substitutions under constraints imposed by amino acid sequence, reading frame, start and stop codons, and protein function. Comparative methods can detect codon-level conservation, synonymous versus nonsynonymous substitution patterns, and preservation of an ORF across related species. For long proteins, these signals are strong because there are many codons. For smORFs, the signal is weaker because the sequence is short and chance conservation can be misleading.

Small coding regions are also biologically suited to de novo emergence. A short ORF can appear by mutation in a previously noncoding transcript, become weakly translated, and then acquire function if the peptide provides a selective advantage. The early stages of such evolution may leave little conservation across distant species. Some young smORFs may be species-specific or lineage-specific. Others may be translated neutrally and disappear. This continuum complicates annotation: a candidate smORF can be real in the sense of being translated, but not yet a conserved gene with clear function.

Annotation systems historically biased against smORFs because minimum length thresholds were practical. Very short protein-coding genes are difficult to predict, difficult to validate, and easy to confuse with random ORFs. Modern annotation increasingly incorporates ribosome profiling, proteomics, conservation, transcript models, and manual curation. The best practice is not to force every candidate into a binary label. Instead, annotation can record multiple fields: ORF coordinates, transcript isoform, predicted start codon, translation evidence, peptide evidence, conservation, disease relevance, and functional validation status.

Coding-potential algorithms use features such as ORF length, codon usage, nucleotide composition, sequence conservation, and protein-domain matches. These tools are useful for triage but should not be treated as biological verdicts. A low coding-potential score may reflect a true lncRNA, a recently evolved microprotein, a short disordered peptide, or an ORF expressed only in an unobserved transcript isoform. A high score may reflect a random ORF with codon composition resembling coding sequence or contamination by an unannotated protein-coding transcript. Machine-learning classifiers inherit the biases of their training sets, especially the underrepresentation of short, lineage-specific, non-AUG, and condition-specific ORFs.

Evolutionary evidence is also complicated by overlapping reading frames. A nucleotide change in an overlapping ORF can alter two peptide sequences or one peptide and one RNA regulatory element. Conservation may therefore reflect the canonical protein, an RNA structure, splicing enhancer, microRNA site, or RBP motif rather than the alternative peptide. Conversely, an alternative ORF can be constrained in a way that is hard to detect because most substitutions are already limited by the main coding sequence. Functional testing remains necessary even when evolutionary signals are suggestive.

The annotation of lncRNA-encoded and circRNA-encoded peptides should be conservative in naming. Calling a locus "protein-coding" may obscure a genuine RNA-level function, while calling it "noncoding" may obscure a genuine peptide. A dual annotation, such as "lncRNA locus with translated smORF" or "circRNA-associated ORF with peptide evidence," is often more informative until mechanism is resolved. The goal is not to defend old labels but to describe the molecule, evidence, and function precisely.

**Table 77.3. Annotation Language for Candidate ORFs.** Precise annotation language should state transcript isoform, coordinates, start codon, frame, translation evidence, peptide evidence, conservation, and function instead of using ambiguous phrases such as "the lncRNA is translated."

| Weak wording | Better wording | Reason the better wording is stronger |
| --- | --- | --- |
| **"The lncRNA is translated."** | "Transcript isoform X contains smORF Y, and Ribo-seq in cell state Z shows frame-specific footprints over that ORF." | It names the isoform, ORF, evidence type, and biological context instead of assigning a vague property to an entire locus. |
| **"This ncRNA encodes a protein."** | "The candidate smORF has endogenous peptide evidence and ORF-disrupting/rescue data; RNA-mediated functions of the transcript remain separately annotated." | It separates peptide production and peptide-mediated function from possible RNA-level mechanisms. |
| **"Ribo-seq proves a microprotein."** | "Ribo-seq supports translation of the candidate ORF; peptide detection and functional perturbation remain separate evidence levels." | It avoids treating ribosome occupancy as proof of product stability or biological function. |
| **"No peptide was detected, so the ORF is not real."** | "Standard proteomics did not detect a unique peptide under the sampled conditions; sensitivity may be limited for short, hydrophobic, unstable, or cell-state-specific products." | It records negative evidence without overclaiming from a method with known false negatives. |
| **"The circRNA ORF is translated."** | "Back-splice-junction-specific footprints or peptides support translation from the circular isoform, with controls excluding the linear host transcript." | It distinguishes circular and linear templates, which is the central artifact risk for circRNA peptide claims. |
| **"The variant is noncoding."** | "The 5′ UTR variant creates or strengthens a uORF that may reduce downstream coding-sequence translation without changing mRNA abundance." | It preserves regulatory translation effects that coding-region-only variant language can miss. |
| **"The ORF is not conserved, so it is not functional."** | "The ORF lacks deep conservation; this weakens evolutionary support but does not exclude a young, lineage-specific, or condition-specific microprotein." | It treats conservation as strong positive evidence when present, not as a universal requirement. |
| **"Start-codon editing proves peptide function."** | "Start-codon editing supports peptide-mediated function only if RNA abundance, splicing, localization, and relevant RNA motifs are preserved and peptide rescue restores the phenotype." | It states the controls needed to separate peptide effects from RNA and editing artifacts. |

## Technology, Computational, Clinical, and Engineering Links

## 77.6. Disease and immune relevance

Small ORFs and noncanonical translation matter in disease for at least four reasons. First, uORFs can regulate the amount of canonical protein produced from a disease-associated mRNA. A variant that creates a new uORF, strengthens an upstream start codon, weakens a termination context, or changes reinitiation can reduce translation of the main coding sequence without changing mRNA abundance. Such variants can be missed if clinical interpretation focuses only on missense, nonsense, splice-site, and canonical coding mutations.

Second, microproteins can be disease modifiers or disease drivers. A microprotein that regulates a calcium pump, mitochondrial complex, signaling factor, transcriptional regulator, or metabolic enzyme can affect tissue physiology. In cancer, lncRNA-encoded peptides and circRNA-associated peptides have been proposed to influence proliferation, metabolism, invasion, immune evasion, and therapy response. The chapter references include recent reviews on lncRNA-encoded peptides in cancer and disease-associated circRNA and lncRNA peptides, but disease claims require careful primary-source confirmation because the field contains many correlative studies.

Third, noncanonical ORFs can contribute to immune recognition. The immune system samples intracellular proteins through proteasomal degradation, peptide transport, and presentation by major histocompatibility complex molecules. Peptides derived from alternative ORFs, uORFs, lncRNAs, circRNAs, endogenous retroelements, or stress-induced translation events can become antigens if they are processed and displayed. This possibility is important for cancer immunology because tumor-specific transcription, aberrant splicing, mutation, or epigenetic derepression can create peptides not broadly present in normal tissues. However, immunogenicity requires more than peptide existence: the peptide must be processed, presented by a compatible human leukocyte antigen molecule, recognized by T cells, and not eliminated by tolerance or immune suppression.

Fourth, disease states alter translation landscapes. Stress responses, viral infection, inflammation, hypoxia, nutrient limitation, and oncogenic signaling can change initiation factor activity, ribosome loading, mRNA localization, and proteasome activity. These changes can increase translation of upstream ORFs, non-AUG ORFs, internal ORFs, or normally repressed transcripts. A noncanonical ORF detected in diseased tissue may therefore be a cause, consequence, biomarker, or adaptive response. Experimental design must separate these possibilities.

Immune and inflammatory contexts also raise a distinction between RNA-mediated and peptide-mediated mechanisms. Many lncRNAs regulate immune cell differentiation, cytokine expression, chromatin states, or RNA-binding protein networks without requiring peptide products. Conversely, an lncRNA-associated ORF may produce a peptide that modulates immune signaling. The available chapter references provide general ncRNA immune context, and source-specific microprotein immune mechanisms should be cited at the claim level. For final disease statements, evidence should specify whether the claim is based on patient association, cell-line perturbation, animal model, proteomic detection, immunopeptidomic detection, or clinical response.

> **Box 77.1. How to Write a Precise Claim About an lncRNA-Associated ORF**
>
> A precise claim names the transcript isoform, ORF coordinates, start codon, evidence type, biological context, and remaining uncertainty. The box should show example statements at increasing evidence levels, from sequence candidate to functional microprotein.

Therapeutic design should also account for unintended coding potential. Synthetic circRNAs, engineered lncRNA-like molecules, mRNA therapeutics, and RNA vaccines may contain accidental ORFs or alternative initiation sites if sequence design does not control reading frames. In some cases translation is intended; in other cases it is a safety and interpretation issue. Codon optimization, untranslated region design, chemical modification, RNA structure, and delivery context can all influence whether a short ORF is translated. This topic connects to chapters on RNA therapeutics, vaccine immunobiology, and mRNA design.

## 77.7. Evidence thresholds and false positives

The most useful way to interpret noncanonical ORFs is as an evidence ladder. The lowest rung is sequence plausibility: the RNA contains an in-frame start and stop codon, perhaps with a favorable initiation context. This is necessary for ordinary translation but extremely weak as evidence because short ORFs occur frequently by chance. The next rung is transcript evidence: the ORF lies within an expressed, correctly spliced RNA isoform. This proves that the substrate exists but does not prove translation.

The next rung is ribosome evidence. Ribosome profiling with frame periodicity, start-site support, and ORF-specific signal supports translation more strongly than bulk ribosome association. However, ribosome evidence should be described precisely. "Ribosome occupancy over the candidate ORF" is weaker than "initiating and elongating ribosomes map to this ORF in the expected frame, with signal reduced after start-codon mutation." A short ORF with a handful of reads in one replicate should not be annotated as a functional microprotein.

Peptide evidence is a higher rung but not the top. A unique endogenous mass spectrometry peptide, validated spectrum, targeted confirmation, or endogenous tag can show that a peptide exists. Still, peptide existence does not establish function. Some peptides may be byproducts of pervasive translation and rapid degradation. Immunopeptidomic detection proves presentation, not necessarily autonomous cellular activity.

Functional evidence requires perturbation and rescue. The strongest experiments disrupt peptide production while preserving RNA expression and then restore the phenotype with the peptide or an ORF-preserving allele. Useful controls include synonymous mutations that preserve peptide sequence but alter RNA sequence, start-codon mutations paired with alternative start rescue, stop-codon insertion with nonsense-mediated decay controls, frameshift mutations outside known RNA motifs, and measurement of transcript abundance, localization, and splicing. For lncRNAs and circRNAs, the experiment should test whether the RNA molecule itself remains present and whether the putative peptide is lost.

False positives arise from several sources. Annotation artifacts occur when transcript isoforms are incomplete or when reads from a protein-coding isoform are assigned to a noncoding transcript. Mapping artifacts occur in repeats, pseudogenes, paralogs, and short exons. Ribosome profiling artifacts arise from nuclease bias, RNA structure, nonribosomal RNP protection, drug-induced stalls, run-off during lysis, and insufficient periodicity. Proteomic artifacts arise from inflated search databases, shared peptides, poor spectra, post-translational modifications not included in the search, and weak false discovery control. Functional artifacts arise when ORF-disrupting mutations alter RNA motifs, promoter activity, splicing, mRNA stability, or overlapping regulatory elements.

![Figure 77.4. Artifact Sources in Noncanonical ORF Discovery](../assets/figures/chapter1072_figure4.png)

**Figure 77.4. Artifact Sources in Noncanonical ORF Discovery.** Candidate smORFs can be inflated by incomplete transcript models, ambiguous read mapping, nonribosomal RNP protection, proteomic database expansion, and ORF-disrupting mutations that affect RNA features. The figure should pair each artifact with a practical control.

False negatives are just as important. Small peptides can be missed because they lack tryptic peptides, are membrane embedded, are unstable, are expressed in rare cells, are translated only during stress, or are rapidly processed into antigens. Some smORFs use non-AUG initiation or noncanonical transcript isoforms that standard pipelines ignore. A rigorous chapter on coding potential should therefore avoid both extremes: it should not promote every footprint as a protein, and it should not dismiss all unsupported candidates as impossible.

> **Box 77.2. Why Small Peptides Are Hard to See**
>
> Microproteins are difficult to detect because short sequence length yields few unique proteomic peptides, hydrophobic helices can be lost during extraction, rapid turnover reduces abundance, and cell-state specificity can place the peptide outside the sampled condition.

> **Box 77.3. Dual-Function RNA Loci**
>
> Some loci may produce an RNA molecule with regulatory activity and a peptide with separate activity. The box should explain how start-codon mutations, stop codons, RNA-preserving synonymous edits, and peptide rescue can distinguish mechanisms.

## Recent Consensus

Recent reviews converge on several cautious positions. First, the boundary between coding and noncoding transcript annotation is more porous than older gene catalogs implied. Some annotated lncRNAs and circRNAs contain translated ORFs, and some of those ORFs produce functional peptides. Second, the proportion of transcripts with ribosome association is much larger than the proportion with demonstrated functional microproteins. Third, uORFs are widespread regulatory elements whose main biological effect often comes from translation of the uORF rather than from an independently acting uORF peptide. Fourth, ribosome profiling, proteomics, conservation, and perturbation studies are complementary; no single method is sufficient for all claims.

The consensus is also methodological. Candidate noncanonical ORFs should be reported with genomic coordinates, transcript IDs, start codon, stop codon, frame, peptide sequence if known, expression context, ribosome evidence, peptide evidence, conservation evidence, and functional tests. Claims should avoid ambiguous phrases such as "the lncRNA is translated" when the evidence actually supports only one isoform, one smORF, or one condition. Disease and therapeutic claims require especially careful wording because many studies are based on expression correlations or overexpression systems.

## Open Questions, Controversies, Deprecated Models, and Common Misconceptions

Open questions:

- How much pervasive translation is biologically meaningful? Ribosomes initiate at many weak sites, especially under stress or when initiation factors are perturbed. Some of these events may generate raw material for evolution or immune surveillance without having an immediate cellular function. Others may regulate RNA fate through ribosome movement, collisions, or decay pathways. The field lacks a single threshold that separates functional translation from translational noise across all organisms and cell types.
- How can researchers classify dual-function loci? A transcript may act as an RNA and also encode a peptide. Reclassifying such a locus as simply protein-coding can hide RNA mechanisms; retaining only the noncoding label can hide peptide mechanisms. Future annotations will likely need separate feature-level labels for transcript, ORF, peptide, and function.

Controversies:

- A third controversy concerns circRNA translation. Engineered circRNAs can translate efficiently when designed with strong initiation elements, and some endogenous circRNAs have credible translation evidence. Yet many claimed endogenous circRNA peptides remain vulnerable to linear RNA contamination, weak peptide evidence, or overexpression artifacts. Strong circRNA translation claims should include junction-specific translation or peptide evidence and controls for the corresponding linear transcript.

Common misconceptions:

- "Noncoding RNA means never translated." Noncoding is an annotation category, not a physical prohibition against ribosome association.
- "Ribosome profiling proves a functional protein." Ribosome profiling supports translation, especially with periodicity and initiation evidence, but function requires peptide-level and perturbation evidence.
- "No mass spectrometry peptide means no microprotein." Standard proteomics has limited sensitivity for small, hydrophobic, unstable, condition-specific, or low-abundance peptides.
- "Any peptide from an lncRNA disproves RNA function." A locus can have RNA-mediated and peptide-mediated functions, and experiments must separate them.
- "Conservation is required for all real microproteins." Conservation is strong evidence when present, but young or lineage-specific smORFs may be real and functional without deep conservation.
