Chapter 117. Viral RNA Structures, Regulatory Elements, Translation, and Immune Evasion

Scope Note

This chapter explains how RNA-virus genomes contain structured elements that direct replication, transcription, translation, packaging, and immune evasion. It covers RNA-virus untranslated regions, frameshift elements, internal ribosome entry sites, packaging signals, subgenomic RNAs, host shutoff, and structure-probing evidence. The ownership is explicitly RNA-virus and retroviral context: DNA-virus alternative processing, noncoding transcripts, latency RNAs, and host shutoff belong to Chapter 112. This chapter bridges genome strategies (Chapter 115) with polymerase enzymology (Chapter 116), condensation (Chapter 118), innate sensing (Chapter 108), and retroviral mechanisms (Chapter 120).

Executive Summary

The genomes of RNA viruses are compact, often encoding fewer than a dozen proteins, yet they orchestrate replication, transcription, translation, assembly, and immune evasion with remarkable efficiency. This efficiency depends on structured RNA elements that function as regulatory modules embedded in the genome. The 5′ and 3′ untranslated regions contain promoters, replication enhancers, translation initiation elements, and genome cyclization signals. Internal ribosome entry sites allow cap-independent translation initiation, reprogramming the host translation machinery to favor viral messages. Programmed ribosomal frameshifting expands coding capacity by altering the reading frame at defined positions, producing two proteins from one open reading frame at a fixed ratio. Pseudoknots, stem-loops, kissing loops, and higher-order architectures create binding sites for viral and host proteins, thermodynamic switches that sense and respond to the viral life cycle, and surfaces that activate or evade innate immune sensors.

Packaging signals, often embedded in structured regions of the genome, ensure that full-length viral RNA, and not cellular mRNA, is selectively encapsidated into assembling virions. In retroviruses, genome dimerization through a kissing-loop mechanism brings two copies of the genome together in the virion. Subgenomic RNAs allow ordered, temporally controlled expression of structural and accessory genes from a single genomic RNA template. Noncoding viral RNAs, including viral microRNAs and long noncoding RNAs, modulate host and viral gene expression. Host shutoff mechanisms, executed by viral proteases, RNases, and cap-snatching endonucleases, degrade or render nonfunctional host mRNAs, redirecting the translation apparatus to viral RNAs.

Viruses evade innate immune detection through several RNA-centered strategies. Double-stranded RNA replication intermediates are sequestered inside membrane-bound replication organelles. The 5′ end of viral RNA can be capped by virally encoded capping enzymes or by cap-snatching, mimicking host mRNA and avoiding detection by RIG-I and IFIT proteins. Internal RNA modifications, such as 2′-O-methylation of the cap, further reduce innate sensor activation. Viral noncoding RNAs and structured elements can act as decoys that bind and inhibit PKR, RIG-I, and OAS, preventing their activation by viral RNA. At the same time, these very structures serve as targets for antiviral discovery: structure probing defines druggable pockets, small molecules that bind structured viral RNA can block function, and antisense oligonucleotides designed against conserved structural elements can inhibit replication.

This chapter integrates structural, biochemical, genetic, and virological evidence to build a unified picture of how viral RNA genomes function as regulatory machines. The chapter emphasizes mechanistic diversity across virus families, experimental evidence standards, and the distinction between established mechanisms and hypotheses still under active investigation.

Concept Inventory

  • Internal ribosome entry site (IRES): a structured RNA element, typically in the 5′ UTR, that recruits ribosomes for translation initiation without requiring a 5′ cap or cap-dependent initiation factors. Viral IRESs are classified into several types distinguished by their RNA structure, required initiation factors, and ribosome recruitment mechanism. The classification is based on phylogeny rather than a continuous variable, and some IRES types require different complements of canonical translation initiation factors.
  • Programmed ribosomal frameshifting: a translation mechanism in which the ribosome shifts reading frame at a defined position in the mRNA with a characteristic efficiency, typically between 1 and 50 percent. The shift is directed by a slippery sequence (often a heptanucleotide of the form X XXY YYZ) followed by a stimulatory RNA structure, most commonly a pseudoknot. Frameshifting produces two proteins from one open reading frame: the product of the zero-frame translation and the product of the shifted frame.
  • Pseudoknot: an RNA tertiary structure formed when nucleotides in the loop of a stem-loop base-pair with a sequence outside the loop, either upstream or downstream. Pseudoknots are distinguished from simple stem-loops by the intercalation of two helical stems, creating a more compact and often more stable structure. In viral RNAs, pseudoknots function as frameshift stimulators, translation regulators, replication signals, and protein-binding platforms.
  • Packaging signal (psi: a structured RNA element that directs selective incorporation of full-length viral genomic RNA into assembling virions. Packaging signals can be located at the 5′ end of the genome (as in many retroviruses), at multiple sites along the genome (as proposed for some positive-strand RNA viruses), or associated with specific RNA structures in the untranslated regions. The term “packaging signal” originally referred to the HIV-1 psi element, but the concept now applies across many virus families despite structural and mechanistic differences.
  • Genome dimerization: the noncovalent association of two copies of the genomic RNA into a dimer, which is then packaged into retroviral virions. Dimerization is initiated by the dimerization initiation site, a conserved stem-loop near the 5′ end of the genome that forms a kissing-loop complex with the corresponding stem-loop on a second genomic RNA copy. Dimerization is mechanistically coupled to packaging and is essential for the production of infectious retroviral particles.
  • Subgenomic RNA (sgRNA): a viral mRNA produced by internal transcription initiation from a negative-sense antigenomic template (in positive-strand RNA viruses that use sgRNA transcription) or by discontinuous synthesis that fuses a leader sequence to the body of an mRNA (in coronaviruses and other nidoviruses). Subgenomic RNAs permit temporally and quantitatively regulated expression of structural and accessory genes from a single genomic template.
  • Transcription regulatory sequence (TRS): a conserved sequence motif in coronavirus genomes that directs the discontinuous synthesis of subgenomic mRNAs. The leader TRS at the 5′ end of the genome base-pairs with body TRSs located upstream of each gene, enabling the fusion of the common leader sequence to each subgenomic mRNA body during negative-strand RNA synthesis.
  • Cap-snatching: a mechanism by which segmented negative-strand RNA viruses, including influenza virus and bunyaviruses, acquire short 5′-capped oligonucleotides from host pre-mRNAs or mRNAs and use them as primers for viral mRNA synthesis. The viral polymerase carries an endonuclease domain that cleaves host transcripts, and the resulting capped fragments prime transcription of viral mRNAs.
  • Host shutoff: a virus-induced global suppression of host gene expression, executed through degradation of host mRNA, inhibition of host transcription, interference with host mRNA processing and export, cleavage of translation initiation factors, or competition for ribosomes. Host shutoff redirects cellular resources to viral gene expression and suppresses the innate immune response by preventing synthesis of interferon and interferon-stimulated gene products.
  • Cis-acting replication element (cre): an RNA structure within the viral genome that is required in cis for RNA replication, typically functioning as a promoter, enhancer, or recognition site for the viral replication complex. The term originally referred to a specific stem-loop in picornavirus genomes that serves as a template for VPg uridylylation, but it now encompasses a wider range of structured replication signals.
  • Minus-strand synthesis promoter: the RNA element at or near the 3′ end of a positive-strand RNA virus genome that is recognized by the viral RNA-dependent RNA polymerase to initiate synthesis of the complementary negative-strand RNA. The promoter typically includes sequences and structures near the 3′ terminus, and its activity can be regulated by genome cyclization interactions with the 5′ end.
  • RIG-I ligand: a viral RNA structure or species that is recognized by RIG-I (retinoic acid-inducible gene I), a cytosolic innate immune receptor. RIG-I recognizes short double-stranded RNA with a 5′ triphosphate or diphosphate and a blunt end, features typical of many viral genomes and replication intermediates. MDA5 ligand is longer double-stranded RNA recognized by MDA5 (melanoma differentiation-associated protein 5). The two receptors cooperate to sense different RNA virus families.
  • PKR activation by viral dsRNA: occurs when the double-stranded RNA-activated protein kinase PKR binds double-stranded RNA longer than approximately 30 base pairs, dimerizes, autophosphorylates, and then phosphorylates the translation initiation factor eIF2-alpha, shutting down translation in the infected cell. Many viruses have evolved structured RNA antagonists of PKR.

What to Know Before Reading This Chapter

The reader should be comfortable with basic RNA structure concepts: base pairing, stem-loops (hairpins), internal loops, bulges, junctions, and the distinction between secondary structure (base-pairing pattern) and tertiary structure (three-dimensional arrangement). The reader does not need to be an expert in RNA folding algorithms, but they should understand that RNA structure is often dynamic — an RNA can sample multiple conformations — and that structure can be probed experimentally using chemical reagents that react differently with paired and unpaired nucleotides.

This chapter assumes familiarity with the general replication cycle of RNA viruses: entry, translation, replication, transcription, assembly, and release. The genome strategies chapter (Chapter 115) provides the foundation. The reader should understand that positive-strand RNA virus genomes can serve directly as mRNA upon entry, that negative-strand RNA viruses must carry their polymerase into the cell, and that retroviruses reverse-transcribe their RNA genome into DNA before integration. The reader should also know that RNA-dependent RNA polymerases (RdRPs) operate on RNA templates and that their fidelity, processivity, and template specificity are covered in Chapter 116.

The concept of a translation initiation factor (eIFs) is essential for understanding IRESs and host shutoff. The reader should know that canonical cap-dependent translation initiation begins with recognition of the 5′ cap by eIF4E, followed by assembly of the eIF4F complex, recruitment of the 43S preinitiation complex, and scanning to the AUG start codon. IRES-mediated initiation bypasses some or all of these steps.

The reader should distinguish between a “structure” (a specific RNA fold) and a “structural element” (a functionally characterized region that acts through a structure). Not every predicted stem-loop in a viral genome is functional; some RNA structures are functionally validated by mutagenesis, compensatory mutation analysis, chemical probing, and functional assays, while others are computational predictions awaiting experimental confirmation.

Innate immune sensing of RNA is treated comprehensively in Chapter 108. This chapter discusses immune evasion mechanisms that depend on viral RNA structures and modifications. The reader should know that RIG-I, MDA5, TLR3, TLR7, TLR8, PKR, OAS, and IFIT proteins are the principal RNA-sensing and response pathways, even if the molecular details are reviewed in the immune chapters.

117.1. UTRs, cis-acting replication elements, and structured regulatory RNAs

The 5′ and 3′ untranslated regions of RNA virus genomes are not inert spacers. They are densely populated with RNA structures that control translation, regulate replication, protect the RNA termini, and mediate interactions with host and viral proteins. The UTRs can comprise 5 to 15 percent or more of the genome length and are often the most conserved regions across isolates of a given virus species, reflecting the structural and functional constraints on these sequences.

The 5′ UTR of picornaviruses, exemplified by poliovirus, illustrates the density of regulatory information in a relatively short region. In poliovirus, the 5′ UTR is approximately 740 nucleotides and contains a cloverleaf structure at the extreme 5′ end, an internal ribosome entry site of type 1 (described in Section 117.2), and a spacer region. The cloverleaf is a conserved cruciform RNA structure that binds the viral protein 3CD (the protease-polymerase precursor) and the host poly(rC)-binding protein PCBP2. This ribonucleoprotein complex is required not for translation but for the initiation of negative-strand RNA synthesis. The cloverleaf serves as a replication promoter by recruiting the viral replication complex to the 5′ end, after which genome circularization through a protein bridge (mediated by the host poly(A)-binding protein PABP and viral protein 3AB) brings the polymerase to the 3′ end for initiation. The functional separation of the 5′ UTR into a replication-dedicated 5′ terminal structure and a translation-dedicated IRES is a recurring architectural theme in positive-strand RNA viruses.

The 3′ UTR of positive-strand RNA viruses contains the minus-strand synthesis promoter. In flaviviruses, the 3′ UTR is unusually long and structured: for dengue virus, it extends approximately 450 nucleotides and contains conserved stem-loops, dumbbell structures, and a terminal 3′ stem-loop that forms the actual promoter. Subgenomic flavivirus RNAs (sfRNAs), discussed in Section 117.4, are generated by incomplete degradation of the 3′ UTR by the host 5′-to-3′ exonuclease XRN1, which stalls at a pseudoknot structure. The sfRNA functions as a decoy for innate immune sensors and as a regulator of viral replication. The 3′ UTR, therefore, encodes both a replication promoter and a noncoding RNA that modulates the host response.

Genome cyclization is a critical regulatory mechanism in flaviviruses. Complementary sequences near the 5′ and 3′ ends of the genome base-pair to form a panhandle structure that brings the ends together. In dengue virus, the 5′ upstream AUG region (5′ UAR), the 5′ cyclization sequence (5′ CS), and the 3′ cyclization sequence (3′ CS) form approximately 10 to 15 base pairs of complementarity. Cyclization is required for negative-strand synthesis because genome circularization positions the viral RdRP, which binds near the 5′ end, at the 3′ terminus for initiation. Mutations that disrupt cyclization reduce or abolish replication, and compensatory mutations that restore base pairing restore replication. Genome cyclization is not unique to flaviviruses; picornaviruses, coronaviruses, and other positive-strand RNA viruses use protein-bridged or RNA-RNA circularization to coordinate 5′- and 3′-end functions.

Cis-acting replication elements (cre) were first characterized in picornaviruses. The picornavirus cre is a conserved stem-loop, typically located in the coding region of a nonstructural protein rather than in the UTRs, that serves as the template for uridylylation of the viral protein VPg (3B). VPg-pUpU, the uridylylated form, functions as the protein primer for both positive-strand and negative-strand RNA synthesis. The RdRP (3D) binds the cre stem-loop and uses the first adenosine of a conserved AAACA motif in the loop as the template for adding two uridine residues to VPg. Mutation of the cre AAACA motif abolishes VPg uridylylation and genome replication, even when the mutant cre is moved to a different location in the genome, demonstrating that the cre acts in cis but its genomic position is flexible within limits.

Figure 117.1. Regulatory Architecture of Viral Untranslated Regions and Cis-Acting Replication Elements

Figure 117.1. Regulatory Architecture of Viral Untranslated Regions and Cis-Acting Replication Elements. The untranslated regions of RNA virus genomes encode replication promoters, translation elements, cyclization signals, and immune decoy RNAs. Each family has solved the same regulatory problems with structurally distinct solutions.

The 5′ and 3′ UTRs of negative-strand RNA viruses are structured differently because the genomic RNA is not translated. Instead, the 3′ and 5′ termini of both genomic (negative-sense) and antigenomic (positive-sense) RNA contain promoters for the viral RdRP. In influenza virus, the 5′ and 3′ termini of each of the eight genomic RNA segments are partially complementary and form a partially double-stranded panhandle structure that is the promoter for the viral polymerase. The promoter is bound by the PB1 subunit via the 5′ end and by the PB2 and PA subunits, and this structure is conserved across influenza A, B, and C viruses. The partially double-stranded nature of the promoter, with the 5′ end forming a hook structure that inserts into a pocket of the polymerase, is essential for transcription and replication and has made the promoter an attractive target for antiviral design.

The 3′ UTR of hepatitis C virus provides another example of structured regulation. The HCV 3′ UTR comprises a variable region, a poly(U/UC) tract, and a highly conserved 3′ terminal region called the 3′X region, which contains three stem-loops (SL1, SL2, and SL3). SL2 is a kissing-loop interaction with another region of the genome that contributes to genome circularization. The 3′X region is essential for replication, and its structure, rather than its exact sequence in some positions, is critical: mutations that disrupt base pairing in SL2 and SL3 reduce replication, and compensatory mutations that restore the structure rescue replication, demonstrating that the RNA secondary structure itself carries the functional information.

A common misconception is that UTR sequences that are conserved across isolates must encode functional RNA structures. While conservation is a useful clue, it can also reflect the recent common ancestry of the isolates rather than structural constraint, and some conserved sequences are conserved as binding sites for proteins rather than for their RNA structure. The rigorous test for a functional RNA structure is compensatory mutagenesis: mutations that disrupt the predicted base pairing should reduce function, and additional mutations that restore base pairing (with a different nucleotide sequence) should restore function. This kind of evidence is available for the flavivirus cyclization sequences, the picornavirus cre, the HCV 3′X region, and the influenza promoter, but for many computationally predicted structures in viral UTRs, such evidence is incomplete.

Boundary cases exist where a sequence functions both as a coding region and as a structural element. The picornavirus cre is the most dramatic example: the AC-rich motif that serves as the VPg uridylylation template is embedded in the reading frame of a nonstructural protein. Mutations in the cre that change the amino acid sequence of the protein without altering the RNA structure can be tolerated, demonstrating that the RNA structure rather than the protein sequence is the source of the cre function. But such dual-function regions impose constraints: the sequence must simultaneously encode a functional protein domain and form an RNA structure, which is thought to limit the evolutionary rate of these regions relative to purely coding or purely structural regions.

The existing reference list does not provide specific primary references for these topics.

117.2. Frameshift elements, IRESs, pseudoknots, and translation control

RNA viruses use noncanonical translation mechanisms to overcome the constraints of their compact genomes and to outcompete host mRNAs for the translation machinery. The two most important strategies are internal ribosome entry, which enables cap-independent translation initiation, and programmed ribosomal frameshifting, which allows two proteins to be produced from one open reading frame at a defined ratio. Both strategies depend on precisely folded RNA structures whose disruption abolishes function.

Internal ribosome entry sites are structured RNA regions, usually 150 to 450 nucleotides in length, that recruit ribosomes to internal positions on the mRNA without requiring a 5′ cap or the full set of canonical initiation factors. The discovery of IRES elements came from studies of poliovirus and encephalomyocarditis virus (EMCV) in the late 1980s, when it was shown that the uncapped picornavirus RNA was translated efficiently in infected cells while host cap-dependent translation was inhibited. Bicistronic reporter assays, in which two open reading frames are separated by the putative IRES, provided the experimental standard: if the second cistron is translated, the intervening sequence must direct internal ribosome entry.

IRES elements are classified into four broad types based on their RNA structure, the set of initiation factors they require, and their phylogenetic distribution. Type 1 IRESs (poliovirus, rhinovirus) require the full set of canonical initiation factors except eIF4E. They are approximately 450 nucleotides long, have a Y-shaped secondary structure with multiple domains, and their activity depends on several RNA-binding proteins, including PCBP2, the polypyrimidine tract-binding protein PTB, and the Lupus La autoantigen. Type 2 IRESs (EMCV, foot-and-mouth disease virus) are shorter, approximately 450 nucleotides, and are active with fewer factors. Type 3 IRESs (hepatitis A virus) are even shorter and have a simpler structure. Type 4 IRESs, found in the Dicistroviridae (including cricket paralysis virus CrPV), are approximately 190 nucleotides and initiate translation without any initiation factors, directly binding the ribosome, occupying the P site with a pseudoknot structure that mimics initiator tRNA, and bypassing the need for initiator Met-tRNAi entirely. The CrPV IRES, determined by cryo-EM at high resolution, showed that the IRES RNA adopts a structure in the ribosomal decoding center that mimics a tRNA-mRNA complex, making it the most minimal translation initiation machinery known.

The hepatitis C virus IRES, which does not fit neatly into the picornavirus typology, has been characterized extensively by structural biology and biochemistry. The HCV IRES is approximately 340 nucleotides and forms a structure with three major domains (II, III, and IV). Domain III contains a pseudoknot-rich core that binds the 40S ribosomal subunit directly, positioning the AUG start codon in the mRNA-binding channel without scanning. Domain II induces a conformational change in the 40S subunit that positions the mRNA correctly for eIF2-GTP-Met-tRNAi binding. The HCV IRES requires eIF2, eIF3, and eIF5 but not eIF4E, eIF4G, eIF4A, or eIF1, distinguishing it from cap-dependent initiation and from type 1 picornavirus IRESs. The HCV IRES has been a leading target for small-molecule antiviral development, with compounds identified that bind domain IIa and lock the IRES in a conformation incompatible with 40S subunit binding.

Figure 117.2. IRES Types, Frameshift Elements, and Noncanonical Translation Control

Figure 117.2. IRES Types, Frameshift Elements, and Noncanonical Translation Control. Internal ribosome entry sites recruit ribosomes without a 5′ cap, with structural complexity inversely related to initiation factor dependence. Frameshift pseudoknots mechanically resist ribosome translocation to direct a fixed fraction of ribosomes into an alternative reading frame.

Programmed ribosomal frameshifting is the second major translation control mechanism. In retroviruses, including HIV-1 and Rous sarcoma virus, the Gag and Gag-Pol polyproteins are translated from the same mRNA. Translation of the Gag open reading frame terminates at a stop codon, but at a frequency of approximately 5 percent, the ribosome shifts one nucleotide backward (a -1 frameshift) before reaching the stop codon and continues translating in the Gag-Pol reading frame. The result is a fixed ratio of Gag (structural proteins) to Gag-Pol (enzymatic proteins including protease, reverse transcriptase, and integrase) of approximately 20 to 1. This ratio is essential for virion morphogenesis: too little Gag-Pol produces noninfectious particles lacking enzymes, and too much Gag-Pol is toxic and prevents assembly.

The frameshift signal consists of two elements: a slippery heptanucleotide sequence (U UUU UUA in HIV-1, where the reading frame is indicated by the triplet spacing) followed by a stimulatory RNA structure separated by a spacer of 5 to 9 nucleotides. The stimulatory structure is most commonly a pseudoknot in retroviruses and in many other RNA viruses. The slippage mechanism is understood in outline: during translocation, the ribosome encounters the stimulatory structure, which resists unwinding and causes the ribosome to pause. During the pause, the peptidyl-tRNA and aminoacyl-tRNA in the P and A sites can re-pair with the mRNA in the -1 frame because of the homopolymeric character of the slippery sequence. When translocation resumes, the ribosome continues in the new frame. The pseudoknot is thought to resist unwinding through its topology — the intercalation of two stems creates a more compact structure than a simple stem-loop — though the exact mechanical pathway from pseudoknot resistance to frameshift is debated. Single-molecule force spectroscopy, optical tweezer studies, and mutational analyses have shown that the mechanical stability of the pseudoknot, rather than its thermodynamic stability, correlates best with frameshift efficiency.

Coronaviruses use a similar -1 frameshift mechanism to produce the replicase polyproteins pp1a and pp1ab. The SARS-CoV-2 frameshift signal includes a slippery sequence (U UUA AAC) and a three-stemmed pseudoknot. Frameshift efficiency in SARS-CoV-2 is approximately 20 to 45 percent, higher than the 5 percent observed in HIV-1, producing a lower Gag-Pol-to-Gag ratio shift that is equivalent to a higher replicase-to-structural-protein ratio. The SARS-CoV-2 frameshift pseudoknot has been determined by cryo-EM in complex with the translating ribosome, showing how the pseudoknot contacts the ribosomal entry tunnel and resists helicase-mediated unwinding. This structure has defined the frameshift site as a target for antiviral compounds that increase or decrease frameshift efficiency, either of which could reduce viral fitness.

Other translation control mechanisms used by RNA viruses include termination-reinitiation (where ribosomes translate a short upstream open reading frame, terminate, and then reinitiate at a downstream AUG), leaky scanning (where some ribosomes bypass the first AUG and initiate at a downstream AUG), and non-AUG initiation. Influenza virus uses a ribosomal scanning mechanism but steals capped primers from host mRNAs through cap-snatching, described in Section 117.4. The caliciviruses, including norovirus, have a genome-linked VPg protein at the 5′ end that interacts with translation initiation factors and functions analogously to a cap, and this mechanism blends cap-dependent and cap-independent strategies.

A common misconception is that all structured RNA regions in viral genomes must influence translation directly. Many structures in viral coding regions have no translational role; they may influence RNA stability, protect against nucleases, serve as packaging signals, or be nonfunctional folds maintained by genetic drift. The standard for proving translational function is the demonstration that mutations in the structure change the efficiency of translation (or frameshift) of a reporter and that compensatory mutations restore it.

Another misconception is that IRES activity can be inferred from the presence of a long, structured 5′ UTR. While many IRESs are in structured 5′ UTRs, the structure alone does not prove IRES function. The bicistronic reporter assay remains the standard, but this assay can produce false positives if cryptic promoter activity, splicing, or RNA cleavage generates monocistronic messages. Best practice requires demonstrating that the second cistron is translated when the first cistron is efficiently translated and that the IRES functions in the context of the full-length viral genome.

The existing references include HIV regulatory reviews (Pavlakis 1990; Lesnik 2002) and do not cover the frameshift or IRES literature.

117.3. Packaging signals, genome dimerization, and assembly cues

Packaging is the process by which viral genomic RNA is selectively incorporated into assembling virions. The packaging problem is nontrivial: an infected cell contains a large excess of cellular RNAs — mRNA, rRNA, tRNA, and noncoding RNAs — and the virus must distinguish its genome from this background with high selectivity. For most RNA viruses, the solution is a packaging signal: one or more structured RNA elements in the genome that bind the viral structural proteins (nucleocapsid protein, capsid protein, or Gag polyprotein) with sufficient affinity and specificity to direct encapsidation.

The best-characterized packaging signal is the HIV-1 psi (Ψ) element, located in the 5′ UTR, downstream of the primer binding site and upstream of the Gag start codon. In HIV-1, the full-length genomic RNA is also the mRNA for Gag and Gag-Pol. Unspliced genomic RNA is exported from the nucleus by the Rev-Rev response element (RRE) pathway. Once in the cytoplasm, two copies of the genomic RNA dimerize through the dimerization initiation site (DIS), a conserved stem-loop near the 5′ end of the genome. The DIS forms a kissing-loop complex: six nucleotides in the loop of the DIS stem-loop (the palindromic sequence GCGCGC in HIV-1 subtype B, or GUGCAC in subtype C) base-pair with the complementary sequence on a second genomic RNA. The kissing-loop complex is then converted into a more stable extended duplex by the nucleocapsid domain of Gag, which acts as an RNA chaperone to resolve the initial loop-loop interaction into a linear intermolecular duplex.

The HIV-1 psi element spans several stem-loops (SL1 through SL4), with SL1 containing the DIS, SL2 containing the major splice donor site, SL3 containing the core psi element recognized by the Gag nucleocapsid domain, and SL4 encoding the start codon for Gag. The Gag polyprotein binds psi through its nucleocapsid domain, which contains two CCHC zinc finger motifs that recognize the structure and sequence of SL3 and flanking regions. Mutations in the zinc fingers or in the psi stem-loops reduce or abolish selective packaging of genomic RNA, leading to the production of noninfectious particles filled with cellular RNA. The specificity of Gag for genomic RNA over cellular mRNA is determined by the structured psi ensemble: correctly folded psi presents a high-affinity surface that is not found in cellular mRNAs, which are less structured in their 5′ UTRs and lack the DIS-palindromic sequence that drives dimerization.

Temporally regulated structural transitions in the HIV-1 5′ UTR control the balance between translation, dimerization, and packaging. The 5′ UTR can adopt at least two alternative conformations: a monomeric, translation-competent conformation (in which the DIS is occluded) and a dimeric, packaging-competent conformation (in which the DIS and psi are exposed). The equilibrium between these conformations is sensitive to the concentration of Gag, suggesting that an increase in Gag concentration late in infection shifts the population of genomic RNAs from translation to packaging. This conformational switch is a general principle: the same RNA sequence can encode different functional states through alternative folding, and viral structural proteins or host factors can drive the equilibrium toward the state needed at a given life cycle stage.

Figure 117.3. HIV-1 5′ UTR Structural Ensemble: Dimerization, Packaging, and Conformational Switching

Figure 117.3. HIV-1 5′ UTR Structural Ensemble: Dimerization, Packaging, and Conformational Switching. The HIV-1 5′ UTR encodes alternative functional states through conformational switching. In the monomer state the RNA is translated; in the dimer state the DIS mediates intermolecular base pairing and Gag binds psi to direct selective encapsidation.

Selective packaging of genomic RNA over subgenomic RNA is a related challenge. In viruses that produce subgenomic RNAs, such as alphaviruses (Sindbis virus, Semliki Forest virus) and coronaviruses, only the full-length genomic RNA is packaged, even though subgenomic RNAs share the 3′ terminus. In alphaviruses, a packaging signal has been mapped to a region of the nonstructural protein-coding sequence that is absent from the subgenomic RNA because subgenomic transcription starts downstream. In coronaviruses, a packaging signal has been identified in the nonstructural protein-coding region near the 5′ end, and the subgenomic RNAs, which carry a common leader but not this packaging signal, are excluded from virions. Recently, the concept of multiple dispersed packaging signals across the genome has been proposed for some viruses, including SARS-CoV-2, where multiple short sequence motifs may collectively specify packaging.

In negative-strand RNA viruses, packaging is coupled to genome replication. The genomic RNA is never free in the cytoplasm; it is always coated with nucleoprotein as part of a helical ribonucleoprotein complex. Packaging signals for negative-strand RNA viruses are less well-defined than for retroviruses but are thought to reside in the terminal promoter regions that are recognized by the polymerase and nucleoprotein during encapsidation. In influenza virus, the packaging of eight distinct genomic segments into one virion depends on segment-specific packaging signals at the 5′ and 3′ ends of each segment, and some segments interact with one another during assembly, ensuring that the full complement of segments is packaged.

The distinction between a packaging signal and a replication signal is not always sharp. In HIV-1, the psi element is near the 5′ replication signals (primer binding site, 5′ terminal repeat), and the TAR element (transactivation response element) at the extreme 5′ end participates in both transcription and in some aspects of packaging. In picornaviruses and flaviviruses, the 5′ UTR structures that direct replication may also contribute to packaging, and the boundary between replication and packaging functions can be difficult to resolve because both processes depend on the same regions.

A common misconception is that every retrovirus packages two genomic RNA copies as a matter of design for diploidy-driven recombination. While recombination is a consequence of having two genomes in the virion, and template switching between the two copies can rescue damaged genomes during reverse transcription, the primary pressure for dimeric packaging may be the need to package two reverse transcriptase molecules — one associated with each genomic RNA — to initiate reverse transcription efficiently. The two-genome strategy may also provide redundancy against RNA damage and a mechanism for recombination-mediated repair.

The existing references include the HIV RRE (Lesnik 2002), HIV regulation (Pavlakis 1990), and HIV RNA structure/immune papers (Hughes 2026, Stunnenberg 2021), but they do not provide the core packaging and dimerization references.

117.4. Subgenomic RNAs, noncoding viral RNAs, and host shutoff

Subgenomic RNAs expand the coding capacity of RNA virus genomes by enabling ordered expression of multiple proteins from a single genomic template. The mechanisms by which subgenomic RNAs are produced differ among virus families, but all achieve the same outcome: structural and accessory proteins are produced from shorter, often more abundant mRNAs, while the genomic RNA is preserved for replication and packaging.

In the order Nidovirales, which includes coronaviruses (Coronaviridae) and arteriviruses, subgenomic mRNA synthesis occurs by discontinuous transcription. During negative-strand RNA synthesis from the positive-sense genomic template, the viral RdRP pauses at body transcription regulatory sequences (TRSs) located upstream of each open reading frame. The nascent negative-strand RNA dissociates from the template and re-anneals to the leader TRS at the 5′ end of the genome through base pairing between the body TRS complement and the leader TRS. This discontinuous step fuses the common leader sequence, which includes the 5′ cap and a short upstream open reading frame, to the body of the subgenomic mRNA. The result is a nested set of 3′-coterminal subgenomic mRNAs, each with an identical 5′ leader and a variable-length body encoding one or more genes. The relative abundance of each subgenomic mRNA is determined by the efficiency of TRS recognition, the distance from the 5′ end, and other factors.

In coronaviruses, the TRS core sequence is conserved (typically ACGAAC in SARS-CoV-2), and the degree of complementarity between the leader TRS and each body TRS influences the abundance of the corresponding subgenomic mRNA. The nucleocapsid protein is produced from the shortest and most abundant subgenomic mRNA, while the spike protein is produced from a longer, less abundant subgenomic mRNA. This transcriptional gradient allows the virus to tune the stoichiometry of structural proteins for efficient assembly.

In alphaviruses (Togaviridae), subgenomic RNA synthesis is simpler: the viral RdRP initiates transcription at an internal subgenomic promoter located in the negative-strand antigenomic template, between the nonstructural and structural open reading frames. The resulting subgenomic mRNA (26S RNA in Sindbis virus and Semliki Forest virus) is produced in molar excess over the genomic RNA (49S RNA), ensuring high expression of the structural proteins required for assembly. The subgenomic promoter is a conserved RNA sequence and structure in the negative strand that is recognized by the viral replication complex.

In rhabdoviruses (e.g., vesicular stomatitis virus) and paramyxoviruses (e.g., Sendai virus), subgenomic mRNA synthesis occurs by a start-stop transcription mechanism. The viral RdRP enters the genome at a single 3′ promoter and transcribes each gene sequentially, pausing at gene-end signals to release the mRNA and reinitiating at gene-start signals for the next gene. The result is a transcriptional gradient: the 3′-proximal genes (typically N, P) are transcribed at higher levels than the 5′-proximal genes (L polymerase), because polymerase drop-off occurs at each gene junction. This gradient provides another solution to the stoichiometry problem.

Noncoding viral RNAs extend beyond subgenomic mRNAs. Flaviviruses produce a subgenomic flavivirus RNA (sfRNA) as described in Section 117.1. The sfRNA, approximately 300 to 500 nucleotides long, is generated when the host 5′-to-3′ exonuclease XRN1 stalls at a pseudoknot structure in the 3′ UTR during degradation. The sfRNA accumulates to high levels in infected cells and functions as a decoy for innate immune sensors, contributing to pathogenesis. In West Nile virus, sfRNA antagonizes the type I interferon response and is required for full virulence in mouse models.

Herpesviruses, though DNA viruses, encode many viral microRNAs, but among bona fide RNA viruses, viral miRNAs are less common. Some cytoplasmic RNA viruses — including dengue virus, West Nile virus, and hepatitis C virus — have been reported to produce small RNAs with miRNA-like properties, but their biogenesis from a cytoplasmic RNA genome in the absence of nuclear Drosha processing has been a source of controversy. RNA viruses that replicate in the nucleus, such as influenza virus and bornaviruses, might have better access to the miRNA machinery. The related concept of virus-derived small interfering RNAs (vsiRNAs), which are produced by Dicer from viral double-stranded RNA in the context of antiviral RNAi, should not be confused with viral miRNAs; vsiRNAs are host defense molecules, not virus-encoded regulatory RNAs, although some viruses have evolved suppressors that interfere with the RNAi pathway.

Host shutoff is a virus-induced suppression of host gene expression that frees the translation apparatus, nucleotide pools, and energy resources for viral use. Viral proteases, RNases, and other accessory proteins execute host shutoff through several mechanisms. Picornaviruses encode the 2A protease (enteroviruses and rhinoviruses) or the L protease (aphthoviruses, including foot-and-mouth disease virus), which cleaves eIF4G, a subunit of the cap-binding complex eIF4F. Cleavage of eIF4G separates the cap-binding domain (which remains associated with eIF4E on capped mRNAs) from the ribosome-recruitment domain, abolishing cap-dependent translation while preserving IRES-dependent translation of the viral RNA, which does not require eIF4E. This mechanism is a clean switch: the virus inactivates host translation without touching its own translation.

Influenza virus uses a different shutoff mechanism. The viral polymerase binds the C-terminal domain of RNA polymerase II and cleaves nascent host transcripts through the PA endonuclease domain for cap-snatching, as described in Section 117.1. In addition, the viral NS1 protein inhibits host mRNA processing and export. Together, these activities suppress host gene expression at the transcriptional and post-transcriptional levels. Influenza virus also expresses the PA-X protein, which preferentially degrades host mRNA through its endonuclease domain, contributing to host shutoff.

Coronaviruses encode nsp1, a protein that binds the 40S ribosomal subunit and inactivates it by inserting its C-terminal domain into the mRNA entry channel, blocking host mRNA accommodation while sparing viral mRNAs that carry a specific 5′ leader sequence. SARS-CoV-2 nsp1 also induces endonucleolytic cleavage of host mRNAs near their 5′ ends, accelerating their degradation. This dual mechanism — ribosome inactivation and host mRNA cleavage — is a powerful host shutoff strategy.

In rhabdoviruses, the matrix protein inhibits host transcription by RNA polymerase II and interferes with mRNA export. In herpesviruses (not RNA viruses, but instructive for comparison), the virion host shutoff protein (vhs) is an RNase packaged in the virion that degrades host mRNA upon entry, before any viral gene expression. The diversity of host shutoff mechanisms across virus families demonstrates that suppression of host gene expression is a universal requirement, but the molecular strategies vary from protease cleavage of translation factors to RNase-mediated degradation to polymerase inhibition.

Common misconceptions about subgenomic RNAs include the belief that all subgenomic RNAs are noncoding or that they are always defective viral genomes. Subgenomic mRNAs are functional, coding mRNAs produced by the viral transcription machinery; they are not defective genomes. A second misconception is that host shutoff means global, undifferentiated suppression. Some viruses selectively shut off host mRNAs while preserving translation of specific host factors required for viral replication. The picornavirus 2A protease cleaves eIF4G but spares cellular mRNAs that contain IRES-like elements, and the virus can shut off interferon and cytokine expression while maintaining translation of certain host mRNAs.

The existing references include Lesnik 2002 (HIV RRE) and Pavlakis 1990 (HIV regulation), which are relevant backgrounds but not primary references for these mechanisms.

117.5. Structure probing, antivirals, and immune evasion

Viral RNA structures are validated as drug targets by structure probing, which experimentally maps base pairing and dynamics, and by functional assays that connect structural features to viral replication. Chemical probing methods — SHAPE (selective 2′-hydroxyl acylation analyzed by primer extension), DMS (dimethyl sulfate), CMCT, and in-line probing — react with RNA nucleotides in a way that depends on their structural context, typically with unpaired or flexible nucleotides being more reactive. These chemistries, read out by primer extension (for SHAPE, DMS) or by mutational profiling coupled to next-generation sequencing (for high-throughput approaches), produce reactivity profiles that can be used to constrain computational structure prediction, derive secondary structure models, and detect structural changes upon protein binding or environmental shifts.

SHAPE reagents, most notably 1M7 (1-methyl-7-nitroisatoic anhydride) and related compounds, acylate the 2′-hydroxyl of all four ribonucleotides with a reactivity that correlates strongly with local nucleotide flexibility. Unpaired and flexible nucleotides are more reactive; highly paired and constrained nucleotides are less reactive. SHAPE data, when used as pseudo-energy constraints in folding algorithms, enable secondary structure models for RNAs of several thousand nucleotides with accuracy comparable to phylogenetic analysis, and SHAPE has been applied to full-length viral genomes including HIV-1, HCV, dengue virus, and SARS-CoV-2.

DMS methylates the N1 of adenine and the N3 of cytosine at positions that participate in Watson-Crick base pairing when those positions are protected by pairing. DMS probing is performed on native RNA or on RNA after denaturation, and the difference in modification between the two conditions reveals which nucleotides are base-paired in the native structure. DMS has been used extensively to map viral RNA structures, including the HIV-1 5′ UTR, the HCV IRES, and the SARS-CoV-2 frameshift pseudoknot.

High-throughput approaches (SHAPE-MaP, DMS-MaPseq, and others) combine chemical probing with mutational profiling: the chemical adduct is read out as a mutation during reverse transcription, and deep sequencing of the cDNA yields a per-nucleotide modification rate that serves as a structural constraint. These methods have enabled genome-scale structure mapping of viral RNAs in intact virions, in infected cells (in-cell SHAPE), and under conditions that distinguish different functional states.

Figure 117.4. Chemical Probing of Viral RNA Structure

Figure 117.4. Chemical Probing of Viral RNA Structure. Chemical probing maps RNA secondary structure by measuring nucleotide flexibility. SHAPE reactivity reports local nucleotide dynamics; DMS reports Watson-Crick base-pairing status. Reactivity-constrained folding produces experimentally validated secondary structure models.

Small-molecule antivirals that target viral RNA structures have lagged behind protein-targeted antivirals because RNA has been considered a difficult drug target. RNA has fewer distinctive binding pockets than proteins, its surface is dominated by a negatively charged phosphate backbone that favors nonspecific electrostatic binding, and the conformational dynamics of RNA can obscure binding sites. Nevertheless, several strategies have produced validated RNA-targeting antiviral leads. The transactivation response element (TAR) of HIV-1, an RNA stem-loop at the 5′ end that binds the Tat protein and cyclin T1 to activate transcription, has been targeted by small molecules that compete with Tat binding and by arginine-rich peptides that bind the TAR bulge region. The HIV-1 RRE, a structured RNA in the env coding region that binds the Rev protein and mediates nuclear export of unspliced and singly spliced viral mRNAs, has been targeted by aminoglycoside-based and peptide-based compounds.

The HCV IRES has been among the most productive RNA-targeting antiviral targets. A benzimidazole-based compound was identified that binds domain IIa of the HCV IRES and stabilizes an alternative conformer that is incompatible with 40S ribosomal subunit binding, inhibiting translation in sub-micromolar concentrations in replicon assays. The SARS-CoV-2 frameshift pseudoknot has been targeted by a small molecule identified in a high-throughput screen that reduces frameshift efficiency, and by compounds that bind the pseudoknot and increase its mechanical stability, potentially trapping the ribosome at the frameshift site. The influenza virus promoter RNA, a partially double-stranded structure at the 5′ and 3′ termini of each segment, has been targeted by small molecules that bind the promoter and inhibit the viral polymerase.

Antisense oligonucleotides (ASOs) and small interfering RNAs (siRNAs) represent a different targeting strategy: instead of binding a structured RNA to block function, they hybridize to a target sequence by Watson-Crick base pairing. ASOs designed against conserved viral RNA structures — the IRES, the frameshift element, the packaging signal, the 5′ UTR — can block translation or replication by steric hindrance or by RNase H-mediated cleavage. The challenge for ASO-based antivirals is delivery: viral RNA is inside infected cells, and reaching those cells in vivo requires efficient formulation, tissue targeting, or local delivery. For respiratory viruses such as SARS-CoV-2 and respiratory syncytial virus, inhaled ASOs may overcome some delivery barriers, but systemic delivery for hepatitis, hemorrhagic fever, or neurotropic viruses remains a significant obstacle. The RNA-targeting approach has the advantage that conserved RNA structures are often less variable than protein epitopes, making resistance development slower.

Table 117.1. Immune Evasion Mechanisms Targeting RNA Sensing Pathways. Viruses evade innate immune sensing of RNA through structural sequestration, end modification, decoy RNAs, and protein antagonists. Each mechanism targets a specific sensor, and many viruses combine multiple evasion strategies.

Evasion mechanism RNA structural basis Virus examples Host sensor targeted Evidence standard
Double-stranded RNA sequestration in replication organelles Membrane-bound spherules, DMVs, or webs shield dsRNA from cytoplasm All positive-strand RNA viruses MDA5, PKR, OAS EM tomography, immuno-EM, RNase protection
5′ cap addition (cap-0, cap-1, 2′-O-methylation) Methylated cap structure mimics host mRNA Flaviviruses, coronaviruses, rhabdoviruses RIG-I (cap-0 vs cap-1 discrimination), IFIT1 Structural, biochemical, IFIT1 binding assays
Cap-snatching from host mRNAs Cleaved 5′ capped fragments from host pre-mRNA prime viral transcription Orthomyxoviridae, Bunyavirales RIG-I Biochemical, crystal structures of cap-binding and endonuclease domains
RNA decoy (VA RNA I for PKR) Highly structured ~160 nt Pol III transcript binds PKR but fails to activate it Adenovirus PKR Genetic (VA RNA deletion), biochemical, structural
RNA decoy (flavivirus sfRNA for RIG-I and other sensors) Highly structured subgenomic flavivirus RNA accumulates and binds RIG-I and other sensors Flaviviruses (dengue, Zika, WNV) RIG-I, MDA5, PKR sfRNA-knockout virus, biochemical binding, IFN reporter assays
Structured RNA protein antagonism (HIV-1 TAR) HIV-1 TAR RNA binds Tat for transcriptional activation; influences innate sensing balance HIV-1 MDA5 (indirect, via transcription start site heterogeneity) Primary: Hughes et al. 2026
Protein-based dsRNA sequestration (influenza NS1) Influenza NS1 dimer binds dsRNA and blocks RIG-I activation through direct competition Influenza A virus RIG-I Genetic (NS1 mutant viruses), structural, biochemical
Protein pseudosubstrate (vaccinia K3L) Vaccinia K3L mimics eIF2α and binds PKR without being phosphorylated Vaccinia virus PKR Structural (K3L–PKR co-crystal), genetic, biochemical
Viral RNase-mediated host mRNA degradation Viral endoribonuclease degrades host mRNAs, reducing sensor and IFN production Influenza (PA-X), coronaviruses (nsp15), herpesviruses (SOX, vhs) Multiple (indirect) Genetic (RNase mutant viruses), RNA-seq of host transcriptome

Table 117.1. Immune Evasion Mechanisms Targeting RNA Sensing Pathways

Evasion mechanism RNA structural basis Virus examples Host sensor targeted Evidence standard Key references
dsRNA sequestration in replication organelles Membrane-bound spherules, DMVs, or webs shield dsRNA from cytoplasm All positive-strand RNA viruses MDA5, PKR, OAS EM tomography, immuno-EM, RNase protection reference-target:
5′ cap addition (cap-0, cap-1, 2′-O-methylation) Methylated cap structure mimics host mRNA Flaviviruses, coronaviruses, rhabdoviruses RIG-I (cap-0 vs cap-1 discrimination), IFIT1 Structural, biochemical, IFIT1 binding assays reference-target:
Cap-snatching from host mRNAs Cleaved 5′ capped fragments from host pre-mRNA prime viral transcription Orthomyxoviridae, Bunyavirales RIG-I Biochemical, crystal structures of cap-binding and endonuclease domains reference-target:
RNA decoy: VA RNA I Highly structured ~160 nt Pol III transcript binds PKR but fails to activate it Adenovirus PKR Genetic (VA RNA deletion), biochemical, structural reference-target:
RNA decoy: sfRNA Highly structured subgenomic flavivirus RNA accumulates and binds RIG-I and other sensors Flaviviruses (dengue, Zika, WNV) RIG-I, MDA5, PKR sfRNA-knockout virus, biochemical binding, IFN reporter assays reference-target:
Structured RNA protein antagonism: TAR HIV-1 TAR RNA binds Tat for transcriptional activation; influences innate sensing balance HIV-1 MDA5 (indirect, via transcription start site heterogeneity) Primary: Hughes et al. 2026 doi__10.1073_pnas.2522948123
Protein-based dsRNA sequestration: NS1 Influenza NS1 dimer binds dsRNA and blocks RIG-I activation through direct competition Influenza A virus RIG-I Genetic (NS1 mutant viruses), structural, biochemical reference-target:
Protein pseudosubstrate: K3L Vaccinia K3L mimics eIF2α and binds PKR without being phosphorylated Vaccinia virus PKR Structural (K3L–PKR co-crystal), genetic, biochemical reference-target:
Viral RNase-mediated host mRNA degradation Viral endoribonuclease degrades host mRNAs, reducing sensor and IFN production Influenza (PA-X), coronaviruses (nsp15), herpesviruses (SOX, vhs) Multiple (indirect) Genetic (RNase mutant viruses), RNA-seq of host transcriptome reference-target:

Figure 117.5. Viral RNA Immune-Evasion Mechanisms Mapped to the Sensing Step They Defeat

Figure 117.5. Viral RNA Immune-Evasion Mechanisms Mapped to the Sensing Step They Defeat. Viral RNA evasion does not erase RNA ancestry; it changes which molecular feature is exposed, which sensor can engage, or whether binding matures into productive signaling.

Immune evasion by viral RNA is achieved through at least four mechanisms: structural sequestration, end modification, RNA decoys, and active suppression. Structural sequestration is the oldest mechanism: viral double-stranded RNA replication intermediates are generated inside membrane-bound replication organelles (spherules, double-membrane vesicles, membranous webs; see Section 116.2) that physically separate the dsRNA from cytoplasmic sensors such as MDA5, PKR, and OAS. The membrane barrier is imperfect, and some dsRNA may leak into the cytoplasm or be released during viral egress, but the compartmentalization substantially reduces sensor activation.

End modification is the second mechanism. The 5′ ends of many viral mRNAs are capped by viral capping enzymes (in flaviviruses, coronaviruses, rhabdoviruses, and paramyxoviruses) or by cap-snatching from host mRNAs (in influenza virus and bunyaviruses). The cap structure — 7-methylguanosine linked by a 5′-to-5′ triphosphate bridge — is recognized by the host cap-binding complex and is a signal of “self” RNA. RIG-I recognizes uncapped, 5′-triphosphate-bearing RNA as “nonself” and activates the antiviral response. By capping their RNAs, viruses evade RIG-I. However, many viruses take this evasion one step further by adding a 2′-O-methyl group to the first nucleotide (cap-1 structure), which prevents recognition by IFIT1, a host protein that sequesters RNAs lacking 2′-O-methylation and blocks their translation. Coronaviruses encode two methyltransferases (nsp14, nsp16) to achieve the full cap-1 structure, and mutant viruses lacking 2′-O-methylation are attenuated and more sensitive to interferon.

RNA decoys are a third mechanism. Certain viral noncoding RNAs can bind and inhibit innate immune sensor proteins, acting as competitive antagonists. The adenovirus VA RNA I (a ~160-nucleotide structured RNA transcribed by RNA polymerase III) is a well-established PKR inhibitor: it forms a structured domain that binds PKR with high affinity but does not support PKR activation, effectively sequestering PKR and preventing it from phosphorylating eIF2-alpha. VA RNA I was one of the first examples of an RNA-mediated immune evasion strategy. In HIV-1, the TAR RNA element at the 5′ end of all transcripts has been reported to bind and inhibit PKR, although this function is less well established than VA RNA I PKR inhibition. Structured viral RNAs that contain imperfect double-stranded regions may bind RIG-I or MDA5 in conformations that do not trigger oligomerization and signaling, and the sfRNA of flaviviruses is proposed to function as a RIG-I decoy in addition to its other roles.

Active suppression of innate immune signaling constitutes the fourth mechanism. The influenza NS1 protein binds double-stranded RNA and sequesters it from RIG-I, and also directly interacts with RIG-I and the downstream adaptor MAVS to inhibit signaling. The SARS-CoV-2 macrodomain (part of nsp3) removes poly(ADP-ribose) from proteins, interfering with PARP-mediated antiviral signaling. These protein-based mechanisms are not RNA-centered, but they operate on the same dsRNA detection pathways and are thus integrated with RNA structure biology.

Several viral proteins specifically antagonize PKR. The influenza NS1 protein binds dsRNA and prevents PKR activation. The hepatitis C virus NS5A and the Ebola virus VP35 proteins bind dsRNA and inhibit PKR. The vaccinia virus E3L protein (from a DNA virus but mechanistically informative) contains a dsRNA-binding domain that sequesters dsRNA from PKR. Some viruses encode proteins that act as pseudosubstrates, binding PKR and blocking phosphorylation of eIF2-alpha without involving RNA. The vaccinia virus K3L protein mimics eIF2-alpha and acts as a competitive inhibitor of PKR.

The involvement of Nod-like receptors in viral RNA sensing has been less studied than RIG-I/MDA5/PKR, but evidence for NLRX1, a mitochondrial Nod-like receptor, suggests it can bind RNA and modulate antiviral signaling.

Box 117.1. Common Misconceptions About Viral RNA Structures

  • “All computationally predicted RNA structures are functional” — Structure prediction alone is not proof. Functional validation requires mutagenesis, compensatory mutagenesis, or chemical probing with functional correlation.
  • “Phylogenetic conservation equals functional necessity” — The s2m element of SARS-CoV-2 is highly conserved across coronaviruses yet dispensable for replication in cell culture and mice.
  • “IRES elements are found only in viruses” — Cellular IRESs exist in many eukaryotic mRNAs, particularly those encoding stress-response and cell-cycle proteins.
  • “The HIV-1 psi sequence is a simple linear signal” — psi is a structurally complex ensemble of stem-loops; the sequence outside the structural context is not a functional packaging signal.
  • “Host shutoff is always complete and indiscriminate” — Many viruses preserve translation of host factors they require; shutoff efficiency and selectivity vary.
  • “Viral RNA that activates RIG-I is always detrimental” — The balance between immune activation and evasion is context-dependent; some immune activation may slow clearance or shape the cytokine environment in ways that benefit viral transmission.
  • “SHAPE reactivity directly equals unpaired status” — SHAPE reactivity reports local nucleotide flexibility, which correlates with but is not identical to unpaired status; highly flexible paired nucleotides can show intermediate reactivity.
  • “A consensus secondary structure is the only structure” — RNA is dynamic; ensemble probing captures the most probable structure, not all functional states.

A recurring misconception is that structured viral RNA is necessarily immunostimulatory and that RNA structure always promotes innate immune activation. The truth is bidirectional: some structured viral RNAs (long dsRNA, 5′-triphosphate-bearing panhandle structures) are potent RIG-I and PKR activators, while other structures (capped, 2′-O-methylated, sequestered inside replication organelles, formed as highly compact pseudoknots that evade sensor binding) suppress activation. The net immune outcome depends on the balance between activating and evasive structures, the timing and localization of their exposure, and the host cell’s sensor expression and signaling state.

Another misconception concerns structure probing: the resulting structure model represents an average over all conformations in the population and over the probing time. RNA structures are dynamic, and probing data represent a Boltzmann-weighted average of the structural ensemble; the model derived from the probing data is the most probable structure under the experimental conditions, not necessarily the only structure. Functional states that represent a minority of the population may be invisible to ensemble probing but detectable by single-molecule methods.

The existing references cover HIV RRE (Lesnik 2002), HIV regulation (Pavlakis 1990), HIV TSS-MDA5 sensing (Hughes 2026), HIV TAR immune variation (Stunnenberg 2021), NLRX1 RNA binding (Hong 2012), and SARS-CoV-2 stem-loop dispensability (Jiang 2023). Additional references are needed across all topics.

Experimental Foundations and Evidence Standards

The evidence that a viral RNA structure is functional comes from a combination of experimental approaches. Chemical probing defines the secondary structure in solution. Phylogenetic covariation analysis identifies base-paired positions that co-vary across related isolates, providing independent evidence for structure. Mutagenesis of individual nucleotides or base pairs tests whether the structure is required for function. Compensatory mutagenesis — the most rigorous test — shows that a mutation that disrupts a base pair reduces function, while a second mutation that restores the base pair (with different nucleotides) restores function. Structural biology (cryo-EM, X-ray crystallography, NMR) resolves the three-dimensional fold and often confirms the secondary structure model and adds tertiary contacts.

For IRES elements, the standard functional assay is the bicistronic reporter: the first cistron (e.g., Renilla luciferase) is translated by cap-dependent scanning, and the second cistron (e.g., firefly luciferase) is translated only if the intervening sequence functions as an IRES. Controls include eliminating the first cistron translation to rule out readthrough, showing that the IRES activity is independent of the first cistron, and demonstrating that the construct does not produce a cryptic monocistronic mRNA by splicing or cleavage. For frameshifting, dual-reporter systems in which the two reading frames encode different reporters provide quantitative frameshift efficiency measurements, and in vitro translation in rabbit reticulocyte lysate or wheat germ extract provides a simpler system for mechanistic dissection.

For packaging signals, the standard approach combines mutagenesis of the putative packaging signal with measurement of genomic RNA incorporation into virions (by RT-qPCR or Northern blot) and particle infectivity. The distinction between reduced packaging and reduced genome replication must be made: if the mutation reduces genome replication, less genomic RNA is available for packaging, which can mimic a packaging defect. Complementation experiments that provide the replication function in trans while the packaging signal remains in cis can separate these effects.

For host shutoff, the standard measurements are metabolic labeling of nascent proteins with radioactive amino acids (classical) or puromycin incorporation (modern), combined with mRNA quantification by RNA-seq or RT-qPCR to determine whether reduced translation is accompanied by reduced mRNA abundance (indicating mRNA degradation) or not (indicating translational inhibition). Ribosome profiling, which maps ribosome-protected mRNA fragments, provides transcriptome-wide resolution of translation changes during shutoff.

For immune evasion, the gold standard is the comparison of a wild-type virus with a mutant virus lacking the candidate evasion mechanism in cells or animals that are competent or deficient in the relevant sensor pathway. For example, a cap-methylation-deficient coronavirus is compared with wild-type virus in wild-type versus IFIT1-knockout cells: if IFIT1 restricts the mutant but not the wild-type, IFIT1 evasion is established.

The main interpretive hazards include: confounding RNA structure with RNA sequence when compensatory mutagenesis is not performed, inferring function from computational structure predictions alone, mistaking a consensus secondary structure as the only structure when RNA can sample multiple conformations, and neglecting the possibility that a structure-probing or mutagenesis result reflects a change in RNA stability or protein binding rather than the intended parameter. Vigorous structure-function claims require multiple lines of evidence.

Biological Contexts and Cross-Chapter Boundaries

This chapter explains how viral RNA genomes function beyond coding sequence. It bridges the genome strategies chapter (Chapter 115), which describes what viral genomes look like, with the polymerase chapter (Chapter 116), which explains how they are copied, the packaging chapter (Chapter 118), which explains how they are condensed and assembled, and the retrovirus chapter (Chapter 120), which treats reverse transcription and integration. The immune sensing chapters (Chapter 108) cover the host side of the RIG-I/MDA5/PKR story, which this chapter approaches from the viral evasion perspective. The structural probing methods connect to the RNA structure prediction and probing methods chapters (Chapter 60-Chapter 63). The IRES and frameshifting mechanisms connect to the translation chapters (Chapter 67-Chapter 70).

The most important cross-chapter boundaries are as follows. Chapter 116 covers the polymerase that recognizes the replication promoters and elem described in this chapter; the promoter structures (flavivirus cyclization, picornavirus cloverleaf, influenza promoter) should be read together with the polymerase enzymology. Chapter 118 covers the biophysical principles of viral RNA condensation and packaging; this chapter provides the RNA structural side of packaging signals and dimerization, which is the ligand for the condensation machinery described in Chapter 118. Chapter 108 covers the host sensors (RIG-I, MDA5, PKR) that detect the RNA structures described here; this chapter describes how viruses modify, sequester, or decoy their RNA to evade those sensors. Chapter 119 covers viroids, which are infectious RNA circles without protein-coding capacity, and their replication depends entirely on RNA structures that mimic and exploit host machinery; the boundary is that viral RNA regulatory elements and viroid structural elements share principles of RNA-based regulation.

Recent Consensus

The following positions represent recent consensus. RNA virus UTRs are densely populated with functional structures that control replication, translation, and packaging, and these structures are often the most conserved genome regions. Internal ribosome entry sites are established as the mechanism by which many positive-strand RNA viruses and some cellular mRNAs initiate cap-independent translation. Programmed -1 ribosomal frameshifting is the primary mechanism by which retroviruses and many other RNA viruses control the ratio of structural to enzymatic proteins, and the stimulatory pseudoknot functions by resisting unwinding during ribosome translocation. The HIV-1 psi and DIS elements are the best-understood packaging signal system, and the kissing-loop dimerization mechanism is a shared feature of retroviruses. Subgenomic RNA synthesis by discontinuous transcription in coronaviruses, by internal initiation in alphaviruses, and by sequential start-stop in rhabdoviruses are well-established mechanisms. Host shutoff is a conserved strategy across virus families, executed by diverse molecular mechanisms. Viral double-stranded RNA sequestration in membrane-bound replication organelles and 5′ end capping and methylation are the two most broadly used innate immune evasion mechanisms. RNA structures are viable antiviral targets, as demonstrated by compounds that bind the HCV IRES, the HIV-1 TAR and RRE, and the SARS-CoV-2 frameshift pseudoknot.

The SARS-CoV-2 frameshift pseudoknot structure has been established at high resolution by cryo-EM, and its role in drug targeting is an active area. The dispensability of the highly conserved stem-loop II motif (s2m) in the SARS-CoV-2 3′ UTR for replication in cell culture and in the mouse model has been demonstrated, challenging the assumption that all deeply conserved viral RNA structures are essential.

Open Questions, Controversies, Deprecated Models, and Common Misconceptions

Open questions:

  • To what extent do RNA virus genomes encode functional RNA structures beyond the UTRs and known regulatory elements? Genome-scale structure probing (SHAPE-MaP, DMS-MaPseq) has revealed many structured regions in viral coding sequences that are evolutionarily conserved, but functional validation through mutagenesis and compensatory mutagenesis lags far behind structure mapping. Whether these coding-region structures are functional regulatory elements, neutral folds, or structural noise is unresolved for most of them.
  • How do viral RNA conformational dynamics control transitions among translation-, dimerization-, and packaging-competent states? Ensemble structure probing captures the average structure, but many viral RNAs, particularly the HIV-1 5′ UTR, exist in multiple conformations that correspond to different functional states. The mechanisms by which the virus and host factors shift the conformational equilibrium, and the kinetics of these shifts, are incompletely understood. Single-molecule methods such as FRET, optical tweezers, and nanopore sequencing are beginning to address these questions but have so far been applied to only a few viral RNAs.
  • Can antiviral strategies targeting viral RNA structures achieve clinical efficacy without unacceptable toxicity? Small molecules that bind structured viral RNA must compete with abundant cellular RNAs for binding, and off-target RNA binding can lead to toxicity. The clinical success of ribosome-targeting antibiotics (which bind rRNA) demonstrates that RNA can be a selective drug target, but the therapeutic index for RNA-targeting antivirals remains to be established in clinical trials.
  • What is the role of RNA modifications in viral immune evasion beyond 5′ cap methylation? N6-methyladenosine (m6A) modifications have been detected on several viral RNAs, including HCV, HIV-1, influenza, and SARS-CoV-2, and are proposed to modulate viral RNA stability, translation, and innate immune recognition. Whether m6A acts primarily as a proviral modification, an antiviral host mark, or both depending on context is unresolved.

Common misconceptions:

  • “All predicted RNA secondary structures in viral genomes are functional.” Structure prediction alone is not sufficient proof; many predicted structures may be artifacts of the prediction algorithm or biologically irrelevant folds.
  • “IRES elements are found only in viruses.” Cellular IRESs exist in many eukaryotic mRNAs, particularly those encoding stress-response proteins, and the discovery of cellular IRESs followed the discovery of viral IRESs.
  • “The HIV-1 psi sequence is a simple linear signal.” Psi is a structurally complex ensemble of stem-loops whose function depends on correct folding; the sequence alone, in an unstructured context, is not a packaging signal.
  • “Host shutoff is always complete and indiscriminate.” Many viruses selectively preserve translation of specific host factors they require, and shutoff efficiency varies among virus families, cell types, and infection stages.
  • “Viral RNA that activates RIG-I is always harmful to the virus.” Some activation of innate immunity may be tolerated or even beneficial at certain infection stages, and the balance between activation and evasion is context-dependent.