Chapter 142. Small RNA Target Prediction, Regulatory Networks, and Quantitative Post-Transcriptional Models

Scope Note

This chapter explains how computational models predict targets of regulatory small RNAs, how those predictions are converted into post-transcriptional regulatory networks, and why quantitative evidence is essential before treating a predicted interaction as a biological mechanism. The emphasis is on microRNAs (miRNAs), small interfering RNAs (siRNAs), PIWI-interacting RNAs (piRNAs), and bacterial small RNAs (sRNAs), with attention to the different pairing rules, protein cofactors, cellular compartments, and validation standards that separate these systems. The chapter treats target prediction as a probabilistic inference problem rather than as a lookup table: sequence complementarity, accessibility, conservation, expression, Argonaute occupancy, perturbation response, stoichiometry, and benchmark design all affect whether a predicted site is likely to mediate regulation in a real cell.

Executive Summary

Small-RNA target prediction asks whether a short guide RNA can recruit a regulatory effector complex to a target RNA under cellular conditions. For animal miRNAs, the dominant features are seed pairing, site context, target accessibility, 3′ pairing, site conservation, target abundance, and the presence of Argonaute-bound sites in relevant cells. For plant miRNAs and many siRNAs, near-perfect complementarity and cleavage potential are more prominent, although translational repression and noncanonical interactions complicate any simple rule. For piRNAs, prediction depends on pathway context: many piRNAs direct PIWI-associated recognition of transposon transcripts through extensive complementarity, but piRNA populations are numerous, repetitive, strand-biased, and tissue-specific, so computational target assignment can be confounded by genome repetitiveness and incomplete annotation. Bacterial sRNA prediction is different again because many trans-encoded bacterial sRNAs use partial base pairing, chaperones such as Hfq or ProQ, and target-site occlusion near ribosome-binding sites or coding-sequence regions; bacterial predictions must account for RNA structure, chaperone binding, transcript boundaries, and mRNA decay coupling.

Predicted target sites do not by themselves define regulatory networks. A regulatory edge becomes stronger when the small RNA and target are co-expressed, the guide-loaded effector binds near the predicted site, target RNA or protein changes after perturbing the small RNA, the site is sufficient in a reporter and necessary in the endogenous transcript, and mutational rescue restores regulation. Network models use these data to infer regulatory modules, but they must separate direct effects from indirect downstream expression changes. Quantitative post-transcriptional models add parameters for guide abundance, target abundance, site affinity, repression efficiency, decay rates, translation rates, and competition among targets. These parameters matter because a rare target cannot usually sequester a highly abundant miRNA enough to change the repression of other targets, whereas a highly expressed transcript with many effective sites, or an engineered sponge, can sometimes act as a meaningful competitor.

The competing endogenous RNA (ceRNA) hypothesis is best treated as a quantitative, context-dependent model rather than a universal explanation for correlated transcript expression. A ceRNA claim requires more than shared predicted miRNA sites and positive correlation between transcripts. Strong evidence should show stoichiometric plausibility, expression in the same compartment and cell type, dependence on the relevant miRNA and binding sites, dose-response behavior in the endogenous expression range, and exclusion of shared transcriptional regulation or cell-state confounding. Many ceRNA reports remain plausible hypotheses rather than established mechanisms because target-site occupancy, absolute molecule counts, and site-specific rescue are missing.

Benchmarking is the discipline that prevents prediction models from becoming catalogs of attractive false positives. A useful benchmark separates training and testing by genes, miRNA families, species, experiments, or time; uses direct interaction data and perturbation response data for different questions; controls for 3′ UTR length, expression, conservation, and GC content; reports precision-recall behavior rather than only accuracy; and compares predictions against matched negative sets. Because small-RNA regulation is often modest, redundant, and context-dependent, target prediction should usually be interpreted as prioritization for validation, not proof of regulation.

Concept Inventory

  • miRNA target site: a region of a target RNA that can be recognized by a miRNA-loaded Argonaute complex. In animals, many canonical sites pair to miRNA positions 2-7 or 2-8, called the seed region; additional pairing, local sequence context, and RNA accessibility affect efficacy.
  • Seed match: Watson-Crick complementarity between the target and the miRNA seed. Seed matches are useful predictive features but are not sufficient evidence for regulation because short motifs occur frequently in transcriptomes.
  • Site context: features around a predicted small-RNA binding site, including location in a 3′ untranslated region (3′ UTR), local AU-richness, distance from stop codon or poly(A) tail, site clustering, structural accessibility, isoform usage, and nearby RNA-binding proteins.
  • Guide-loaded effector complex: a protein-small RNA complex, such as Argonaute-miRNA, Argonaute-siRNA, PIWI-piRNA, or an Hfq-bound bacterial sRNA complex, that mediates target recognition and repression.
  • Target-site accessibility: the probability that the target nucleotides are not sequestered in intramolecular RNA structure or protein-bound states that prevent pairing. Accessibility can be estimated computationally, but in-cell structure and RNP occupancy may diverge from in vitro folding predictions.
  • Direct target: a target RNA whose regulation depends on a specific small-RNA interaction site. Directness requires site-level evidence, not only expression change after perturbing the small RNA.
  • Indirect target: an RNA or protein that changes after small-RNA perturbation because a direct target changed upstream, because cell state changed, or because the perturbation produced off-target or immune effects.
  • ceRNA: a transcript proposed to regulate another transcript by competing for shared miRNAs. This chapter treats ceRNA as a quantitative hypothesis requiring stoichiometric and site-specific evidence.
  • Small-RNA off-target effect: regulation of unintended RNAs by a designed siRNA, miRNA mimic, antisense-like small RNA, or guide RNA. Seed-mediated siRNA off-targeting in animals is a major concern for experiments and therapeutics (Bereczki et al. 2025).
  • Benchmark leakage: a modeling artifact in which training and test data are not independent, often because homologous transcripts, related miRNA family members, replicated assay records, or database-derived labels appear in both sets.

What to Know Before Reading This Chapter

A small RNA is a short regulatory RNA, usually about 20-30 nucleotides for miRNAs, siRNAs, and piRNAs and often 50-300 nucleotides for bacterial sRNAs, that guides a protein or RNP complex to another nucleic acid. The physical object being predicted is not only a base-paired duplex. It is a cellular encounter among a guide RNA, an effector protein, a target RNA isoform, competing RNAs, RNA-binding proteins, ribosomes, decay factors, and a specific biological state. The same 7-nucleotide seed match can be irrelevant in one cell because the target isoform lacks the 3′ UTR, inaccessible in another because a protein masks the site, and functional in a third because the miRNA, target, and repression machinery are all abundant in the same compartment.

Readers should distinguish three related questions. First, does the guide RNA physically bind or pair with the candidate target? Second, does binding cause a measurable molecular change, such as accelerated mRNA decay, translational repression, cleavage, or altered ribosome occupancy? Third, does that molecular change matter for a phenotype, pathway, disease state, or therapeutic outcome? Prediction tools often address only the first question and approximate the second. Network studies and disease papers often leap to the third. This chapter emphasizes the evidentiary steps between those questions.

This chapter also assumes familiarity with transcript isoforms. A predicted site is meaningful only if the relevant RNA isoform contains the site in the cell type being studied. Alternative polyadenylation can shorten 3′ UTRs and remove miRNA sites; alternative splicing can create or delete coding-sequence or untranslated-region sites; bacterial transcript boundaries can change under stress; and repetitive transposon fragments can obscure piRNA target assignment. Target prediction therefore depends on annotation as much as on pairing rules.

142.1. miRNA target prediction features and algorithm classes

MicroRNA target prediction began with a simple observation that remains central: animal miRNAs often repress mRNAs through short sites complementary to the miRNA seed region, especially nucleotides 2-7 or 2-8 from the miRNA 5′ end. A 6- to 8-nucleotide sequence is short enough to occur many times by chance, so seed matching is a sensitive but nonspecific feature. Predictive models therefore add biological filters that ask whether a seed match lies in a plausible context, is conserved in related species, is accessible in the target RNA, is supported by Argonaute binding, and is associated with mRNA or protein repression after miRNA perturbation. The prediction problem is thus not “find all reverse complements of a short sequence”; it is “rank possible Argonaute-mediated regulatory sites for a particular cell state and transcript annotation.”

The canonical animal miRNA site classes are useful vocabulary but should not be treated as absolute categories. A 6mer site pairs to positions 2-7; a 7mer-m8 site pairs to positions 2-8; a 7mer-A1 site pairs to positions 2-7 and contains an adenine opposite miRNA position 1; and an 8mer site combines positions 2-8 pairing with an A1 feature. These classes usually rank from weaker to stronger in average repression, but individual site performance varies. Supplementary pairing to the miRNA 3′ region can strengthen some sites, offset imperfect seed pairing in some noncanonical sites, or contribute to target-directed miRNA degradation when pairing is extensive enough in the right architecture. A scientifically careful prediction therefore states which site class is being called, whether noncanonical sites are allowed, and how the model treats 3′ pairing and site multiplicity.

Figure 142.1. Feature grammar of animal miRNA target prediction

Figure 142.1. Feature grammar of animal miRNA target prediction. Prevents readers from equating seed matching with target validation.

Site context describes why two identical seed matches can behave differently. Sites in long 3′ UTRs are common, but many are weak because they are embedded in structured regions, densely occupied by RNA-binding proteins, or far from regions where Argonaute repression is most efficient. AU-rich sequence around the site can increase accessibility. Multiple sites for the same miRNA or for related miRNA family members can produce stronger repression than a single isolated site, although spacing and cooperative effects are not universal. Sites very near the stop codon or poly(A) tail can have different activity from internal sites. Coding-sequence and 5′ UTR sites can be functional, but they are often harder to interpret because ribosomes, initiation factors, RNA structure, and coding constraints alter accessibility and evolutionary conservation. The best models treat site location as a quantitative feature rather than as a binary permission rule.

Conservation is powerful because functional miRNA sites can be retained across related species despite the high background rate at which short motifs appear and disappear. A conserved seed match in orthologous 3′ UTRs, especially for a conserved miRNA family, deserves a higher prior probability than an isolated nonconserved match. Conservation is not a proof of targeting, however. Some true targets are lineage-specific, condition-specific, viral, or recently evolved. Some conserved motifs are maintained for reasons unrelated to the miRNA under study, such as binding by another RBP, codon constraints, overlapping regulatory elements, or regional nucleotide composition. Conservation filters can also underperform in rapidly evolving 3′ UTRs and in organisms with poor comparative genomic sampling. Thus conservation is best used as one feature in a ranking model, not as a universal evidence threshold.

Accessibility modeling tries to estimate whether the target nucleotides are available for pairing. Thermodynamic tools can compute the energy cost of opening local structure and the energy gained by guide-target hybridization. These methods are useful because small-RNA binding must compete with intramolecular base pairing in the target RNA. But accessibility estimates are approximations. In a cell, the target RNA may be bound by RBPs, scanned by ribosomes, localized to a granule, modified chemically, or folded co-transcriptionally into structures not captured by equilibrium predictions. Probing data, such as DMS- or SHAPE-derived accessibility, can improve context when available, but even probing measures an ensemble that may not match the exact RNP state recognized by Argonaute. Accessibility should therefore adjust prediction confidence rather than replace experimental validation.

Algorithm classes differ mainly in how they combine these features. Rule-based tools apply explicit criteria such as seed match class, conservation, free energy, and site location. Scoring models assign weights to features, often learned from reporter assays, mRNA fold-change data, or proteomic response after miRNA perturbation. Machine-learning models use labeled examples to learn nonlinear relationships among sequence, structural, conservation, expression, and biochemical features. Deep-learning models can learn motifs and local context directly from sequence or multi-modal features, but they can also learn dataset biases if benchmark splits are weak. Review-level discussions of computational miRNA identification and noncoding RNA target prediction provide useful background, but several landmark miRNA target-prediction tool papers are missing from the local bibliography and need expert curation (Rajendiran et al. 2018; Jiang et al.

Table 142.1. Features used in animal miRNA target-prediction algorithms. Helps readers separate sequence, context, evidence, and response features.

Feature Biological rationale Typical data source Common failure mode Whether feature is context-dependent
Seed class 6mer, 7mer-A1, 7mer-m8, and 8mer sites capture the dominant animal miRNA recognition grammar. Mature miRNA sequence and annotated target UTR sequence. Short seed motifs occur frequently and create many sequence-only false positives. Yes; site strength varies with transcript isoform, cell state, and surrounding features.
3′ supplementary pairing Pairing outside the seed can strengthen some sites or mark unusual noncanonical interactions. Guide-target alignment, thermodynamic hybridization, and site-level validation. Overweighting 3′ pairing can rescue weak seed calls without functional evidence. Yes; architecture and extent of pairing can change repression or trigger target-directed miRNA decay.
AU-rich context AU-rich flanks often make a site more accessible to Argonaute-loaded guide pairing. Local sequence composition around the candidate site. AU-richness can track UTR composition rather than causal accessibility. Yes; RBP occupancy, structure, and isoform choice can override local sequence composition.
Site location 3′ UTR sites are common, but coding-region and 5′ UTR sites can also function in restricted contexts. Transcript annotation, UTR boundaries, coding sequence, and alternative polyadenylation maps. Using a reference isoform can call sites absent from the expressed transcript. Yes; translation, stop-codon distance, poly(A)-proximal context, and cell-type-specific UTRs matter.
Site number and spacing Multiple effective sites can increase repression or competition capacity. Site scans across expressed isoforms and miRNA-family seed groups. Counting many weak or inaccessible sites can inflate gene-level scores. Yes; spacing, saturation, cooperative effects, and site masking are not universal.
Conservation Conserved sites have a higher prior probability when guide and UTR orthology are reliable. Comparative genomics across orthologous UTRs and conserved miRNA families. Conserved motifs may reflect RBP binding, coding constraints, or nucleotide composition. Yes; lineage-specific, viral, condition-specific, and poorly sampled systems can be missed.
Accessibility Sites must compete with target RNA structure and RNP occupancy before pairing can occur. Folding energy, opening-energy estimates, SHAPE or DMS probing, and RBP maps. Equilibrium predictions may not match in-cell folding, translation, localization, or modifications. Yes; accessibility is state-specific and should adjust confidence rather than validate a site alone.
Argonaute CLIP AGO occupancy near a candidate site raises confidence that the target region is physically contacted. AGO HITS-CLIP, PAR-CLIP, eCLIP, CLEAR-CLIP, CLASH, or related chimera data. CLIP peaks can reflect crosslinking bias, abundance, indirect association, or incomplete condition coverage. Yes; occupancy is cell type, condition, antibody, nuclease, and library dependent.
miRNA and target expression Direct repression requires guide-loaded effector and target isoform to be present together. Small-RNA-seq, RNA-seq, isoform quantification, proteomics, and localization data. Relative expression can hide molecule counts, ligation bias, isoform absence, or compartment mismatch. Yes; absolute abundance and compartment are central for repression and competition.
Perturbation response Target RNA, translation, or protein changes after guide perturbation support regulatory relevance. Mimic, inhibitor, genetic deletion, time-course RNA-seq, ribosome profiling, or proteomics. Overexpression, incomplete inhibition, family compensation, and indirect cascades can mimic direct effects. Yes; dose, timing, guide family redundancy, and assay layer strongly affect interpretation.

A practical miRNA target-prediction workflow begins by specifying the guide, transcript annotation, and biological context. If the guide is a known miRNA, mature miRNA arm choice and family membership should be checked because 5p and 3p arms can have different seeds and targets. If the target transcript is a human gene, the analyst should choose the relevant transcript isoforms, including cell-type-specific 3′ UTRs. The workflow then scans for canonical and optionally noncanonical sites, computes feature scores, overlays expression and Argonaute occupancy when available, ranks sites and genes, and separates “site-level prediction” from “gene-level prediction.” A gene with one strong site and a gene with ten weak sites are not equivalent, and a gene-level score can hide which isoform or site is responsible.

Expression data constrain predictions in two ways. A miRNA cannot directly repress an mRNA in a cell where the miRNA-loaded Argonaute complex is absent, and an mRNA cannot be regulated through an isoform or site that is not expressed. Absolute abundance is more informative than relative expression because repression and competition are stoichiometric processes. Small-RNA-seq read counts can be distorted by ligation bias and RNA modifications, while mRNA-seq can obscure isoform usage and subcellular localization. Even with these limitations, context-specific expression filters usually remove many biologically impossible predictions and should be reported explicitly.

Argonaute CLIP-family data provide a bridge between sequence prediction and physical occupancy. Crosslinking and immunoprecipitation methods can identify regions of RNAs bound by Argonaute in a specific cell type. Ligation-based variants can recover chimeric reads that join the miRNA and target fragment, giving stronger site-level evidence. These data are not complete maps of all functional interactions. Crosslinking efficiency varies, highly expressed transcripts can dominate, ligation creates biases, and some functional interactions may be missed. A CLIP peak near a seed match raises confidence; a seed match without CLIP support may still be functional if the experiment lacked sensitivity or the relevant condition was not assayed. Prediction models should treat CLIP as context-specific evidence, not as a universal gold standard.

The most important caution for miRNA target prediction is that gene lists are easy to overinterpret. A single miRNA family can have thousands of seed matches across the transcriptome, but only a subset are occupied and an even smaller subset produce measurable regulation in a particular state. Enrichment of seed matches among downregulated genes after miRNA transfection is strong evidence that the miRNA is active, but it does not prove that every downregulated gene is direct. Conversely, a high-scoring predicted target that does not change after perturbation may be buffered at the protein level, expressed in the wrong isoform, masked by competing RBPs, or regulated only under stress. The correct output of a prediction model is therefore a ranked, evidence-annotated hypothesis set.

Box 142.1. A Seed Match Is a Hypothesis, Not a Target

Render-ready box content

A seed match is best read as an address label, not a delivery receipt. The sequence says that a guide could in principle inspect that RNA segment. Before calling the gene a target, ask four questions: Is the expressed isoform carrying the site in the cell under study? Is a loaded Argonaute complex for the guide present at compatible abundance and localization? Is the site accessible rather than hidden by structure, ribosomes, or RNA-binding proteins? Does perturbing the guide alter the endogenous RNA, translation, or protein through that site? A ranked prediction can be useful before all answers are known, but the claim label should stay precise: “candidate site,” “occupied site,” “responsive transcript,” and “validated direct target” mean different things. This vocabulary prevents a long seed-match list from becoming a false mechanism.

142.2. siRNA, piRNA, and bacterial sRNA target prediction

Small interfering RNAs and miRNAs share Argonaute-mediated recognition, but their prediction rules differ because experimental and therapeutic siRNAs are often designed to cleave or strongly repress a specified target while minimizing unintended repression elsewhere. In many animal RNA interference experiments, the intended siRNA target is nearly perfectly complementary to the guide strand, whereas off-target repression can occur through miRNA-like seed pairing to unrelated transcripts. This dual logic means an siRNA design algorithm must evaluate both on-target potency and off-target risk. On-target features include guide-strand selection, thermodynamic asymmetry at duplex ends, target accessibility, GC content, avoidance of internal repeats or extreme stability, and compatibility with the relevant Argonaute. Off-target features include seed matches in 3′ UTRs, partial complementarity in coding regions, unintended near-perfect matches, immune-stimulatory motifs, chemical modifications, and tissue-specific transcript expression. Recent review work emphasizes that network-aware and artificial-intelligence approaches are increasingly used to mitigate small-RNA off-target effects, but the local bibliography does not yet include the major siRNA design and specificity papers needed for full curation (Bereczki et al.

Therapeutic siRNA prediction has additional constraints. A chemically modified siRNA delivered by GalNAc conjugation to hepatocytes or by a lipid nanoparticle has a different exposure profile from a transient transfection reagent in cultured cells. Chemical modifications can reduce immune sensing, improve stability, alter Argonaute loading, and reduce some off-target effects, but modifications can also change potency and strand bias. A therapeutic specificity model must consider the target tissue, expected intracellular concentration, duration of exposure, target knockdown needed for efficacy, and tolerated knockdown of off-target genes. A low-affinity off-target interaction may be irrelevant at experimental concentrations that resemble treatment, but important in an overexpression screen. Prediction scores should therefore be calibrated against dose and delivery context rather than treated as fixed properties of the guide sequence.

Figure 142.2. Distinct prediction objectives for miRNAs, siRNAs, piRNAs, and bacterial sRNAs

Figure 142.2. Distinct prediction objectives for miRNAs, siRNAs, piRNAs, and bacterial sRNAs. Makes clear why one prediction model cannot be transferred across small-RNA classes without changing assumptions and labels.

Plant miRNAs and plant siRNAs often show more extensive complementarity to their targets than typical animal miRNAs, and many guide cleavage events can be predicted by near-perfect pairing with positional mismatch rules. This makes computational target calling more alignment-like in plants, but not trivial. Target accessibility, transcript isoform annotation, sliced RNA decay products, translational repression, and noncanonical mismatches still matter. Degradome or parallel analysis of RNA ends data can identify cleavage products at the expected guide-pairing position and can strongly support direct cleavage. However, degradome signal is condition-specific, depends on RNA abundance and decay kinetics, and does not capture pure translational repression. A plant small-RNA prediction record should therefore distinguish cleavage-supported targets from predicted pairing-only targets.

Piwi-interacting RNAs are typically 24-31 nucleotide small RNAs associated with PIWI-clade Argonautes, especially in animal germlines. Many piRNAs guide recognition of transposon transcripts and other repetitive RNAs, supporting genome defense. Prediction is difficult because piRNA populations are large, piRNA clusters can produce many related guides, and transposon sequences are repetitive and fragmented in genome assemblies. A piRNA can have many genomic matches, and a transposon-derived transcript can contain many potential sites. Prediction must account for strand bias, ping-pong signatures, phased piRNA production, PIWI protein specificity, tissue and developmental stage, and whether the target is a nascent transcript, mature RNA, or chromatin-associated RNA. Extensive complementarity is often more informative than a short seed-like feature, but rules vary by lineage and PIWI complex.

PiRNA target prediction also illustrates the danger of incomplete target catalogs. If the genome assembly collapses repeats or omits polymorphic transposon insertions, a piRNA may appear to target a unique annotated gene when the true target is an unannotated repeat fragment. Conversely, genuine genic piRNA regulation has been reported in some systems, but genic predictions require careful exclusion of embedded transposon sequence, antisense transcription, and mapping artifacts. Evidence is stronger when PIWI binding, small-RNA orientation, target cleavage or transcriptional silencing, and genetic perturbation of the piRNA pathway agree.

Bacterial sRNA target prediction is a separate problem because bacterial regulatory sRNAs are usually longer than miRNAs and often interact with mRNAs through discontinuous base-pairing regions. Many trans-encoded bacterial sRNAs are imperfectly complementary to targets and require RNA chaperones such as Hfq or ProQ to stabilize interactions, remodel structures, or recruit decay factors. A common outcome is altered translation initiation: the sRNA can block the ribosome-binding site, expose an occluded ribosome-binding site, or change mRNA decay by recruiting or preventing RNase action. The predicted target site may lie in the 5′ UTR, around the start codon, inside a coding region, or in an intercistronic region of a polycistronic transcript. Because bacterial transcription units and 5′ ends are condition-dependent, accurate transcript boundary annotation is essential.

Table 142.2. Target-prediction contrasts across small-RNA pathways. Prevents inappropriate transfer of animal miRNA assumptions to siRNA, piRNA, and bacterial systems.

Pathway Guide length Main effector Typical pairing pattern Common target region Output Major prediction caveats
Animal miRNA Usually about 20-24 nt. AGO-loaded miRNA-induced silencing complex. Seed-centered partial pairing, with context and occasional 3′ supplementary pairing. Mostly 3′ UTRs, with some coding-region or 5′ UTR sites. Modest mRNA destabilization and translational repression. Seed matches are common; isoform, expression, accessibility, and CLIP context determine usefulness.
Therapeutic or experimental siRNA Usually about 21-23 nt duplex guides. AGO2-containing RNA-induced silencing complex in many animal systems. Near-perfect pairing for intended target plus miRNA-like seed pairing for off-targets. Intended transcript site, unintended 3′ UTR seed matches, and near-perfect off-targets. Target cleavage or strong knockdown; possible seed off-target repression. Dose, delivery, chemistry, guide-strand loading, immune motifs, and tissue expression alter specificity.
Plant miRNA/siRNA Often 20-24 nt, depending on pathway. Plant AGO complexes. More extensive complementarity than typical animal miRNAs; positional mismatches matter. Coding regions or UTRs of target transcripts; some chromatin-linked small-RNA targets. Slicing, translational repression, or pathway-specific silencing. Degradome signal is condition-dependent and cleavage evidence does not cover pure translational repression.
piRNA Commonly 24-31 nt. PIWI-clade Argonaute proteins. Extensive complementarity, strand bias, ping-pong or phased signatures in some systems. Transposon and repeat-derived RNAs; possible nascent or chromatin-associated transcripts. Slicing, transposon repression, transcriptional silencing, or pathway amplification. Repeat mapping, incomplete assemblies, tissue specificity, and PIWI-complex diversity can misassign targets.
Bacterial Hfq-associated sRNA Often 50-300 nt sRNA with a shorter pairing region. Hfq-associated sRNA-mRNA complexes and decay or translation factors. Imperfect, sometimes discontinuous base pairing aided by Hfq. 5′ UTRs, ribosome-binding sites, start-codon regions, coding segments, or operon junctions. Translation repression or activation and altered mRNA stability. Transcript boundaries, structure, Hfq binding, and stress condition must match the prediction.
Bacterial ProQ-associated sRNA Often 50-300 nt regulatory RNA. ProQ-associated sRNA or mRNA complexes. Pairing and structural features can differ from Hfq-centered rules. Structured UTRs, coding regions, or mRNA segments detected in ProQ interaction maps. Translation or stability changes, including activation in some contexts. Hfq-trained models may miss ProQ targets; validation needs ProQ occupancy and perturbation data.

Bacterial prediction tools use hybridization energy, seed-like sRNA regions, accessibility, conservation, and chaperone-binding information. Conservation can be helpful when related bacteria share sRNAs and target regions, but it is complicated by rapid evolution of intergenic regions and by horizontal gene transfer. Hfq-bound sRNAs often have accessible seed regions and target AU-rich or structurally exposed mRNA segments; ProQ-associated RNAs can have distinct structural features.

The mechanistic endpoint also differs among small-RNA classes. A mammalian miRNA site often produces modest repression by mRNA destabilization and translational effects. A designed siRNA may produce strong Argonaute2-mediated cleavage of a fully complementary target. A piRNA may trigger slicing, transcriptional silencing, or transposon pathway amplification. A bacterial sRNA may repress translation, activate translation, expose decay sites, protect an mRNA, or coordinate an operon response. The prediction model must therefore define the expected regulatory output. A tool trained to predict mammalian miRNA downregulation is not transferable to bacterial sRNA activation without changing both features and labels.

Box 142.2. Before Transferring a Predictor Across Pathways

Render-ready box content

Do not move a predictor to a new small-RNA pathway until the target, effector, and label have been redefined. For animal miRNAs, a positive label may mean modest repression from seed-centered 3′ UTR pairing. For a therapeutic siRNA, the same seed behavior is an off-target liability, while near-perfect complementarity to the intended transcript is desired. For piRNAs, repeat-aware mapping, PIWI identity, strand bias, and germline context may dominate. For bacterial sRNAs, the meaningful target may be a start-codon region whose accessibility is changed by Hfq or ProQ. A model can share subroutines, such as hybridization energy or accessibility scoring, across systems, but it cannot share assumptions unchanged. First specify the biological output, then choose features and validation labels that match that output.

142.3. Network inference, repression models, and target-site competition

A small-RNA regulatory network is a graph in which nodes represent small RNAs, target RNAs, proteins, pathways, cell states, or phenotypes, and edges represent regulatory relationships. In a simple diagram, a miRNA points to many mRNAs and each mRNA receives input from several miRNAs. In a real cell, the edges have weights, signs, delays, and context restrictions. A miRNA-target edge can mean predicted seed pairing, Argonaute occupancy, mRNA decay after perturbation, protein repression, or phenotypic dependence; these are not equivalent evidence types. Network inference therefore begins by defining what an edge means and by keeping separate layers for physical interaction, molecular response, and functional consequence.

Integrated transcriptome and small-RNA sequencing studies often construct networks by combining differentially expressed miRNAs, inversely correlated mRNAs, and target-prediction databases. This approach can generate useful hypotheses, especially in non-model organisms or developmental systems where targeted experiments are limited. For example, integrated small-RNA and transcriptome analysis has been used to propose miRNA-mediated regulatory networks during floral bud break in Prunus mume (Zhang et al. 2022). Such studies are valuable as discovery maps, but their edges are usually not established direct targets unless site-level or perturbation evidence is added. Inverse correlation can arise from direct miRNA repression, shared upstream transcriptional regulation, changing cell-type composition, developmental timing, stress response, or batch effects.

Figure 142.3. Evidence-layered small-RNA regulatory network

Figure 142.3. Evidence-layered small-RNA regulatory network. Counters overinterpretation of dense network diagrams in integrated omics papers.

The first quantitative layer is a repression model for a single site. A minimal model specifies the concentration of guide-loaded effector, the concentration of target RNA containing the site, the binding affinity or occupancy probability, the repression efficiency when occupied, and the decay or translation parameters of the target. If the small RNA is abundant and the site has high affinity, occupancy can be high and repression measurable. If the target site is inaccessible or the guide is scarce, the same sequence match may have little effect. If a transcript has multiple independent sites, repression can be approximately additive over modest ranges, but saturation, cooperative binding, and site masking can create nonlinear behavior. A model that lacks abundance cannot distinguish a rare but high-affinity interaction from a common regulatory edge.

Perturbation data add directionality but also indirect effects. Overexpressing a miRNA mimic can downregulate many transcripts with seed matches, including targets that would not be regulated at endogenous miRNA levels. Inhibiting a miRNA can reveal targets that are repressed under baseline conditions, but inhibitor chemistry, incomplete inhibition, and compensation by related miRNA family members complicate interpretation. Genetic deletion of a miRNA locus can be cleaner, yet developmental compensation or cell-state changes may dominate late measurements. Time-resolved data are particularly useful: direct mRNA destabilization often appears earlier than secondary transcriptional cascades, while protein-level effects may lag behind mRNA effects. Network inference should therefore weight early, site-enriched responses more strongly than late broad pathway changes when assigning direct edges.

Target-site competition arises because many RNAs can bind the same small RNA family. Competition is easiest to understand through mass action. If the total amount of active miRNA-loaded Argonaute is limited, adding more effective target sites can reduce occupancy of other targets. The magnitude of competition depends on the abundance of the competitor, the number and affinity of its sites, the abundance of the miRNA, the abundance and affinity of other targets, and the kinetics of repression and turnover. Competition is negligible when the competitor contributes only a tiny fraction of total effective binding capacity. Competition can become measurable when a competitor is very abundant, has many high-affinity sites, or is artificially overexpressed, as in engineered sponges or target mimics.

Figure 142.4. Mass-action logic of target-site competition

Figure 142.4. Mass-action logic of target-site competition. Prepares readers for the ceRNA stoichiometry section.

Network models often use ordinary differential equations, Bayesian networks, regression models, graphical models, or machine-learning predictors. Each framework answers a different question. Differential-equation models are useful when the analyst wants explicit parameters for production, decay, binding, repression, and competition. Regression models can estimate associations between miRNA expression, site counts, and mRNA abundance across samples, but they are sensitive to confounding by transcriptional programs and cell-type mixture. Bayesian or graphical models can represent uncertainty and conditional dependencies, but inferred edges still need mechanistic validation. Deep-learning models may improve prediction from high-dimensional features, yet they can hide whether a prediction depends on seed biology, expression correlation, or dataset artifacts. Bereczki et al. (2025) discuss network theory and artificial-intelligence approaches for small-RNA off-target mitigation, which is directly relevant to this modeling layer.

A useful network should also represent redundancy. Many miRNAs belong to families sharing the same seed, so perturbing one family member may have little effect if other family members remain active. Conversely, a transcript may contain sites for several unrelated miRNAs and be robustly repressed only when a cell expresses a particular combination of guides. Redundancy appears in bacterial systems too: multiple sRNAs can regulate the same mRNA under different stresses, and one sRNA can coordinate a regulon of transporters, metabolic enzymes, and regulators. A graph that displays only one guide per target can miss this combinatorial control.

Network inference becomes stronger when it combines orthogonal evidence. A high-confidence direct miRNA-target edge may have a conserved 8mer site, Argonaute CLIP support in the cell type, downregulation after miRNA mimic transfection, upregulation after miRNA inhibition, reporter repression by the wild-type 3′ UTR, loss of repression after site mutation, rescue after compensatory miRNA mutation, and protein-level response. Few edges have all this evidence, but the evidence ladder clarifies uncertainty. In disease and drug-response networks, such as studies of exosomal miRNAs in tumor-macrophage communication, network diagrams should distinguish measured signaling events from predicted target edges and from therapeutic interpretation (Yi et al. 2024).

The same principles apply to emerging small-RNA classes such as tRNA-derived small RNAs. A time-resolved study of tRNA-derived small RNA regulation after epileptogenesis can reveal dynamic associations and candidate regulatory modules, but assigning direct targets requires target-recognition rules, binding evidence, and perturbation-response logic specific to that RNA class (Zaheer et al. 2025). The chapter therefore uses “small-RNA regulatory network” broadly but refuses a one-size-fits-all target-prediction grammar.

142.4. ceRNA hypotheses, stoichiometry, and evidence thresholds

The competing endogenous RNA hypothesis proposes that transcripts sharing miRNA response elements can regulate one another by competing for miRNAs. A long noncoding RNA, circular RNA, pseudogene transcript, viral RNA, or protein-coding mRNA could in principle act as a competitor if it carries effective sites for a miRNA that also represses another transcript. The idea is mechanistically plausible because target-site competition exists. The controversy concerns generality and evidence: many reported ceRNA networks are built from shared predicted seed sites and expression correlations, but those features alone do not show that one transcript sequesters enough miRNA to derepress another transcript in a cell.

Stoichiometry is the first threshold. A ceRNA must supply enough effective binding sites to change miRNA availability. “Enough” is not a fixed molecule count; it depends on miRNA abundance, Argonaute loading, site affinity, site accessibility, target turnover, and the abundance of all other competing sites in the cell. A transcript expressed at one copy per cell with one weak site is unlikely to affect repression by a miRNA present in thousands of loaded Argonaute complexes. A circular RNA with dozens of high-affinity sites, a viral transcript expressed to high levels, or an engineered sponge can be much more plausible. Absolute quantification is therefore more informative than fold-change expression. A two-fold increase from one to two molecules per cell may be statistically significant but stoichiometrically irrelevant.

Box 142.3. A Quick ceRNA Stoichiometry Sanity Check

Render-ready box content

A quick ceRNA sanity check asks whether the proposed competitor can change the effective binding-site pool by enough to matter. Estimate four quantities in the same cell state: loaded miRNA molecules, competitor molecules, effective sites per competitor molecule, and the abundance of other high-affinity sites. A transcript present at two copies per cell with one weak site adds little capacity to a miRNA family present in thousands of AGO complexes and already bound by many mRNAs. A circular RNA, viral RNA, or engineered sponge expressed at high copy number with many accessible sites is more plausible. This calculation does not prove competition; it only decides whether the mechanism deserves stronger testing. If absolute abundance, localization, and site effectiveness are unknown, a ceRNA edge should remain a hypothesis generated from correlation and shared sites.

Table 142.3. Evidence thresholds for ceRNA claims. Gives an explicit rubric for evaluating ceRNA papers and network claims.

Evidence requirement Minimum useful measurement Common weak substitute Interpretation if missing
Co-expression Competitor, target, and relevant miRNA measured in the same cell type and condition. Bulk positive correlation between two transcripts. The proposed competition edge is biologically implausible or only hypothesis-generating.
Shared effective sites Specific miRNA response elements with conservation, AGO occupancy, accessibility, or reporter support. Database-predicted seed sharing alone. The model lacks evidence that the same guide can engage both RNAs.
Absolute abundance Molecule counts or calibrated abundance for miRNA-loaded AGO, competitor, target, and other site pools. Fold-change expression without copy-number scale. Stoichiometric plausibility cannot be judged; low-copy competitors may be irrelevant.
Localization Evidence that competitor, target, miRNA, and AGO occupy the same compartment. Whole-cell expression data. A nuclear, cytoplasmic, granule-localized, or extracellular mismatch can invalidate the mechanism.
miRNA dependence Competitor effect disappears or changes when the relevant miRNA or AGO pathway is perturbed. Correlation with miRNA expression. Derepression may reflect transcriptional co-regulation or unrelated cell-state effects.
Site mutation Loss of competitor effect after mutating the proposed sites while preserving RNA level and localization. Reporter repression with an intact overexpressed fragment. The causal site is untested; effects may come from RNA abundance or other elements.
Physiological dose-response Perturbation across endogenous or disease-relevant expression ranges with matched target readout. High-copy plasmid or mimic overexpression. The experiment shows possible competition at high dose, not endogenous ceRNA action.
Confounder control Controls for transcription factors, copy number, cell cycle, immune infiltration, batch effects, and cell mixture. Unadjusted correlation network. Shared upstream regulation can explain the apparent ceRNA relationship.
Phenotypic rescue Site-dependent molecular rescue linked to the claimed phenotype. Phenotype after competitor overexpression alone. The phenotype may be indirect and should not anchor a mechanistic ceRNA claim.

The second threshold is site effectiveness. A ceRNA model should identify the actual miRNA response elements that mediate competition. Shared database predictions are too permissive because short seed matches are common. Evidence is stronger when the proposed competitor contains conserved or CLIP-supported sites, mutation of those sites abolishes the competitor effect, and introducing compensatory mutations in the miRNA restores the relationship. For circular RNAs and long noncoding RNAs, subcellular localization matters. A nuclear lncRNA cannot sponge a cytoplasmic miRNA-loaded Argonaute pool unless the relevant interaction occurs in the nucleus and the effector is present there. A cytoplasmic circular RNA can be a stronger candidate, but circularity alone does not imply sponge function.

The third threshold is dose-response within the physiological range. Many ceRNA-like effects can be forced by overexpressing a transcript far above endogenous abundance. Such experiments show that a sequence can compete when supplied at high dose, not that endogenous transcript variation regulates another gene. A strong study perturbs the ceRNA across endogenous or disease-relevant ranges, measures the relevant miRNA and target abundance, and shows that the effect disappears when miRNA binding sites are mutated. Rescue experiments should preserve overall transcript expression and localization so that loss of effect is attributable to site loss rather than RNA instability or altered cell state.

Figure 142.5. Evidence thresholds for a credible ceRNA mechanism

Figure 142.5. Evidence thresholds for a credible ceRNA mechanism. Converts a controversial literature area into explicit evidence thresholds.

The fourth threshold is exclusion of confounding. Positive correlation between two mRNAs can reflect shared transcription factors, copy-number variation, cell-cycle state, immune infiltration, differentiation, or technical batch effects. Negative correlation between a miRNA and an mRNA can reflect direct repression, but also upstream regulation or cell-type mixture. A ceRNA claim should therefore include controls for transcriptional co-regulation and sample composition. Perturbing the proposed ceRNA should change the target in a miRNA-dependent manner, and perturbing the miRNA should alter or eliminate the ceRNA-target relationship. Network studies that infer ceRNA modules from bulk cancer datasets are hypothesis-generating unless they address these confounders experimentally or with strong causal designs.

Target-site competition can also occur without satisfying the full ceRNA narrative. Artificial miRNA sponges, target mimics, and high-dose therapeutic oligonucleotides can sequester small RNAs or saturate effector proteins. Viral RNAs can present abundant binding sites that alter host miRNA availability. Transfected siRNAs can compete for Argonaute loading or produce seed-mediated repression of unintended targets. These cases are important for biotechnology and therapeutics, but they should not be generalized to ordinary endogenous transcripts without abundance and site-level evidence. The phrase “sponge” should be reserved for cases where sequestration is supported; otherwise “contains predicted sites for” or “may compete for” is more accurate.

The ceRNA field also illustrates how network visualizations can inflate confidence. A ceRNA diagram often connects a lncRNA, several miRNAs, and many mRNAs with arrows or inhibition bars. Unless edge types are distinguished, the reader may infer that every line represents direct measured regulation. A better diagram encodes predicted seed sharing, expression correlation, physical binding, perturbation response, site-mutant evidence, and phenotypic rescue as separate layers or confidence levels. Chapter 142 uses this standard because it prevents weak computational edges from being mistaken for established mechanisms.

Recent small-RNA off-target and network-theory reviews support a cautious quantitative treatment of target competition, especially where artificial guides or therapeutic small RNAs can perturb networks beyond their intended targets (Bereczki et al. 2025). The local bibliography does not yet include the main ceRNA stoichiometry critiques, mathematical models, or high-confidence endogenous examples.

142.5. Experimental validation, benchmarking, and false positives

Experimental validation begins by matching the assay to the claim. A luciferase reporter assay can show that a candidate site is sufficient to confer repression in a synthetic context. It does not prove that the endogenous transcript is regulated, because the reporter may use a different promoter, copy number, UTR architecture, RNA structure, localization, and decay environment. Mutation of the site in the reporter is stronger than reporter repression alone, and compensatory mutation of the guide is stronger still. Endogenous genome editing of the target site is more decisive, especially when it preserves the rest of the transcript and shows loss of small-RNA-dependent regulation in the native locus. For bacterial sRNAs, compensatory mutations in the sRNA and target pairing region are a classic way to demonstrate direct base-pairing logic.

Perturbing the small RNA is necessary but not sufficient. A miRNA mimic can create nonphysiological guide abundance and seed off-target effects. An antagomir or locked nucleic acid inhibitor can have incomplete specificity, different potency across family members, and chemistry-dependent effects. siRNA knockdown of a pathway component can trigger stress or innate immune responses. Deleting a small-RNA gene can alter development or cell composition before target measurement. Therefore validation should include multiple perturbation types when possible, dose-response analysis, time-course measurements, and rescue. For therapeutic siRNAs and mimics, off-target and immune effects are not side details; they are central safety and mechanism questions (Bereczki et al. 2025).

Figure 142.6. Validation ladder for small-RNA target claims

Figure 142.6. Validation ladder for small-RNA target claims. Gives readers an immediately reusable standard for judging claims.

Direct physical interaction assays answer a different part of the evidence ladder. Argonaute CLIP, PIWI CLIP, Hfq or ProQ interaction maps, RNA-RNA ligation methods, and chimeric read assays can identify candidate contact sites. These methods are powerful because they locate interactions in cells, but they have false positives and false negatives. Crosslinking favors certain nucleotides and protein-RNA geometries. Immunoprecipitation can bring along indirect interactors. Ligation efficiency varies by RNA ends and structure. Highly expressed RNAs can appear frequently even when occupancy is weak. Mapping short reads to repeats can be ambiguous, which is especially serious for piRNA targets. Physical binding should be integrated with sequence logic and perturbation response rather than treated as a standalone proof of functional repression.

Molecular response assays include RNA-seq, quantitative PCR, ribosome profiling, proteomics, decay-rate measurements, and polysome profiling. RNA-seq detects mRNA abundance changes, which are prominent for many animal miRNA targets because repression often involves deadenylation and decay. Ribosome profiling can reveal changes in translation efficiency, but altered ribosome occupancy can reflect changes in initiation, elongation, mRNA abundance, stress, or ribosome pausing. Proteomics is close to functional output but less sensitive for small changes and can lag behind RNA regulation. Decay-rate measurements can separate altered RNA stability from altered transcription. A well-supported target claim specifies which molecular layer changed and avoids treating mRNA, ribosome occupancy, and protein abundance as interchangeable.

Benchmark construction determines whether a prediction method is actually useful. Positive examples can come from reporter-validated sites, endogenous site editing, CLIP-supported chimeric interactions, degradome-supported cleavage sites, or perturbation-responsive genes with site evidence. These positives are not equivalent. A model trained on reporter assays may learn reporter-friendly context, while a model trained on CLIP peaks may learn binding rather than repression. Negative examples are even harder. A site with no observed response is not necessarily a true negative if the small RNA was absent, the assay lacked power, or the condition was wrong. Matched negatives should control for UTR length, expression, nucleotide composition, conservation, and assay coverage. Random transcript regions are usually too easy and can exaggerate performance.

Table 142.4. Benchmark hazards in small-RNA target prediction. Makes benchmarking failures concrete for computational readers.

Hazard Example Effect on model performance Mitigation
Easy random negatives Random transcript regions lacking expression or seed matches. Inflates accuracy and AUROC because negatives are biologically obvious. Use matched negatives for UTR length, expression, GC content, conservation, and assay coverage.
miRNA-family leakage Related guides with the same seed split across training and test sets. Model memorizes seed-family behavior instead of generalizing. Split by miRNA family or seed group and report family-held-out performance.
Target-gene leakage Records from the same gene, UTR, or isoform appear in both train and test sets. Transcript-specific features make test performance look better than deployment performance. Split by target gene or transcript and collapse redundant site records before training.
Homologous UTR leakage Orthologous or paralogous UTRs share conserved target-site patterns across species splits. Conservation becomes a hidden identifier for known targets. Use species-aware and homology-aware splits, or evaluate transfer on withheld lineages.
Duplicated database labels The same validated interaction enters multiple databases or assay compilations. Duplicates overweight popular interactions and reduce independence. Deduplicate by guide, gene, site, assay, publication, and genomic interval.
Assay-specific labels CLIP peaks, reporter positives, and perturbation-responsive genes are merged as one positive class. Model learns assay bias, such as binding or reporter context, instead of repression. Define task-specific labels and evaluate binding, molecular response, and site necessity separately.
Expression confounding Highly expressed targets are easier to detect in CLIP, RNA-seq, and proteomics. Model may rank detectability rather than regulatory strength. Match or model expression, require relevant condition coverage, and report calibration.
Repeat-mapping ambiguity piRNA or repeat-derived target reads map to multiple genomic loci. False genic targets or inflated interaction counts can appear. Use repeat-aware mapping, ambiguity flags, and independent pathway evidence.
Unbalanced metrics Many easy negatives make AUROC high while top predictions remain false-positive rich. Reported performance overstates usefulness for validation prioritization. Include precision-recall curves, top-k precision, calibration, and matched negative baselines.

Data splitting is a common source of inflated accuracy. If related miRNAs with the same seed appear in both training and test sets, a model can perform well by memorizing seed-family behavior. If the same gene, UTR, or experimental dataset contributes records to both sets, the model can learn transcript-specific or study-specific features. Homology across species can leak conserved site patterns. A strong benchmark separates data by miRNA family, target gene, species, experiment, and time when appropriate. Performance should be reported with precision-recall curves, calibration, and top-k precision because users often care about the top candidates for validation. Area under a receiver operating characteristic curve can look high when negatives are abundant and easy, even if the top predictions contain many false positives.

False positives arise from biology and computation. Biology produces false positives when predicted sites are inaccessible, unexpressed, masked by proteins, absent from the expressed isoform, or too weak to matter. Computation produces false positives through permissive seed scanning, outdated annotations, mapping ambiguity, duplicated database records, training-test leakage, and uncorrected multiple testing. Study design produces false positives when overexpression perturbs cells beyond physiological ranges, when cell lines differ from the claimed tissue, when bulk expression correlations reflect cell-type mixture, or when a phenotypic assay is interpreted without molecular rescue. The false-positive problem is not a reason to avoid prediction; it is a reason to assign evidence grades.

A practical evidence grade can use five levels. Level 1 is sequence-only prediction with no context filter. Level 2 adds expression, conservation, or accessibility support. Level 3 adds physical binding or interaction-map support in a relevant condition. Level 4 adds perturbation response at RNA, translation, or protein level with site enrichment. Level 5 adds site-specific necessity and rescue at the endogenous locus or an equivalently rigorous system. Many exploratory networks should remain at Levels 1-3. Mechanistic claims about direct regulation should usually require Level 4 or Level 5, especially when the claim supports disease mechanism, therapeutic target selection, or network rewiring.

The final caution is that prediction algorithms encode the evidence available when they were built. A model trained mainly on mammalian 3′ UTR miRNA repression will underrepresent coding-region sites, noncanonical interactions, cell-type-specific isoforms, non-mammalian systems, and RNA classes whose targeting rules remain poorly characterized. A bacterial sRNA model trained on Hfq-dependent repression may miss ProQ-associated activation. A piRNA model trained on one animal germline may not transfer to another lineage. A modern small-RNA target-prediction result should therefore be reported with its domain of validity, feature assumptions, benchmark design, and validation plan.

Recent Consensus

The current consensus is that small-RNA target prediction is useful for prioritization but insufficient for mechanism by itself. Seed pairing remains a dominant feature for animal miRNA and siRNA off-target prediction, extensive complementarity is central for many plant small-RNA and piRNA target calls, and hybridization-accessibility logic is central for bacterial sRNA prediction. Across systems, context determines whether a predicted site is functional: cell-type-specific expression, transcript isoforms, RNP occupancy, target abundance, and pathway state all alter predictions. Network and AI approaches can improve prioritization and off-target mitigation, but they must be benchmarked against independent and biologically matched validation data (Bereczki et al. 2025).

Consensus also supports separating direct small-RNA targeting from downstream expression changes. The field has strong tools for generating candidate networks, including prediction databases, CLIP-family assays, RNA-RNA ligation maps, transcriptomics, proteomics, degradome data, and perturbation screens. No single method is a gold standard for all target classes. Strong inference comes from convergent evidence and from site-specific perturbation.

Open Questions, Controversies, Deprecated Models, and Common Misconceptions

Open questions:

  • How transferable are quantitative rules for noncanonical miRNA sites, deep-learning target predictors, piRNA target prediction, bacterial sRNA prediction, and ceRNA stoichiometry across contexts?

Common misconceptions:

  • “A predicted seed match is not a validated miRNA target.” A downregulated transcript after miRNA overexpression is not automatically direct.
  • “A CLIP peak is binding evidence, not necessarily repression evidence.” A shared miRNA site does not establish a ceRNA relationship. A visually dense network does not imply high-confidence mechanism. A high machine-learning score does not remove the need to know the training data, benchmark split, and biological context.