Chapter 136. RNAi, CRISPR, Perturb-Seq, Reporter Libraries, and Pooled Functional Screens

Scope Note

This chapter explains pooled and arrayed functional screens that perturb RNA molecules, genes, regulatory sequences, or cellular states and recover causal hypotheses by sequencing, sorting, reporter measurement, imaging, or single-cell profiling. The chapter focuses on RNA-relevant screening platforms: RNA interference (RNAi) screens using small interfering RNAs (siRNAs) or short hairpin RNAs (shRNAs); CRISPR knockout, CRISPR interference (CRISPRi), and CRISPR activation (CRISPRa) screens; Cas13 and other RNA-targeting perturbation systems; reporter libraries and massively parallel reporter assay (MPRA)-style sequence-function experiments; Perturb-seq, CROP-seq, and related single-cell pooled screens; and statistical, validation, and reproducibility standards that distinguish useful screening hypotheses from guide effects, seed toxicity, barcode artifacts, bottlenecks, and batch structure. It does not own the iterative binding- or reaction-based enrichment of functional nucleic-acid pools; in vitro selection and SELEX belong to Chapter 137.

Executive Summary

Functional screens convert many candidate mechanisms into a measurable experiment by assigning perturbations to cells or molecules, preserving a record of which perturbation was applied, exposing the system to a phenotype-defining measurement, and using sequencing or another high-throughput readout to infer which perturbations changed the outcome. The shared logic is simple: a perturbation has an intended target, an assay converts a cellular or molecular state into a readout, and statistical analysis compares observed readout changes with the variation expected from sampling, growth, delivery, barcode abundance, guide efficiency, and biological replication. The practical difficulty is that the perturbation identifier is not the same thing as the biological mechanism. A guide RNA, siRNA, shRNA, barcode, or reporter sequence can fail, act incompletely, act on unintended targets, or alter the assay system itself.

RNAi screens use endogenous small-RNA machinery to lower target RNA output. Synthetic siRNAs are often used in arrayed formats, whereas shRNAs and microRNA-adapted hairpins are commonly used in pooled viral libraries. RNAi is valuable when partial, reversible, or transcript-level depletion is biologically appropriate, but RNAi screens are especially vulnerable to microRNA-like off-target effects. The guide-strand seed region can repress many unintended transcripts, and some seed sequences are toxic because they collectively repress survival genes. RNAi hits therefore require independent reagents with different seeds, direct target knockdown measurement, and ideally rescue or orthogonal validation.

CRISPR screens use guide RNAs to direct Cas effectors. CRISPR knockout screens use nuclease-mediated DNA cutting and repair to disrupt genes. CRISPRi and CRISPRa use catalytically inactive Cas proteins with repressive or activating domains to lower or raise transcription without cutting DNA. These modalities are not interchangeable. Knockout screens can produce strong loss-of-function genetics for protein-coding genes, but they can be confounded by DNA-break toxicity, copy-number amplification, exon choice, repair outcome, protein turnover, and effects on overlapping regulatory DNA. CRISPRi and CRISPRa are often better for dosage-sensitive genes, essential genes, promoters, enhancers, and some noncoding loci, but their activity depends on transcription start site annotation, chromatin context, effector expression, and local regulatory architecture.

Cas13 and related RNA-targeting CRISPR systems direct programmable effectors to RNA molecules rather than genomic DNA. These systems are attractive for viral RNAs, transcript isoforms, long noncoding RNAs, circular RNAs, and other cases where the RNA product must be separated from the DNA locus. Their screen interpretation depends on RNA accessibility, isoform structure, target abundance, effector expression, guide design, subcellular localization, and possible collateral or transcriptome-wide effects that vary by effector and context.

Reporter libraries and MPRA-style assays test sequence-function relationships by linking many designed sequences or variants to reporter outputs. They can measure how 5′ untranslated regions, 3′ untranslated regions, splice elements, codons, RNA structures, enhancer fragments, promoters, and disease-associated variants affect expression. The evidence is powerful but context-limited: a reporter hit demonstrates activity in the reporter system, not automatic function at the endogenous locus. Strong interpretation requires multiple barcodes, DNA input normalization, replicate design, focused validation, and when needed endogenous editing or direct transcript measurement.

Perturb-seq, CROP-seq, and related single-cell pooled screens combine perturbation libraries with single-cell RNA-seq so that guide identity and transcriptome state are measured in the same cell or sample. These methods can connect genes and regulatory elements to pathways, cell states, differentiation decisions, stress responses, and RNA programs that simple survival screens would miss. They also add new artifacts: guide capture failure, multiplets, sparse transcript capture, ambient RNA, donor and batch effects, cell-cycle confounding, and shifts in cell-state composition. Reliable single-cell screening treats cells as observations nested within guides, samples, donors, and batches rather than as independent biological replicates.

Across all platforms, a screen hit is a prioritized hypothesis rather than a completed mechanism. Reproducible interpretation rests on library representation, appropriate controls, redundant reagents, prespecified contrasts, models suited to the statistical unit, correction for multiple testing, direct target engagement, and independent validation.

Concept Inventory

  • Pooled functional screen: an experiment in which many perturbations are mixed, applied to a population of cells or molecules, and tracked by sequencing perturbation identifiers such as guide RNAs, shRNAs, barcodes, or reporter tags.
  • Arrayed screen: a screen in which each perturbation is tested in a separate well, sample, or physical compartment. Arrayed screens are less scalable than pooled screens but often better for imaging, dose control, and direct phenotyping.
  • RNA interference (RNAi): small-RNA-guided post-transcriptional silencing in which an Argonaute-loaded guide RNA directs repression or cleavage of target RNA.
  • siRNA: a short RNA duplex, commonly designed to produce one guide strand that enters Argonaute and targets a complementary transcript.
  • ShRNA: a vector-expressed short hairpin RNA processed inside the cell into siRNA-like guide strands.
  • Seed sequence: the short guide-strand region, often described as positions 2-8, whose partial pairing can repress transcripts in a microRNA-like manner.
  • Seed toxicity: a reagent-level RNAi artifact in which a guide seed sequence broadly represses essential or survival-related transcripts independent of the annotated target gene.
  • CRISPR knockout screen: a screen in which guide RNAs direct a Cas nuclease to genomic DNA, and DNA repair creates alleles that may disrupt gene function.
  • CRISPRi: CRISPR interference, usually dCas9 fused to a repressive domain such as KRAB and guided near promoters or regulatory elements to reduce transcription.
  • CRISPRa: CRISPR activation, usually dCas9 or guide-recruitment architectures coupled to transcriptional activation domains to increase expression.
  • Cas13 screen: a screen using an RNA-guided RNA-targeting CRISPR effector to bind, cleave, deplete, image, or modulate RNA products.
  • Reporter library: a set of reporter constructs in which designed sequences or variants are linked to measurable reporter RNA, reporter protein, or barcode output.
  • MPRA: massively parallel reporter assay, a family of experiments that tests many regulatory sequences or variants in one reporter framework.
  • Perturb-seq: a single-cell pooled perturbation strategy that links perturbation identity to transcriptome-wide single-cell RNA-seq phenotypes.
  • CROP-seq: a guide-capture architecture for single-cell CRISPR screening; it is a specific method family rather than a generic term for all single-cell perturbation screens.
  • Hit: a perturbation, gene, sequence, guide, barcode, or state signature that passes predefined screen criteria. A hit is not final proof of mechanism.
  • Library representation: the number and distribution of cells, molecules, constructs, or sequencing reads assigned to each library member at each experimental step.
  • Orthogonal validation: follow-up using independent reagents, different perturbation modalities, rescue constructs, endogenous-locus assays, or direct molecular readouts.

What to Know Before Reading This Chapter

The reader should distinguish four linked but separable layers: perturbation, target engagement, phenotype, and inference. A perturbation is the reagent or construct introduced into the system: an siRNA, shRNA, sgRNA, Cas13 guide, reporter variant, barcode, or integrated cassette. Target engagement is the molecular change that the perturbation actually produces, such as reduced RNA abundance, indel formation, transcriptional repression, transcriptional activation, RNA cleavage, or altered reporter expression. A phenotype is the measured output, such as growth, survival, fluorescence, viral replication, splicing, protein abundance, reporter RNA abundance, or a single-cell transcriptome. Inference is the statistical and biological claim drawn from many perturbations and measurements.

This chapter uses a running example: a screen for host RNA regulators of antiviral response in human cells. RNAi could transiently knock down candidate RNA-binding proteins or viral RNA cofactors. CRISPR knockout could identify host genes required for viral replication or interferon-mediated survival. CRISPRi could partially repress essential genes without DNA cutting, and CRISPRa could activate interferon-stimulated genes. Cas13 could target viral RNA segments or host long noncoding RNAs. MPRA-style reporter libraries could test how 5′ untranslated regions, 3′ untranslated regions, codon composition, or RNA structures control translation and stability during immune activation. Perturb-seq could reveal whether each perturbation induces an interferon-high state, a stress response, a cell-cycle block, or a specific RNA-processing program. The example shows why the same biological question can require different perturbation layers and different validation standards.

136.1. RNAi screens and siRNA/shRNA design

An RNAi screen is a functional screen that uses small-RNA-guided silencing to reduce expression of many candidate target transcripts. The molecular effector is usually Argonaute, a small-RNA-binding protein that uses a guide strand to recognize RNA through base pairing. In experimental RNAi, the guide strand is supplied by a synthetic siRNA duplex or produced from a DNA-expressed shRNA. Synthetic siRNAs are often transfected into cells in arrayed plates, one reagent or reagent pool per well. shRNAs are often delivered by lentiviral vectors in pooled libraries, with the shRNA sequence itself or an associated barcode recovered by sequencing after selection.

The basic design question is whether a lower amount of a target RNA or its encoded product changes a phenotype. In the antiviral-response example, a library could target host RNA-binding proteins and ask which knockdowns reduce viral replication. A depleted viral reporter signal after knockdown of a host factor might indicate that the factor supports viral RNA synthesis, viral RNA stability, translation, or particle production. The first inference is only that an RNAi reagent changed the assay output. A stronger claim needs direct measurement that the intended host transcript was reduced and that independent reagents targeting the same host gene produce similar effects.

Figure 136.1. RNAi Screen Workflow from Reagent Design to Knockdown Validation

Figure 136.1. RNAi Screen Workflow from Reagent Design to Knockdown Validation. RNAi screens test whether reducing RNA output changes a phenotype. Synthetic siRNAs are often used in arrayed wells, while shRNAs are often delivered as pooled viral libraries. The workflow emphasizes guide-strand selection, target-site choice, library representation, phenotype measurement, direct RNA or protein knockdown measurement, independent reagents, and rescue when feasible.

RNAi reagent design begins with strand selection. An siRNA duplex has two strands, but usually only one strand is intended to act as the guide. Thermodynamic asymmetry at the duplex ends, sequence composition, and processing context influence which strand enters Argonaute. If the passenger strand loads, the reagent can target a different transcript set. shRNAs add another layer because the expressed hairpin must be transcribed, folded, exported or processed as appropriate for the expression system, cleaved by small-RNA biogenesis machinery, and loaded into Argonaute. Hairpin stems, loops, promoters, vector copy number, and microRNA-adapted scaffolds can all affect potency and toxicity.

Target-site choice also matters. A useful RNAi target site must be present in the transcript isoform expressed in the screen system. A guide designed against an exon absent from the relevant isoform will fail even if the gene annotation is correct in another cell type. A guide designed against a polymorphic region can lose complementarity in a particular cell line or donor. Target RNA structure and RNA-binding proteins can reduce accessibility, although accessibility rules are less deterministic for RNAi than for simple hybridization because Argonaute can sample many accessible or transiently accessible regions. Guides near repetitive sequence, paralogous genes, pseudogenes, or conserved domains can blur whether the screen perturbed one transcript or a family.

The major pedagogical point is that RNAi normally creates knockdown, not knockout. Knockdown means the target RNA or protein is reduced but not necessarily absent. Partial depletion can be a strength. It can reveal dosage-sensitive biology, permit study of essential genes, and mimic therapeutic reduction more closely than complete gene disruption. Partial depletion can also be a limitation. Stable proteins may persist after transcript knockdown; abundant RNAs may not fall below a functional threshold; feedback may compensate for reduced expression; and residual enzyme activity may be enough to preserve pathway flux. A negative RNAi screen result therefore does not prove that a gene is irrelevant unless target engagement and phenotypic sensitivity are established.

Arrayed and pooled RNAi screens answer different practical questions. Arrayed screens allow microscopy, time courses, multiple doses, and direct well-level measurements. They are useful for phenotypes such as viral inclusion formation, RNA localization, stress granule assembly, or splicing reporter changes. Pooled shRNA screens are efficient for growth, survival, drug resistance, fluorescence sorting, and long selections. They require low multiplicity of infection so most cells receive one hairpin, sufficient cells to preserve representation, and sequencing of initial and final libraries. If representation collapses during infection or selection, later read counts reflect sampling history rather than biology.

RNAi controls should match the screen mechanism. Non-targeting controls estimate delivery and sequencing background but do not model seed-mediated repression. Positive controls confirm that the assay can detect known pathway effects. Multiple independent siRNAs or shRNAs per gene are essential, but the reagents should not share the same seed sequence or the same unrecognized off-target pattern. A strong RNAi hit is supported by consistent phenotypes from independent guide strands, measurable knockdown of the intended RNA and ideally protein, dose or time dependence, and rescue by an RNAi-resistant cDNA or transcript when rescue is biologically feasible.

RNAi remains important even in the CRISPR era. Some cell types tolerate RNA transfection better than viral CRISPR delivery. Some biological questions require transient perturbation, partial depletion, or avoidance of DNA damage. Viral RNAs, transcript isoforms, and some noncoding RNAs can be more directly addressed by RNA-level perturbation than by genomic cutting. However, RNAi evidence should be interpreted as small-RNA perturbation evidence, not automatically as clean gene-specific loss of function.

Table 136.1. RNAi Reagent Design Variables and Controls. siRNA and shRNA sequence, strand bias, expression, dose, coverage, and control design affect both on-target activity and seed-mediated artifacts; screen interpretation requires reagent-level and target-level controls.

Design variable Why it matters Common failure mode Control or mitigation
Guide-strand selection Determines which strand enters Argonaute Passenger-strand off-targets Thermodynamic design, strand-bias testing, independent reagents
Target-site isoform Ensures the guide targets the expressed RNA Guide targets absent exon or polymorphic site Use system-specific transcript annotation and direct RNA measurement
Seed sequence Drives many microRNA-like off-target effects Shared seed creates false hits Use seed-aware design and independent seeds
shRNA expression level Controls processing and pathway burden MicroRNA pathway saturation or toxicity Use tuned promoters or microRNA-adapted scaffolds
Knockdown depth Determines whether residual product remains functional False negative for stable or abundant targets Measure RNA and protein where possible
Rescue design Tests target specificity Rescue changes dosage or isoform context Use RNAi-resistant construct matched to biology

136.2. Off-target modeling and seed toxicity

An off-target effect is an unintended effect of a perturbation reagent. In RNAi, the most important off-target class is seed-mediated repression. The guide-strand seed sequence is short enough that partial matches occur in many transcripts, especially in 3′ untranslated regions and accessible transcript regions. This means an siRNA or shRNA can behave partly like a microRNA: it can reduce many RNAs that were never intended targets. The consequence for screens is severe. A phenotype can be reproducible and statistically strong while still being caused by the seed sequence rather than by the annotated target gene.

Seed-mediated off-targeting is not random noise. It has sequence structure. Guides with the same seed can share transcriptome-wide effects even when they are assigned to different genes. Transcript abundance, site accessibility, RNA-binding proteins, 3′ untranslated region length, and cell state influence the magnitude of repression. A guide can therefore produce a context-specific phenotype through unintended repression of a pathway active in one cell type, stress condition, or viral infection state. This is why controls for RNAi screens should include seed-aware analyses rather than only non-targeting or scrambled controls.

Seed toxicity is a high-impact subset of seed-mediated off-targeting. A toxic seed is a seed sequence that broadly represses transcripts required for cell growth, survival, or homeostasis. In a negative-selection shRNA screen, toxic-seed hairpins drop out of the population and can mimic essential-gene hits. In a positive-selection screen, a seed that suppresses apoptosis, interferon signaling, or cell-cycle arrest can enrich even when the annotated target gene has nothing to do with the phenotype. Seed toxicity is not necessarily universal; the same seed can have different consequences in a cancer cell line, a primary immune cell, or a differentiated neuron because the expressed 3′ untranslated regions and survival circuits differ.

Figure 136.2. Seed-Mediated RNAi Off-Targeting and Seed-Toxicity Logic

Figure 136.2. Seed-Mediated RNAi Off-Targeting and Seed-Toxicity Logic. A guide-strand seed can pair with many unintended transcript sites and repress a network of RNAs in a microRNA-like manner. If a seed represses survival genes, multiple siRNAs or shRNAs with different annotated targets can behave as toxic reagents. Seed-aware analysis asks whether phenotypes track independent target-directed reagents or shared seed behavior.

Off-target modeling starts at the reagent level. Analysts can test whether phenotypes correlate better with intended gene identity or with guide seed identity. They can compare independent reagents with different seeds, examine enrichment of seed matches in downregulated transcripts, and ask whether a hit remains after removing reagents with known toxic or promiscuous seeds. When expression data are available, the transcriptome-wide pattern caused by an RNAi reagent can be compared with the expected downregulation of seed-matched transcripts. In arrayed screens, plate position, transfection toxicity, and concentration effects should be separated from sequence-specific effects.

Seed-aware design reduces but does not eliminate risk. Designers avoid guides whose seeds are predicted to target many essential transcripts or whose sequence composition is associated with toxicity. They also avoid immune-stimulatory motifs, repetitive regions, and passenger-strand loading risk. Chemical modifications can improve stability or reduce innate immune activation in synthetic siRNAs, but modifications can also change potency and strand bias. shRNA expression level must be controlled because excessive hairpin production can compete with endogenous microRNA processing and Argonaute loading.

The evidence basis for seed off-targeting comes from the combination of molecular small-RNA biology, expression profiling after siRNA treatment, computational seed-match analysis, and screen-level behavior. Direct biochemical proof that one seed site causes a screen phenotype is usually available only for focused follow-up. For screening-scale interpretation, the practical standard is convergence: independent reagents with distinct seeds, direct target knockdown, target-pathway consistency, and rescue. If those tests fail, the safer interpretation is “RNAi reagent phenotype” rather than “gene function.”

Do not overgeneralize the off-target problem in either direction. RNAi screens are not useless; many well-designed screens have generated biologically valuable hypotheses. But an RNAi hit is not strengthened merely by a small p value or a large fold change. The proper question is whether the phenotype follows the intended target after accounting for seed sequence, strand selection, knockdown efficiency, reagent concentration, and assay sensitivity.

136.3. CRISPR knockout, CRISPRi, and CRISPRa screens

CRISPR screens use programmable guide RNAs to direct Cas proteins to nucleic acid targets. In a CRISPR knockout screen, a single-guide RNA directs a Cas nuclease such as Cas9 to genomic DNA. The nuclease creates a double-strand break, and cellular repair can generate insertions, deletions, larger rearrangements, or other alleles that disrupt gene function. A pooled knockout screen typically delivers a library of guides to Cas-expressing cells at low multiplicity, allows editing and phenotypic selection, then counts guide abundance. Guides depleted from a final population can mark genes required for growth or survival; guides enriched after drug, toxin, viral infection, or immune pressure can mark resistance factors or pathway repressors.

The molecular unit in a knockout screen is not the RNA but the genomic locus. This distinction matters for RNA biology. If a guide cuts an early constitutive exon of a protein-coding gene, the resulting alleles may abolish protein function. If a guide cuts an alternatively spliced exon, a nonessential domain, a duplicated region, or a region subject to exon skipping, the gene may retain function. If a guide cuts near a noncoding RNA locus, the phenotype may reflect disruption of a promoter, enhancer, antisense transcript, host-gene intron, local chromatin state, or DNA damage response rather than loss of the RNA molecule itself. For long noncoding RNAs and enhancer RNAs, knockout-like edits must be interpreted with locus biology in mind.

Box 136.1. Locus Perturbation Is Not Automatically RNA Perturbation

Before calling a noncoding-locus hit an RNA-function hit, separate the tested layer from the biological claim. A Cas9 deletion can remove promoters, enhancers, splice sites, insulators, DNA motifs, or overlapping transcription units. CRISPRi can repress transcription through local chromatin changes. Cas13, RNAi, or antisense oligonucleotides can lower the RNA product without changing the DNA sequence. These perturbations are complementary, not interchangeable.

A stronger RNA-product claim asks four questions: Was the RNA molecule measurably depleted or altered? Did nearby gene expression, chromatin, or DNA damage change independently of RNA abundance? Do independent RNA-targeting reagents reproduce the phenotype? Can an RNA rescue restore the phenotype without replacing the DNA locus? If the answers diverge, the safer conclusion is a locus-associated phenotype, not a stand-alone RNA-product mechanism.

CRISPR knockout has characteristic artifacts. Cutting in amplified cancer-genome regions can cause toxicity because many DNA breaks occur at once. Guides vary in on-target cutting efficiency because of sequence, chromatin accessibility, target position, and mismatch tolerance. Repair outcomes vary across loci and cell types. Some in-frame indels preserve protein activity, while some frameshifts produce truncated products with residual or neomorphic effects. Protein turnover can delay phenotypes; essential genes can drop out before a condition-specific question is measured. These are not minor details. They determine whether guide depletion means loss of the intended gene product, generic DNA-break cost, or selection on edited allele types.

CRISPRi changes the perturbation layer from DNA damage to transcriptional repression. In many mammalian systems, catalytically inactive Cas9 is fused to a KRAB repressor domain and guided near a transcription start site. The complex recruits repressive chromatin machinery and lowers transcription. CRISPRi is useful when complete knockout is lethal, when dosage is the relevant variable, when genes have stable proteins but transcript reduction can be measured, or when a noncoding locus should be repressed without cutting DNA. For an antiviral screen, CRISPRi could partially reduce expression of essential RNA-processing genes and identify dosage-sensitive effects on viral replication without selecting immediately for complete gene loss.

CRISPRa performs the opposite kind of transcriptional modulation. dCas9 or guide-scaffold systems recruit activation domains near promoters or enhancers to raise expression. CRISPRa screens can identify genes whose increased expression promotes antiviral resistance, drives differentiation, induces toxicity, or rewires RNA-processing programs. They are especially valuable for gain-of-function questions that would be awkward to test by cDNA overexpression because the endogenous locus may preserve isoform choice, regulatory context, and transcript processing. However, CRISPRa only works when the locus is activatable in the cell state. It can activate neighboring transcripts, cryptic promoters, or regulatory cascades, so direct RNA measurement is part of the assay rather than a luxury.

Figure 136.3. Molecular Layers Perturbed by CRISPR and RNA-Targeting Modalities

Figure 136.3. Molecular Layers Perturbed by CRISPR and RNA-Targeting Modalities. Programmable perturbation systems act at different molecular layers. Cas9 knockout cuts genomic DNA and relies on repair. CRISPRi represses transcription, and CRISPRa activates transcription without DNA cutting. Cas13 and RNAi act on RNA products through different effectors. The layer perturbed determines the causal claim, artifact profile, and validation requirement.

Guide design differs across CRISPR modalities. Knockout guides often target early coding exons or functional domains while avoiding common variants, paralogs, and off-target DNA sites. CRISPRi guides usually tile a window around the transcription start site, where repression efficiency is most predictable. CRISPRa guides often target promoter-proximal regions or enhancer-linked sites where activator recruitment can engage transcription. For noncoding genes, guide placement should be chosen to distinguish RNA-product depletion, transcriptional interference, and local DNA regulatory effects. Tiling designs and multiple guide classes are often more informative than a small set of single guides.

The best CRISPR screen designs include positive controls, non-targeting or safe-targeting controls, multiple guides per gene or element, independent biological replicates, and initial and final library sequencing. They also include target-engagement measurement for selected hits: indel formation, protein loss, RNA reduction, or RNA induction as appropriate. A knockout screen hit should not be automatically equated with a drug target. A gene required for growth in one cell line may be a context-specific dependency, a DNA-break artifact, a synthetic lethal interaction with an unrecognized genotype, or a gene whose inhibition would be toxic in normal tissue. Screening is prioritization; validation supplies the causal and translational weight.

136.4. Cas13 and RNA-targeting perturbation systems

Cas13 systems are CRISPR effectors that target RNA rather than DNA. A Cas13 guide RNA pairs with an RNA target, and the effector can bind or cleave that RNA depending on the system and engineered construct. For RNA biology, the conceptual attraction is direct: Cas13 can perturb the RNA product without cutting the genomic locus. This is valuable for viral RNAs, transcript isoforms, long noncoding RNAs, circular RNAs, repetitive RNAs, and overlapping loci where a DNA edit or deletion would confound RNA-product function with regulatory DNA function.

In a Cas13 screen, each guide is designed against an RNA sequence. The readout can be target RNA abundance, cell survival, fluorescence, viral replication, reporter activity, or a single-cell transcriptome. For example, a pooled Cas13 library targeting a viral genome could ask which viral RNA segments or structures are most vulnerable to RNA-guided depletion. A host-focused library could target long noncoding RNAs induced by interferon and measure whether knockdown changes antiviral gene expression. Unlike a Cas9 knockout screen, the expected immediate target engagement is RNA depletion or RNA binding, not indel formation.

Cas13 guide design has several dependencies. The target sequence must be present in the relevant RNA isoform and exposed enough for guide pairing and effector access. RNA structure, RNA-binding proteins, translation, localization, and RNP assembly can make otherwise valid target sequences ineffective. Highly abundant RNAs may require strong guide activity to produce useful depletion. Nuclear RNAs, cytoplasmic RNAs, mitochondrial RNAs, and viral replication-complex RNAs may experience different effector access. Guide performance can therefore be strongly context-dependent even when sequence complementarity is perfect.

Cas13 effectors also have system-specific behaviors. Some Cas13 proteins can show collateral RNA cleavage after target recognition under some conditions, although the magnitude and practical consequences depend on effector family, expression level, guide, cell state, target abundance, and assay design. Engineered nuclease-dead variants can be used for RNA binding, imaging, recruitment, or base-editing-like functions when fused to other domains. These applications expand the perturbation space but add new controls: a binding-only effector can sterically block RNA processing, transport, or translation; a recruited domain can have local and indirect effects; and high effector expression can perturb cell physiology.

Table 136.2. Perturbation Modalities, Strengths, and Dominant Artifacts. RNAi, CRISPR knockout, CRISPRi, CRISPRa, Cas13, and reporter perturbations act on different molecules and yield different readouts; modality-specific artifacts require orthogonal validation.

Modality Primary molecule perturbed Typical readout Strengths Dominant artifacts or caveats Strong validation
siRNA RNAi Mature or cytoplasmic RNA through Argonaute Knockdown phenotype, imaging, reporter, growth Fast, arrayable, transcript-level, partial depletion possible Seed off-targets, passenger activity, immune activation, incomplete knockdown Independent seeds, RNA/protein measurement, RNAi-resistant rescue
shRNA RNAi Vector-expressed hairpin processed into small RNA Pooled enrichment or depletion Stable pooled delivery and long selections Seed toxicity, processing variation, microRNA pathway saturation Multiple hairpins, seed-aware analysis, orthogonal perturbation
CRISPR knockout Genomic DNA coding sequence or locus sgRNA enrichment or depletion Strong loss-of-function genetics DNA-break toxicity, copy-number artifacts, repair heterogeneity Multiple guides, edit/protein validation, rescue
CRISPRi Promoter or regulatory DNA controlling transcription Reduced RNA expression and phenotype Reversible repression; useful for dosage and essential genes TSS annotation, chromatin context, incomplete repression Target RNA measurement and guide tiling
CRISPRa Promoter or enhancer-proximal regulatory DNA Increased RNA expression and phenotype Gain-of-expression screening Activatability limits, neighboring or cryptic activation RNA/protein induction and independent guides
Cas13 RNA product RNA knockdown, reporter, survival, single-cell state RNA-product targeting without DNA cutting RNA accessibility, effector behavior, possible collateral effects Independent guides, target RNA measurement, guide-resistant rescue
Reporter library or MPRA Designed regulatory sequence in reporter context Barcode, RNA, protein, or fluorescence output High-resolution sequence-function mapping Barcode imbalance and context limits Multiple barcodes, input normalization, endogenous follow-up
Perturb-seq or CROP-seq Gene, RNA, or regulatory element plus transcriptome state Single-cell RNA-seq with perturbation identity High-dimensional pathway and state phenotypes Guide capture failure, multiplets, sparse RNA, batch effects Guide QC, target engagement, donor-aware analysis

Cas13 screens should be validated at the RNA level. A guide that changes a phenotype but does not reduce or otherwise engage the target RNA is not strong evidence for target function. Independent guides against different accessible regions of the same RNA should converge. A guide-resistant RNA rescue, when feasible, can distinguish on-target RNA depletion from guide or effector artifacts. For viral targets, escape mutations and sequence variability should be considered. For host RNAs, isoform-specificity should be tested rather than assumed from gene-level annotation.

Cas13 should not be described as a universal replacement for RNAi. RNAi recruits endogenous Argonaute machinery and has microRNA-like off-target biology. Cas13 recruits a CRISPR effector with its own guide rules, expression constraints, and possible collateral effects. Both perturb RNA, but they perturb RNA through different mechanisms. Comparing RNAi, Cas13, antisense oligonucleotides, CRISPRi, and genomic deletion can be especially powerful when asking whether a noncoding RNA molecule functions independently of its DNA locus.

The evidence base for Cas13 screening is still less mature than for RNAi and Cas9 genome-wide screens. Guide-design rules, effector choice, delivery, collateral-effect boundaries, and transcript-context models are active areas of method development. A careful chapter or study should therefore state which Cas13 effector was used, how guides were designed, how target engagement was measured, whether transcriptome-wide side effects were assessed, and whether conclusions are specific to that cell type and expression system.

136.5. Reporter libraries and MPRA-style sequence-function assays

A reporter library tests sequence function by linking many designed sequences or variants to measurable reporter outputs. The perturbation is not necessarily a gene knockdown or knockout. Instead, the experiment asks whether a sequence element changes transcription, splicing, RNA stability, RNA localization, translation, protein output, or another reporter-linked molecular property. An MPRA is a reporter-library family in which many sequences are assayed in parallel, usually by sequencing reporter RNAs, associated barcodes, DNA input libraries, or sorted reporter populations.

For RNA biology, reporter libraries are useful because many RNA regulatory hypotheses are sequence-level hypotheses. A 5′ untranslated region can alter translation initiation through structure, upstream open reading frames, start-codon context, internal ribosome entry, or protein binding. A 3′ untranslated region can alter stability, localization, poly(A)-tail dynamics, or microRNA sensitivity. Exonic and intronic sequences can alter splicing. Coding sequences can alter translation elongation, mRNA stability, codon optimality, RNA structure, and protein output. Reporter libraries can test thousands to millions of designed variants that systematically change these features.

The common workflow has six steps. First, design the library: natural variants, saturation mutagenesis, tiling fragments, random sequences, motif perturbations, codon substitutions, structure-preserving variants, or clinically observed variants. Second, synthesize or clone the library. Third, link each designed sequence to one or more barcodes or sequence it directly. Fourth, deliver the constructs to cells, usually by plasmid transfection, viral delivery, landing-pad integration, or another controlled expression system. Fifth, measure DNA input and reporter RNA, reporter protein, fluorescence, or sorted bins. Sixth, normalize and model output by sequence, barcode, replicate, and covariates.

Figure 136.4. Reporter-Library and MPRA Workflow for RNA Regulatory Sequence-Function Assays

Figure 136.4. Reporter-Library and MPRA Workflow for RNA Regulatory Sequence-Function Assays. Reporter-library experiments test many designed sequences or variants in parallel. Each element is linked to one or more barcodes, cloned into a reporter context, delivered to cells, measured by DNA and RNA sequencing or reporter protein output, and analyzed after input normalization. Multiple barcodes and focused follow-up distinguish sequence activity from barcode artifacts and context effects.

Barcode design is a central issue. Multiple barcodes per sequence help separate true sequence activity from barcode-specific effects, synthesis errors, cloning bias, and amplification differences. DNA input counts estimate how many constructs were present, while reporter RNA counts estimate expression or stability in the assay context. RNA/DNA ratios, barcode aggregation, replicate models, and filters for low-count elements are common. If protein output or fluorescence is measured, RNA-level and protein-level outputs should be distinguished because a sequence could increase RNA abundance while reducing translation efficiency, or vice versa.

Reporter context determines what the assay can claim. A plasmid MPRA does not reproduce native chromatin, three-dimensional enhancer contacts, endogenous transcription, full-length RNA processing, natural RNA localization, or physiological copy number. Integrated reporter assays improve some context issues but still isolate the tested sequence from much of its native locus. A 3′ untranslated region fragment may behave differently when attached to a short reporter than when embedded in a long endogenous transcript with alternative polyadenylation, RNA-binding protein occupancy, and localization signals. A codon-optimization library may report on a particular coding sequence, cell type, and expression level rather than a universal codon rule.

Box 136.2. Reporter Hits Are Contextual Claims

Interpret reporter-library results at the level directly measured. A barcode-normalized reporter effect says that a sequence changed reporter RNA, protein, fluorescence, or sorting behavior in a specific construct and cell context. It does not automatically show that the same sequence controls the endogenous transcript.

A useful evidence ladder is: first, replicate the reporter effect with multiple barcodes and independent library preparations. Second, measure the correct molecular output, such as RNA abundance for stability claims, isoform ratios for splicing claims, or protein-per-RNA output for translation claims. Third, test focused variants that distinguish motif loss, structure disruption, and unrelated sequence changes. Fourth, check whether the endogenous locus expresses the relevant isoform in the relevant cell state. Fifth, use endogenous editing, allele-specific expression, RNA half-life assays, or native-context reporters when the claim concerns disease variants or natural regulation.

Reporter-library evidence is strongest when the tested construct matches the biological question. If the question is whether a synthetic mRNA 5′ untranslated region improves translation in a vaccine-like mRNA, the reporter should use relevant cap, untranslated region architecture, coding sequence, poly(A) context, RNA chemistry, delivery, and cell type as far as feasible. If the question is whether a disease-associated noncoding variant alters endogenous gene regulation, the MPRA can prioritize variants but should be followed by endogenous editing, allele-specific expression, chromatin or RNA measurement, and disease-relevant cell models. If the question is whether an RNA motif controls stability, mutational rescue of the motif and direct RNA half-life measurement strengthen the claim.

MPRA-style experiments increasingly feed computational sequence-function models. Models can learn motif grammar, RNA structure features, codon features, splice regulatory codes, enhancer logic, or untranslated-region effects from large libraries. These models are only as general as their training data. A model trained on synthetic sequences may learn synthesis constraints; a model trained in one cell line may miss cell-type-specific RNA-binding proteins; a model trained on plasmid reporters may overpredict native effects. The right interpretation is not “the model discovered the regulatory code” but “the model learned predictive features for this assay and should be tested in independent contexts.”

Table 136.3. Reporter-Library Designs for RNA Regulatory Questions. Reporter libraries can test untranslated regions, splice elements, RNA stability, translation, localization, or binding, but episomal context, copy number, and synthetic sequence design limit transfer to endogenous loci.

RNA question Library design Primary output Main caveat Follow-up
Does a 5′ UTR tune translation? Natural or synthetic 5′ UTR variants upstream of reporter RNA and protein or fluorescence Reporter initiation context may differ from final mRNA Test final cap, UTR, coding sequence, poly(A), and cell type
Does a 3′ UTR alter stability? 3′ UTR fragments or mutational series downstream of reporter RNA abundance or half-life proxy Missing full-length transcript and localization context Direct RNA half-life and endogenous editing
Which motifs regulate splicing? Minigene or exon/intron variant library Isoform ratios Minigene context can alter splice-site competition Endogenous splice assay or targeted editing
Do codon variants affect RNA and protein output? Synonymous coding variant library RNA abundance and protein output Protein sequence constant but RNA structure and translation context differ Separate RNA stability from translation measurement
Which noncoding variants affect expression? Variant pairs or saturation mutagenesis in reporter RNA/DNA ratio or fluorescence Reporter activity is not native-locus function Allele-specific expression and endogenous variant editing
Which RNA structures matter? Structure-disrupting and compensatory variants Reporter RNA or protein output Sequence changes can alter motifs as well as structure Compensatory rescue and orthogonal structure probing

136.6. Perturb-seq, CROP-seq, and single-cell pooled screens

Perturb-seq and related methods combine pooled perturbation with single-cell RNA-seq. The screen introduces a perturbation library into cells and then uses single-cell sequencing to recover both the transcriptome and the perturbation identity, either from guide capture, expressed guide sequences, barcodes, or compatible construct designs. CROP-seq is one guide-capture architecture in this broader family. CRISP-seq, Mosaic-seq, direct-capture Perturb-seq, multimodal Perturb-seq, and related workflows differ in vector design, guide recovery, assay chemistry, and analysis conventions, but they share the goal of linking perturbations to high-dimensional cell states.

The conceptual advance is that the phenotype is not a single scalar. A bulk screen might say that a guide reduces survival after viral infection. A Perturb-seq experiment can ask whether the same guide lowers interferon-stimulated gene expression, activates unfolded protein response, changes splicing-factor expression, shifts cells into a proliferative state, or induces a macrophage-like antiviral program. This is especially important for RNA biology because RNA-processing factors often have pleiotropic, state-dependent effects that are invisible in simple growth measurements.

Single-cell pooled screens have a nested statistical structure. Cells are measured individually, but cells from the same infection, capture lane, donor, batch, or guide are not independent biological replicates. A common analysis groups cells by guide and sample, creates pseudobulk profiles, or uses models that include sample-level and donor-level structure. Cell-level covariates such as total unique molecular identifiers, mitochondrial RNA fraction, cell-cycle state, stress signatures, guide count, and doublet score may be relevant. Treating thousands of cells from one sample as thousands of independent replicates inflates significance and hides batch dependence.

Box 136.3. What Counts as a Replicate in Perturb-seq?

Single-cell screens create many observations, but not every observation is an independent experiment. Cells from the same infection, selection, capture lane, donor, or culture batch share history. Treating those cells as independent biological replicates can make small batch effects look like precise perturbation effects.

For most Perturb-seq questions, the useful hierarchy is molecules within cells, cells within guide assignments, guides within target genes or elements, and cells or pseudobulk profiles within samples, donors, and batches. A robust design asks whether a perturbation signature appears across independent guides and across independent biological samples, not only across many cells from one sample. Cell-level covariates such as library size, cell cycle, stress, mitochondrial RNA fraction, and doublet score are still important, but they do not replace replication at the sample or donor level.

Guide assignment is a core measurement, not a preprocessing detail. A cell may have no guide detected because guide capture failed. A cell may have multiple guide signals because it received multiple constructs, because of multiplet capture, or because of ambient guide RNA. A high multiplicity design may be intentional for combinatorial perturbations, but single-perturbation interpretation requires low multiplicity and filtering. Guide detection thresholds should balance false assignment against loss of cells. If guide identity is uncertain, downstream expression signatures can be diluted or assigned to the wrong perturbation.

Target engagement in single-cell screens requires explicit measurement. In a CRISPRi Perturb-seq experiment, the target transcript should be reduced in guide-positive cells, unless the target is not captured or the phenotype is indirect. In a CRISPRa experiment, induction should be measured. In a knockout experiment, RNA loss may be unreliable because nonsense-mediated decay, allele state, protein turnover, and transcript capture all intervene. In a Cas13 single-cell screen, target RNA depletion and transcriptome-wide side effects should be assessed. For noncoding RNAs and lowly expressed targets, absence of captured RNA is not proof of absence or successful perturbation.

Single-cell readouts can reveal pathways, but pathway language needs care. Perturbations that cluster by transcriptome signature may act in the same pathway, converge on the same stress response, or merely share cell-cycle or viability effects. A perturbation that enriches a cell state may appear to regulate many genes because it changes cell composition, not because it directly regulates those genes within a stable state. Trajectory and state models can help, but they are model-based summaries. Strong pathway claims require targeted validation, independent perturbations, and direct measurements tied to the proposed mechanism.

Perturb-seq designs are expanding beyond transcriptomes. Multimodal single-cell screens can combine guide identity with chromatin accessibility, surface proteins, protein epitopes, antigen specificity, lineage barcodes, spatial location, or targeted RNA panels. These designs are powerful for immune differentiation, neuronal maturation, cancer drug response, and host-pathogen biology. They also multiply the number of tests and artifacts. A multimodal screen can fail because one modality has low capture, because cells are lost during processing, or because integration methods align unrelated variation. The central discipline remains the same: define the biological question, preserve representation, measure target engagement, model the correct unit, and validate key conclusions.

Figure 136.5. Perturb-seq and CROP-seq Architecture

Figure 136.5. Perturb-seq and CROP-seq Architecture. Single-cell pooled screens assign perturbations to cells and recover guide identity together with transcriptome states. The figure shows perturbation delivery, single-cell capture, guide assignment, transcriptome measurement, quality control, pseudobulk or hierarchical analysis, and validation. Ambiguities include missing guides, multiple guides, doublets, sparse RNA capture, ambient RNA, and donor or batch effects.

136.7. Statistical analysis, hit validation, and reproducibility

Pooled-screen statistics begin with the experimental unit. In a guide depletion screen, the raw data are counts for guides or hairpins before and after selection. In a reporter library, the raw data are DNA and RNA barcode counts, reporter fluorescence bins, or direct sequence counts. In a Perturb-seq experiment, the raw data are cells, molecules, guide assignments, genes, samples, and donors. The correct model depends on which unit supports the claim. A gene-level dependency claim aggregates guide-level evidence. A sequence-function claim aggregates barcode-level evidence to a designed sequence or feature. A single-cell pathway claim must account for cells nested within sample and perturbation structure.

Count data are overdispersed. Sequencing counts vary because of sampling, PCR, growth, library construction, bottlenecks, and biological differences. Negative-binomial models, beta-binomial models, robust rank aggregation, permutation tests, mixed models, and empirical-null approaches all appear in screening analysis because simple fold changes and unmodeled Poisson assumptions are often inadequate. The choice of method should follow the data-generating process. For example, a survival screen with replicated guide counts needs a count model and gene-level aggregation; a reporter assay with multiple barcodes per element needs input normalization and barcode-level random or fixed effects; a Perturb-seq experiment may need pseudobulk contrasts or hierarchical models rather than naive cell-level tests.

Multiple testing is not optional. Genome-scale screens test thousands of genes. MPRA libraries test thousands to millions of variants, motifs, or model features. Single-cell screens can test perturbation-by-gene, perturbation-by-state, and perturbation-by-condition effects. False discovery rate control, prespecified contrasts, independent validation sets, and effect-size thresholds are part of the evidence standard. A tiny adjusted p value can still correspond to a trivial effect, a batch artifact, or an off-target guide. Conversely, a biologically important effect can be missed if the library has weak representation, guides fail, or the assay readout is poorly matched to the mechanism.

Library representation is a reproducibility requirement. Each library member should be present at sufficient abundance during cloning, viral production, infection, selection, sorting, and sequencing. Representation lost early cannot be restored by deeper sequencing later. A bottleneck during fluorescence sorting can create apparent guide enrichment. A low-complexity viral pool can eliminate rare guides before selection. Uneven barcode abundance can dominate reporter-library variance. Representation should be measured at input and relevant intermediate stages, and experiments should use cell numbers appropriate for the intended coverage.

Table 136.4. Statistical Units, Artifacts, and Validation Standards Across Pooled Screens. Guide, cell, donor, construct, and target can each be the relevant statistical unit in pooled screens; models and validation must respect nesting, multiple testing, off-targets, and perturbation heterogeneity.

Biological question Primary statistical unit Model or analysis family Main false-positive mode Minimum validation expectation
Which genes are required for growth? Guide or shRNA counts aggregated to gene Count model, rank aggregation, depletion statistics Bottleneck, seed toxicity, copy-number break toxicity Replicate concordance and independent reagents
Which genes modify drug or infection response? Guide counts by condition Interaction or contrast model Baseline growth effect mistaken for modifier Untreated comparison and dose-response validation
Which RNA variants alter reporter output? Barcode counts aggregated to sequence DNA-normalized RNA/DNA ratio, regression, mixed model Barcode imbalance and cloning bias Multiple barcodes and focused reporter assay
Which perturbations induce a transcriptome program? Cells grouped by guide, sample, and donor Pseudobulk, differential expression, latent-variable models Treating cells as independent replicates Donor-aware analysis and target engagement
Which perturbations share a pathway? Perturbation-level signatures Clustering, network analysis, gene-set scoring Stress or cell-cycle effects mistaken for pathway Orthogonal pathway readout
Does a noncoding RNA act as an RNA molecule? RNA-targeting and locus-targeting perturbations compared Cross-modality comparison DNA regulatory effect mistaken for RNA-product function RNA depletion, locus perturbation, and RNA rescue
Is a hit therapeutically relevant? Target and phenotype across models Stratified dependency and validation panels Cell-line-specific artifact or non-druggable dependency Multiple models, selectivity, toxicity, and target engagement

Controls must be modality-specific. RNAi screens need non-targeting controls, positive knockdown controls, seed-aware controls, and concentration controls. CRISPR knockout screens need non-targeting or safe-targeting controls, positive essential-gene controls, copy-number-aware interpretation, and edit or protein validation. CRISPRi and CRISPRa screens need expression controls and guide-position controls. Cas13 screens need target RNA measurement, effector-only controls, non-targeting guides, and independent guides against accessible regions. Reporter libraries need DNA input sequencing, multiple barcodes, empty or neutral context controls, positive regulatory elements, and synthesis or cloning quality checks. Perturb-seq needs guide-capture controls, doublet filtering, sample replication, donor-aware analysis, and target-engagement checks.

Hit validation should climb a ladder whose height matches the claim. The lowest rung is a statistical hit in the primary screen. The next rungs include replicate concordance, independent reagents, direct target engagement, dose or timing consistency, phenotype reproduction in an arrayed or focused assay, rescue or complementation, orthogonal perturbation, endogenous-locus confirmation, and mechanistic dissection. A prioritization paper may reasonably stop at a lower rung for many hits while validating a subset deeply. A mechanistic claim, clinical claim, or design rule needs higher rungs.

Figure 136.6. Validation Ladder from Statistical Hit to Mechanistic Claim

Figure 136.6. Validation Ladder from Statistical Hit to Mechanistic Claim. Screen hits become causal claims through staged validation. The lowest rung is statistical enrichment, depletion, reporter change, or transcriptome signature. Higher rungs add replicate concordance, independent reagents, direct target engagement, dose or timing consistency, focused assay reproduction, rescue or complementation, orthogonal perturbation, endogenous-locus confirmation, and mechanistic dissection.

Rescue logic is particularly important in RNA biology. For RNAi, an RNAi-resistant cDNA can test whether the phenotype follows the intended coding sequence, but it may not rescue untranslated-region regulation, isoform choice, or expression dosage. For Cas13, a guide-resistant RNA rescue can support RNA-product specificity, but overexpression may change localization and RNP assembly. For CRISPRi repression of a noncoding RNA locus, rescue by an exogenous RNA can help separate RNA function from DNA regulatory effects, but only if the rescue RNA recapitulates processing, localization, and interaction partners. For reporter libraries, a focused reporter assay can validate a sequence effect, while endogenous editing tests native context.

Reproducibility also includes transparent reporting. A screen report should state library design, guide or barcode counts, delivery method, multiplicity, cell numbers, selection duration, sequencing depth, replicate structure, filtering thresholds, statistical model, multiple-testing method, control performance, target-engagement assays, validation criteria, and data availability. For RNA-focused screens, the report should also state transcript annotations, isoform assumptions, RNA measurement methods, and any innate immune activation, stress response, or RNA-processing artifact relevant to the perturbation.

The major misconception is that scale substitutes for causality. It does not. Large libraries can create large false-positive sets when design, representation, and validation are weak. A screen is most powerful when it is used as a structured hypothesis generator and when the follow-up experiment directly tests the proposed molecular mechanism.

Reporting negative and ambiguous results also improves reproducibility. A guide class that fails because the target RNA is not expressed, a reporter library that loses representation during cloning, or a Perturb-seq experiment with poor guide capture can teach later investigators how the biological system and assay interact. These results should be recorded with enough detail to distinguish a true absence of phenotype from inadequate target engagement, low power, or assay mismatch. In RNA-focused screens, this distinction is especially important because transcript isoforms, RNA localization, RNP assembly, innate immune activation, and RNA stability can all change the apparent activity of the same perturbation across contexts.

Biological Contexts, Boundary Cases, and System Choice

The screening system is part of the biological claim. Cancer cell lines are convenient, scalable, and often easy to transduce, but they can carry copy-number amplifications, DNA repair defects, abnormal stress pathways, and altered RNA regulation. Primary cells and organoids can be more relevant to immunity, development, or disease, but they may have limited cell numbers, donor variation, differentiation heterogeneity, and delivery challenges. Viral infection screens depend on viral strain, multiplicity, timing, innate immune competence, and biosafety constraints. Plant, fungal, animal, and microbial systems differ in RNAi machinery, CRISPR delivery, chromatin context, and genetic redundancy.

Perturbation modality should be matched to target biology. A protein-coding gene with a clear loss-of-function question may be suited to CRISPR knockout. A dosage-sensitive essential gene may be better tested by CRISPRi or RNAi. A gain-of-expression question may need CRISPRa or a cDNA library. A viral RNA or noncoding RNA product may need Cas13, RNAi, antisense oligonucleotides, or comparative perturbation across RNA-targeting and locus-targeting modalities. A regulatory sequence grammar question may need MPRA, saturation editing, or endogenous reporter insertion. A cell-state question may need Perturb-seq rather than survival selection.

Boundary cases are common in RNA biology. A noncoding RNA locus can encode an RNA molecule, act as a DNA regulatory element, influence nearby genes by transcription itself, and overlap enhancer or promoter activity. A perturbation that deletes the locus, represses transcription, degrades the RNA, or blocks an RNA motif tests different hypotheses. Similarly, a 3′ untranslated-region variant can affect RNA stability, localization, translation, polyadenylation, microRNA targeting, or none of these in a reporter context. A screen design should make these alternatives visible rather than forcing a single simplified interpretation.

Recent Consensus

The current methodological consensus is pluralistic. RNAi, CRISPR knockout, CRISPRi, CRISPRa, Cas13, reporter libraries, and Perturb-seq answer overlapping but nonidentical questions. CRISPR knockout is often strong for protein-coding loss-of-function genetics, but it is not automatically best for essential genes, amplified loci, noncoding RNAs, dosage-sensitive processes, or RNA-product questions. CRISPRi and CRISPRa are powerful for transcriptional modulation and regulatory interrogation, but they require expression validation and careful guide placement. RNAi remains valuable for partial and transcript-level depletion, but seed-mediated artifacts require special caution. Cas13 is promising for direct RNA targeting, but guide rules and effector behavior must be validated in context. MPRA-style assays are central for sequence-function mapping, but reporter context limits endogenous claims. Perturb-seq has transformed high-dimensional perturbation biology, but it demands careful modeling of guide, cell, donor, sample, and state structure.

Open Questions, Controversies, Deprecated Models, and Common Misconceptions

Open questions:

  • How should Cas13 guide-design rules, collateral-effect boundaries, and transcript-accessibility models be refined?
  • Which single-cell screen analysis models are appropriate for different combinations of perturbations, donors, batches, and multimodal readouts?
  • How can reporter-library models transfer more reliably from synthetic or plasmid contexts to endogenous loci?
  • How can RNAi seed-toxicity prediction become less context-dependent?
  • How can CRISPRi and CRISPRa interpretation at complex noncoding loci separate transcription, RNA-product, enhancer, promoter, and chromatin effects?

Common misconceptions:

  • “A single siRNA phenotype proves target-gene function.” Single-reagent phenotypes can reflect seed effects, toxicity, delivery, or off-target knockdown; target function needs multiple reagents and rescue or orthogonal perturbation.
  • “Multiple RNAi reagents prove specificity when they share a seed or toxicity profile.” Shared off-target mechanisms can make several reagents agree for the wrong reason.
  • “A CRISPR knockout hit at a noncoding locus proves the RNA molecule is functional.” DNA-element disruption, transcriptional interference, chromatin effects, or neighboring genes can drive the phenotype.
  • “A guide depleted from an amplified region is automatically an essential-gene guide.” Copy-number effects and DNA-damage toxicity can deplete guides independently of gene essentiality.
  • “A reporter variant is automatically an endogenous regulatory variant.” Reporters omit chromatin, isoform, RNA-structure, cell-state, and dosage context.
  • “A Perturb-seq cluster is automatically a pathway.” Similar transcriptomic responses can reflect indirect stress, shared growth effects, or cell-state shifts rather than one molecular pathway.
  • “A screen hit is the final causal explanation.” Screen hits are starting points that require validation, mechanism, and context-specific interpretation.