# Chapter 130. Single-Cell, Spatial, Multimodal, Perturbational, and Lineage RNA Technologies and Analysis

## Scope Note

This chapter owns the technologies, computational analysis, benchmarking, and failure modes of single-cell and single-nucleus RNA sequencing, transcriptome-scale spatial platforms, multiplexed spatial imaging and in situ sequencing, multimodal assays, perturbational screens, and lineage-coupled methods. It explains how assays generate evidence, which molecules and compartments are sampled, and how capture efficiency, ambient RNA, normalization, spatial resolution, segmentation, batch structure, annotation, integration, guide recovery, and lineage-barcode behavior constrain analysis. Foundational targeted RNA detection and imaging measurement physics—including blots, nuclease protection, primer extension and RACE, reverse-transcription quantitative PCR (RT-qPCR), digital PCR, single-molecule fluorescence in situ hybridization (smFISH), branched-DNA assays, live-cell tracking, calibration, and detection limits—belong to [Chapter 123](chapter1155.md). Biological discoveries about cell-state programs, tissue niches, developmental transitions, lineage decisions, perturbational mechanisms, and cross-study causal synthesis belong to [Chapter 106](chapter1101.md); biological examples here serve only to explain method outputs and failure modes.

## Executive Summary

Single-cell and single-nucleus RNA technologies convert a heterogeneous biological sample into many barcode-associated RNA count profiles. A cell barcode identifies a partition such as a droplet, well, microwell, or indexed molecular reaction. A unique molecular identifier, or UMI, marks a captured RNA-derived molecule before amplification so that PCR duplicates can be collapsed. These features make high-throughput cell atlases and perturbational screens possible, but they do not make the assay a complete census of every RNA molecule in a cell. The observed matrix contains molecules that survived sampling, were accessible to the assay chemistry, were reverse transcribed, were amplified or sequenced, and could be assigned to a gene or feature.

The distinction between scRNA-seq and snRNA-seq is biological as well as technical. scRNA-seq begins with intact cells and tends to sample cytoplasmic and nuclear polyadenylated RNA according to the protocol. snRNA-seq begins with isolated nuclei and is often better for frozen, archived, fragile, large, or hard-to-dissociate tissues, especially adult brain and other complex organs. Nuclei-based profiles are enriched for nuclear, intronic, nascent, and retained RNAs and can underrepresent cytoplasmic, mitochondrial, localized, and rapidly exported transcripts. A nucleus profile should not be treated as a low-depth version of a whole-cell profile.

Spatial transcriptomics and imaging-based RNA technologies add location. Spatial barcode capture methods preserve tissue coordinates by capturing RNA on slide positions, beads, pixels, or regions of interest and then sequencing the captured material. Multiplexed imaging approaches such as MERFISH, seqFISH, and in situ sequencing localize selected RNA molecules or decoded amplicons inside fixed cells and tissues and can be integrated with cell-atlas-scale analysis. This chapter compares their molecular breadth, spatial precision, tissue compatibility, scaling, segmentation dependence, and platform integration; [Chapter 123](chapter1155.md) supplies the foundational smFISH and branched-DNA measurement model, calibration logic, detection limits, and assay-specific optical controls. No platform simultaneously provides unbiased transcriptome-wide detection, single-molecule subcellular localization, perfect morphology, unlimited throughput, and simple quantification.

Multimodal and perturbational assays make RNA profiles more interpretable by adding selected protein markers, chromatin accessibility, guide identities, sample barcodes, genotypes, lineage labels, or spatial context. CITE-seq measures RNA together with antibody-derived tags for selected proteins. Multiome assays can connect RNA state with chromatin accessibility in the same nucleus. Perturb-seq links CRISPR or other perturbations to transcriptomic phenotypes at single-cell resolution. Lineage-coupled assays recover ancestry or clonal identity alongside cell state. These designs are powerful precisely because the same cell barcode can link several evidence layers, but each extra modality adds its own specificity, sensitivity, background, and sampling problems.

Computational analysis is part of the measurement. Normalization, feature selection, dimensionality reduction, clustering, batch correction, reference mapping, deconvolution, trajectory inference, differential expression, and perturbation modeling all impose assumptions. Cell-type labels are evidence-weighted hypotheses, not algorithmic facts. Batch correction can align shared states, but it can erase real disease, donor, treatment, region, or time-point differences when those variables are confounded with technical batches. The strongest studies define the sampled RNA population, include biological replicates, preserve metadata, validate key labels and spatial claims with orthogonal evidence, and avoid treating clusters, proximity patterns, or guide-associated expression changes as causal conclusions without controls.

Citation status for this draft: the chapter-local `references.md` contains a coverage-repair bibliography with verified metadata for review, method, application, and benchmark records. Claim-level linkage remains incomplete, and several subsection blocks still contain weak application-focused sources; final curation should prioritize landmark method papers and direct benchmarks.

## Concept Inventory

- **Single-cell RNA sequencing:** a family of assays that assign RNA-derived reads or UMI-collapsed molecule counts to individual cells through physical partitioning, molecular barcoding, or both. Most high-throughput scRNA-seq measures sparse gene-level end counts from polyadenylated RNA rather than full transcriptomes.
- **Single-nucleus RNA sequencing:** a single-cell-style RNA profiling approach in which isolated nuclei rather than intact cells are captured, barcoded, and sequenced. snRNA-seq is useful for frozen and fragile tissues but shifts the measured RNA compartment toward nuclear and pre-mRNA signal.
- **Cell barcode:** a molecular index that links reads, UMIs, antibody tags, guide tags, or other features to a partition. A cell barcode identifies a capture event, not necessarily one viable intact cell.
- **Unique molecular identifier:** a short random or semi-random barcode attached to an RNA-derived molecule before amplification. UMIs help collapse PCR duplicates among captured molecules but do not correct molecules that were never captured.
- **Capture efficiency:** the fraction of input molecules, or of a defined RNA class, represented in the final dataset. Capture efficiency varies by chemistry, RNA class, cell type, sample quality, sequencing depth, and computational assignment.
- **Ambient RNA:** RNA in the suspension, droplet, tissue section, or capture field that does not originate from the intended cell, nucleus, or location. Ambient RNA can create false marker expression.
- **Doublet or multiplet:** a barcode-associated profile produced by two or more captured cells or nuclei. Doublets can mimic rare cell states, transition states, or hybrid phenotypes.
- **Spatial transcriptomics:** RNA measurement that retains tissue, cellular, or subcellular coordinates. The term includes spatial barcode capture, region-of-interest sequencing, imaging-based RNA detection, and in situ sequencing.
- **Imaging-based spatial transcriptomics:** targeted, multiplexed optical detection or in situ decoding used to assign RNA molecules or RNA-derived products to cells, subcellular compartments, and tissue coordinates at atlas or platform scale. Foundational single-target and molecule-counting physics are introduced in [Chapter 123](chapter1155.md).
- **Multimodal single-cell assay:** an assay that measures RNA together with another molecular or contextual layer in the same cell or nucleus, such as surface protein, chromatin accessibility, perturbation identity, genotype, lineage, or spatial location.
- **Perturb-seq:** a pooled perturbation design in which perturbation identities, often guide RNAs or linked barcodes, are recovered together with single-cell RNA profiles.
- **Cell-state annotation:** assignment of biological meaning to profiles using marker genes, references, multimodal features, morphology, spatial location, perturbation response, or expert knowledge. Annotation is a claim with evidence, not an intrinsic property of a cluster.
- **Integration and batch correction:** computational alignment or joint modeling of datasets across samples, donors, batches, technologies, modalities, or references. Integration can improve comparability but can also remove real biological differences.

## What to Know Before Reading This Chapter

The reader should understand the bulk RNA-seq workflow from [Chapter 125](chapter1116.md): RNA is captured or extracted, converted to cDNA, prepared as a library, sequenced, aligned or pseudoaligned, assigned to genes or transcripts, normalized, and statistically interpreted. Single-cell methods add a critical early barcoding step. If the barcode is attached before most amplification, reads can later be grouped by cell barcode, UMI, and gene. The familiar gene-by-sample count matrix becomes a gene-by-cell or gene-by-nucleus matrix, but the matrix is still a processed representation rather than a direct image of all RNA molecules.

A zero count in a single-cell matrix is ambiguous. The gene may not have been transcribed in the captured cell; the transcript may have been present but not captured; the molecule may have failed reverse transcription; the read may have been filtered; or the gene model may not have matched the read. This ambiguity is especially important for low-abundance transcription factors, cytokines, receptors, viral RNAs, long noncoding RNAs, splice isoforms, allele-specific RNAs, and transient stress-response transcripts. The chapter therefore treats non-detection as an assay observation, not as proof of molecular absence.

Four running examples make the methods concrete. A dissociated immune-cell sample illustrates scRNA-seq, CITE-seq, doublets, and cell-state annotation because immune cells are often defined by both transcript and surface-protein programs. A frozen adult brain sample illustrates snRNA-seq because intact neurons and glia are difficult to recover from archived tissue, whereas nuclei can often be isolated. A tumor biopsy illustrates spatial transcriptomics because malignant cells, stromal cells, immune cells, necrotic regions, vasculature, and invasive margins are arranged in tissue space. A CRISPR interference Perturb-seq experiment illustrates perturbational interpretation because guide identity, target repression, viability, cell cycle, and downstream transcription must be separated.

## 130.1. scRNA-seq and snRNA-seq chemistries

Single-cell RNA sequencing begins by linking RNA-derived molecules to a partition that is meant to represent one cell. In droplet microfluidic methods, a cell or nucleus suspension is diluted so that most droplets contain no biological particle and a minority contain one cell or nucleus. Barcoded beads or gel particles provide oligonucleotides containing a cell barcode, a UMI, and a capture sequence, commonly oligo(dT) for polyadenylated RNA. After lysis inside the droplet, RNAs hybridize to bead oligonucleotides, reverse transcriptase makes cDNA, and all captured molecules from that droplet inherit the same cell barcode. The sample can then be pooled, amplified, sequenced, and collapsed into UMI counts.

The droplet workflow is powerful because it moves cell identity into DNA sequence early. Once barcoded cDNA is made, thousands or millions of partitions can be processed together. The tradeoff is sparse sampling. High-throughput droplet assays usually recover enough information to classify cells and estimate gene-level expression patterns, but not enough to reconstruct most transcript isoforms or detect every biologically important low-abundance RNA. Droplet methods are therefore well suited to atlas construction, cell-state discovery, perturbational screens, and broad comparative studies, but they are weaker for complete transcript architecture, rare splice junctions, low-copy viral RNAs, and absolute RNA quantification.

Plate-based scRNA-seq starts from single cells sorted or placed into wells. Fluorescence-activated cell sorting, index sorting, microfluidic capture, manual picking, or imaging-guided selection can provide richer metadata about each cell before sequencing. Some plate protocols generate longer or more full-length cDNA than droplet end-counting assays, making them useful for small samples, rare sorted phenotypes, allele-specific questions, immune-receptor reconstruction, and transcript-body information. The cost is lower throughput and higher per-cell expense. A plate profile can be deeper than a droplet profile, but it is still shaped by lysis, reverse transcription, amplification, and sampling bias.

Microwell and split-pool approaches occupy additional design space. Microwell systems physically trap cells and barcoded beads in small wells, often allowing imaging or simpler instrumentation than droplet microfluidics. Split-pool indexing methods avoid isolating each cell in a droplet by repeatedly distributing fixed cells or nuclei into wells and adding combinatorial barcodes. These approaches can scale to very large numbers of cells or nuclei, but barcode collisions, fixation effects, and combinatorial indexing design become central quality issues. The mechanistic lesson is that a single-cell assay is not one method; it is a family of partitioning and indexing strategies.

The difference between 3′ end counting, 5′ end counting, and full-length capture is not cosmetic. A 3′ end-counting assay counts molecules near transcript poly(A) ends and is efficient for gene-level abundance in many eukaryotic samples. A 5′ assay can support immune receptor enrichment and some transcription start-proximal information, but it still usually provides sparse end-count data rather than complete transcript models. Full-length protocols recover more transcript-body sequence and can support isoform, allele, or mutation questions in limited cell numbers. None of these designs removes the need to understand transcript annotation, read assignment, and amplification bias.

snRNA-seq uses isolated nuclei rather than intact cells as the captured biological unit. The practical advantages are substantial. Frozen tissue, archived clinical specimens, adult brain, skeletal muscle, adipose tissue, heart, kidney, and other difficult tissues may yield usable nuclei when intact-cell recovery is poor or biased. Nuclei isolation can reduce dissociation-induced transcriptional responses because nuclei can sometimes be extracted under colder and faster conditions than whole-cell dissociation. It also makes large, branched, fragile, or matrix-embedded cells more accessible.

The biological tradeoff is equally important. A nucleus contains chromatin, nascent RNA, pre-mRNA, retained introns, nuclear-retained long noncoding RNAs, and some mature RNA, but it lacks much of the cytoplasm. Cytoplasmic mRNAs, localized transcripts in neuronal processes, mitochondrial RNAs, ribosome-associated mRNAs, and rapidly exported mature transcripts can be depleted relative to whole-cell scRNA-seq. Many snRNA-seq pipelines count intronic reads to increase sensitivity and capture nascent or pre-mRNA signal, whereas some scRNA-seq analyses focus mainly on exonic reads. That choice changes the apparent abundance of genes and the interpretation of regulatory timing.

The appropriate chemistry depends on the question. A fresh blood immune atlas may favor droplet scRNA-seq plus cell hashing and CITE-seq. A frozen adult human brain atlas may favor snRNA-seq with intronic read counting. A rare sorted stem-cell state may justify plate-based deeper sequencing. An immune repertoire question may require 5′ capture and targeted V(D)J enrichment. A question about RNA localization in dendrites, epithelial polarity, or viral replication compartments may be poorly answered by dissociated scRNA-seq and should move toward spatial or imaging methods. Method choice defines which RNA molecules can become evidence.

> **Box 130.1. Choosing the RNA Technology for the Biological Question**
>
> Ask first what must remain visible. If the key evidence is broad cell-state composition in a fresh suspension, droplet scRNA-seq may be appropriate. If the material is frozen, fragile, archived, or difficult to dissociate, snRNA-seq may preserve representation better, but the evidence will be nuclear-shifted. If the claim depends on tissue architecture, cell neighborhoods, tumor margins, organ layers, or infection foci, a spatial capture method or imaging assay is needed. If the claim depends on subcellular localization or validation of a small marker panel, smFISH-style or multiplexed imaging is often stronger than dissociation. If the question requires immune receptor recovery, selected proteins, perturbation identity, chromatin accessibility, or lineage, choose a multimodal or perturbational design from the start. No downstream analysis can recover a compartment, location, molecule class, or perturbation label that the assay never measured.

Figure 130.1 should show that whole-cell scRNA-seq and snRNA-seq share barcode logic but sample different compartments.

![Figure 130.1. Whole-cell scRNA-seq and snRNA-seq capture different RNA compartments](../assets/figures/chapter1118_figure1.png)

**Figure 130.1. Whole-cell scRNA-seq and snRNA-seq capture different RNA compartments.** Teach that scRNA-seq and snRNA-seq share barcode logic but sample different biological material.

## 130.2. Capture efficiency, normalization, and ambient RNA

Capture efficiency is the fraction of RNA molecules that become usable observations. This fraction is usually far below one in high-throughput single-cell assays. Molecules can be lost because the cell lysed incompletely, the RNA lacked the capture feature required by the protocol, the transcript was degraded, the RNA was inaccessible in a ribonucleoprotein complex, reverse transcription failed, amplification was biased, sequencing depth was insufficient, or the read could not be assigned. UMI counting makes amplification less misleading by collapsing duplicates, but it cannot count molecules that never received a UMI.

Dropout is the common name for non-detection of a transcript in a cell where the transcript may have been present. The term is useful but dangerous if it implies that all zeros are technical. Some zeros represent true absence, especially for lineage-restricted genes or tightly regulated inducible transcripts. Other zeros reflect limited sensitivity. A careful analysis asks which genes are expected to be detectable at the chosen depth, how detection varies across cell types, and whether key conclusions depend on absence claims. For example, failure to detect a cytokine transcript in a T cell is weaker evidence than detection of that cytokine in replicated activated cells, because cytokine mRNAs can be transient and low abundance.

Normalization tries to make barcode profiles comparable despite different sequencing depths, molecule recovery, RNA content, and technical noise. Simple library-size normalization divides each cell by its total counts and often applies a logarithmic transform. More elaborate approaches estimate size factors, model variance, regress covariates, or use negative-binomial and latent-variable frameworks. The core assumption should be stated: is the analysis trying to compare relative composition of captured RNA, approximate per-cell abundance, or identify states after removing technical depth? The same normalization can be appropriate for one question and misleading for another.

Total RNA content is not always nuisance variation. Activated lymphocytes, cycling progenitors, large neurons, polyploid cells, dying cells, hypertrophic cells, and differentiated secretory cells can differ biologically in RNA abundance. If all cells are forced to the same library size, a global increase in RNA production or cell size can be converted into relative decreases for genes that did not change. Conversely, if raw counts are compared without depth adjustment, deeply sequenced or high-capture cells can appear biologically distinct. Spike-ins, sample design, flow cytometry, cell-size measures, or independent RNA quantification can help interpret global shifts, but many atlas-scale datasets lack absolute molecule calibration.

Ambient RNA is RNA outside the intended captured cell or nucleus. It may come from lysed cells, damaged tissue, dead cells, free nuclei, extracellular RNA, ruptured fragile cell types, or carryover during loading and washing. Ambient RNA is often dominated by highly expressed transcripts from abundant or fragile populations. In a liver sample, hepatocyte transcripts can contaminate immune-cell barcodes; in blood, hemoglobin transcripts can affect unrelated cells; in brain, neuronal or glial RNA can enter droplets that contain other nuclei. Ambient RNA can make a rare cell type appear to express markers from a different cell type, or can create low-level viral, mitochondrial, or stress signatures.

Empty droplets and low-quality partitions are part of the same problem. A droplet without a cell can still contain ambient RNA and generate reads. A droplet with a damaged cell can generate a low-complexity profile dominated by mitochondrial, ribosomal, or stress-associated molecules. A droplet with a nucleus may contain nuclear RNA plus ambient cytoplasmic RNA from the tissue suspension. Thresholds based on total counts, detected genes, mitochondrial fraction, ribosomal fraction, intronic fraction, and marker patterns help filter profiles, but thresholds must be tissue-specific. A mitochondrial fraction that indicates damage in one sample may be typical for metabolically active muscle or heart cells.

Doublets and multiplets occur when two or more cells or nuclei share one barcode. They can mimic rare populations by combining markers from two common cell types. A T cell plus monocyte doublet can look like a T cell with antigen-presenting features. A tumor cell plus macrophage doublet can look like a malignant cell with immune markers. During development, doublets can resemble transitional states. Doublet detection methods use expected loading rates, unusually high UMI counts, co-expression of incompatible markers, simulated mixed profiles, genotype differences, or sample multiplexing. These methods reduce but do not eliminate uncertainty.

Some apparent doublets are biological. Phagocytes can contain RNA from engulfed cells. Immune synapses and cell-cell conjugates can bring two cell programs into one capture event. Syncytial tissues and multinucleated cells challenge the idea that one barcode equals one cell. Nuclei from multinucleated muscle fibers and megakaryocytes require special interpretation. The quality-control question is therefore not simply whether a profile is pure, but what biological unit the assay captured and whether that unit matches the question.

Table 130.1 should summarize common artifacts, data symptoms, contexts, controls, and residual cautions.

**Table 130.1. Common single-cell RNA artifacts, symptoms, and controls.** Give readers a practical diagnostic checklist for interpreting cell-level RNA matrices.

| Artifact or confounder | Typical symptoms in data | Most affected contexts | Controls or mitigation | Residual caution |
| --- | --- | --- | --- | --- |
| **Low capture efficiency/dropout** | Sparse counts, many zeros, weak low-abundance genes, variable detection across cells | High-throughput droplet data, low-input samples, transcripts with poor capture or assignment | UMIs, adequate depth, sensitivity benchmarks, replicate-aware interpretation of non-detection | A zero can be true absence or missed detection; absence claims need stronger evidence |
| **Ambient RNA** | Low-level markers from unrelated cell types, free mitochondrial or viral signal, reads in weak barcodes | Fragile, lysed, necrotic, or dissociated tissues; droplet scRNA-seq and snRNA-seq | Empty-droplet modeling, ambient correction, marker plausibility checks, tissue-aware QC | Some extracellular, engulfed, or transferred RNA may be biological, not removable background |
| **Empty droplets** | Low UMI counts with abundant ambient transcripts but no coherent cell program | Droplet and nuclei loading with high background RNA | Barcode-rank inspection, empty-droplet calling, count and gene thresholds | Aggressive calling can discard small or low-RNA cells; permissive calling admits background |
| **Doublets/multiplets** | High UMI counts, mixed incompatible markers, hybrid or rare-looking states | High loading rates, sticky cells, nuclei suspensions, cell-cell conjugates | Loading-rate control, doublet scores, genotype or hash tags, incompatible-marker review | True conjugates, phagocytosis, syncytia, and multinucleated cells can resemble doublets |
| **Dead or dying cells** | Low complexity, mitochondrial or ribosomal dominance, stress and apoptosis signatures | Slow processing, harsh dissociation, necrotic tumors, clinical tissue delays | Viability enrichment, rapid cold handling, QC thresholds, stress-marker review | Removing damaged profiles can bias against fragile or disease-relevant populations |
| **Dissociation-induced stress** | Immediate-early, heat-shock, interferon, or apoptosis programs shared across cell types | Enzymatic digestion, sorting, hypoxia, warm delays, plant or matrix-rich tissues | Shorter protocols, cold-active dissociation, nuclei isolation, imaging validation | Nuclei reduce some stress signals but introduce nuclear-compartment bias |
| **Mitochondrial RNA thresholds** | High mitochondrial fraction interpreted as damage or low-quality capture | Metabolically active heart, muscle, kidney, neurons, damaged suspensions | Tissue-specific thresholds, joint review of genes, counts, introns, and morphology | A fixed cutoff can wrongly remove healthy high-mitochondrial cells or retain damaged cells |
| **Batch effects** | Separation by donor, lane, enzyme, buffer, slide, antibody lot, or guide library | Multi-sample atlases, disease-control comparisons, cross-platform integration | Balanced design, randomization, metadata, biological replicates, cautious integration | Confounded biology and batch cannot be fully separated after collection |
| **Barcode swapping or cross-contamination** | Sample labels conflict with genotype, hash tag, or expected marker pattern | Multiplexed libraries, high-throughput sequencing, pooled donors or perturbations | Demultiplexing, genotype checks, hash-tag thresholds, lane controls | Low-level contamination can survive filtering and mimic weak expression |
| **Overaggressive filtering** | Loss of rare, small, low-RNA, stressed, or lineage-specific cells | Rare populations, frozen tissue, nuclei, developmental or clinical samples | Sensitivity analysis across QC cutoffs, marker recovery checks, replicate review | Cleaner matrices may hide biologically important but technically difficult states |

## 130.3. Spatial transcriptomics platforms

Spatial transcriptomics measures RNA while preserving location. The word "spatial" covers several distinct measurement designs, so the first interpretive step is to define the spatial unit. In some assays the unit is a barcoded spot on a slide that may contain several cells. In high-definition capture assays the unit may be a bead or pixel smaller than a typical cell, but RNA diffusion, tissue thickness, and segmentation still matter. In region-of-interest sequencing, the unit may be a histologically selected area. In imaging assays, the unit may be a fluorescent spot assigned to a cell, subcellular compartment, or tissue coordinate. These units are not interchangeable.

Slide-based spatial barcode capture places tissue sections on surfaces containing oligonucleotides with known positional barcodes. RNA released from the tissue hybridizes to the surface, is reverse transcribed or otherwise captured, and is sequenced. Each read can then be assigned to both a gene and a location. This design can measure many genes across a tissue section while retaining histological context. The limitations are capture efficiency, spatial resolution, tissue permeabilization, RNA diffusion, spot mixing, and dependence on the quality of the section. In early or moderate-resolution formats, one spot can contain multiple cells, so deconvolution is often required.

High-definition bead or pixel-based spatial capture reduces the physical size of capture features. Smaller features can improve apparent resolution, but resolution is not only feature diameter. A thin tissue section can place partial cells over several features. RNA can diffuse during permeabilization. Segmentation may assign molecules to the wrong cell or fail in dense tissue. Sequencing depth per feature may fall as features become smaller. The practical question is whether the platform resolves the biological structure of interest: epithelial layers, tumor margins, immune niches, glomeruli, neuronal layers, viral foci, or subcellular compartments.

Region-of-interest and laser-capture approaches select tissue areas before sequencing. These methods can use pathology, immunostaining, morphology, or manual annotation to choose regions. They are valuable when the question concerns defined compartments such as tumor nests, stroma, inflamed regions, lesions, glomeruli, crypts, or brain layers. They can be compatible with formalin-fixed paraffin-embedded tissue in some workflows. The tradeoff is that the region is an analyst- or instrument-defined unit rather than an unbiased single-cell measurement. A region can contain many cell types, and its RNA profile is a mixture.

Spatial transcriptomics is often paired with histology. The image is not decorative; it is part of the evidence. Tissue preservation, section thickness, staining, registration, and pathological annotation determine whether RNA signals can be interpreted in anatomical context. A tumor spatial map without histology may identify gene-rich regions but not necrosis, invasion, vasculature, or immune exclusion. A brain spatial map without layer or region annotation may miss the anatomical meaning of cell states. Conversely, histology without molecular validation can overinterpret morphology.

Spatial deconvolution uses dissociated scRNA-seq or snRNA-seq references to infer which cell types contribute to a spatial spot or region. This is useful when spatial capture positions contain mixtures. The inference is constrained by the reference. If the reference lacks a cell type, activation state, disease state, or tissue-resident program, the deconvolution model cannot recover it reliably. Dissociation can also alter gene expression, making a reference differ from the intact tissue. Deconvolution should therefore be treated as model-based evidence that benefits from validation by imaging, protein staining, morphology, or targeted RNA assays.

Figure 130.2 should compare spatial platform families by molecular breadth and spatial precision, while labeling resolution as a composite property rather than a single number.

![Figure 130.2. Spatial transcriptomics platform tradeoff map](../assets/figures/chapter1118_figure2.png)

**Figure 130.2. Spatial transcriptomics platform tradeoff map.** Show why platform choice depends on biological question rather than on a single ranking.

Spatial transcriptomics is strongest when location is central to the biological claim. It can test whether immune cells are excluded from malignant nests, whether interferon-stimulated cells cluster near infected foci, whether developmental gene programs align with tissue axes, whether fibroblast states occupy distinct niches, or whether neuronal and glial states map to anatomical layers. It is weaker when used only as a more expensive way to make a gene-expression matrix. Spatial coordinates are valuable because tissues are organized systems, not because every coordinate automatically implies interaction or causality.

## 130.4. Imaging-based RNA methods and in situ sequencing

Imaging-based spatial transcriptomics detects RNA molecules or RNA-derived signals directly in fixed cells or tissues while preserving the coordinates needed for cell-state and tissue-scale analysis. This section owns the scaling of targeted imaging into multiplexed spatial platforms, the assignment of decoded molecules to cells and tissue compartments, integration with single-cell or spatial-capture atlases, and platform-level benchmarking. [Chapter 123](chapter1155.md) owns the foundational measurement physics of smFISH and branched-DNA detection, including probe and standard design, molecule counting, calibration, dynamic range, detection limits, and basic optical artifacts. In brief, smFISH uses multiple fluorescent probes targeting one RNA molecule so that true transcripts appear as diffraction-limited spots; here it serves as the targeted reference point against which higher-plex spatial platforms are compared.

Amplified in situ hybridization can place low-abundance targets into a spatial validation panel by building branched DNA or other signal-amplification structures on probe-bound RNAs. In this chapter, the important downstream question is whether amplified signals can be registered, segmented, and compared across cells, fields, sections, and atlas-scale datasets without being mistaken for directly comparable single-molecule counts. The underlying amplification chemistry, specificity controls, and quantitative consequences are treated in [Chapter 123](chapter1155.md).

Highly multiplexed imaging methods such as MERFISH- and seqFISH-style strategies use combinatorial barcodes, sequential rounds of hybridization or imaging, and error-correcting designs to detect many RNA species in the same specimen. These methods can reach single-cell and subcellular resolution for targeted panels. Their strength is direct localization of chosen genes in tissue architecture. Their limitations include panel design bias, optical crowding, registration error between imaging rounds, photobleaching, autofluorescence, tissue clearing or expansion effects, segmentation uncertainty, and the difficulty of comparing counts across genes with different probe performance.

In situ sequencing reads sequence information inside fixed material. Some designs use padlock probes, rolling-circle amplification, sequencing-by-ligation, sequencing-by-synthesis, or hybridization cycles to decode targeted transcripts or mutations at spatial coordinates. In situ sequencing can distinguish sequence variants, targeted isoforms, or selected gene panels while retaining tissue context. Its limitations include probe specificity, amplification bias, optical density, chemistry efficiency, and image-based base-calling error. It is often best understood as a spatial targeted sequencing assay rather than as a universal replacement for dissociated RNA-seq.

Segmentation is a hidden molecular decision in imaging-based methods. To assign RNA spots to cells, an analysis often needs nuclear stains, membrane stains, cytoplasmic markers, morphology, or learned segmentation models. Dense tissue, irregular cell shapes, thin processes, syncytia, necrosis, and overlapping cells can make segmentation wrong. If a transcript sits near a cell boundary, the assigned cell identity can change with the segmentation algorithm. Subcellular localization claims must also distinguish true RNA position from optical blur, tissue thickness, registration error, and probe diffusion.

Imaging methods are often the best validation for single-cell discoveries. If scRNA-seq identifies a rare cell-state marker, RNA imaging can test whether the marker appears in intact tissue, whether the cells occupy a plausible niche, and whether the signal is present in multiple donors or sections. If spatial barcoding suggests a ligand-receptor niche, multiplexed imaging can test whether the ligand-expressing and receptor-expressing cells are actually adjacent. If a dissociation experiment suggests a stress state, imaging can test whether the same transcript pattern existed before dissociation.

Table 130.2 should compare spatial capture, region-of-interest sequencing, smFISH-style imaging, MERFISH/seqFISH-style imaging, and in situ sequencing.

**Table 130.2. Spatial and imaging RNA technology families.** Compare the major spatial RNA measurement families by output, strengths, and limitations.

| Technology family | Molecular readout | Typical spatial unit | Strengths | Limitations and artifacts | Best-fit questions |
| --- | --- | --- | --- | --- | --- |
| **Slide-based spatial capture** | Sequenced RNA captured by positional oligonucleotides | Barcoded spot or array position, often multicellular | Broad gene measurement with histology and tissue coordinates | Spot mixing, RNA diffusion, permeabilization bias, limited single-cell assignment | Tissue architecture, tumor margins, regional programs, atlas-scale localization |
| **High-definition bead or pixel capture** | Sequenced RNA captured by smaller positional features | Bead, pixel, or subspot feature, sometimes below cell diameter | Finer spatial sampling than moderate-resolution spots | Lower reads per feature, partial cells, segmentation dependence, diffusion still matters | Epithelial layers, glomeruli, immune niches, viral foci, fine tissue boundaries |
| **Region-of-interest sequencing** | RNA from selected tissue regions or compartments | Analyst- or instrument-defined region | Uses pathology, morphology, or immunostaining to guide sampling | Region mixtures, selection bias, limited discovery inside the chosen area | Tumor nests versus stroma, lesions, crypts, brain layers, FFPE-compatible studies |
| **smFISH and amplified targeted imaging** | Probe-detected RNA spots or amplified hybridization signal | Molecule-like spot assigned to cell, compartment, or coordinate | Strong localization evidence for selected transcripts | Probe specificity, optical background, amplification bias, segmentation errors | Validation of rare markers, subcellular localization, tissue presence of nominated genes |
| **MERFISH and seqFISH-style multiplexed imaging** | Sequentially decoded combinatorial probe barcodes | Targeted transcript spots across cells or tissue fields | Multiplexed single-cell or subcellular localization for chosen panels | Panel design bias, registration error, crowding, photobleaching, segmentation uncertainty | Spatial cell states, neighborhoods, ligand-receptor adjacency, targeted atlas validation |
| **In situ sequencing** | In-place decoded cDNA, amplicon, ligation, or synthesis products | Decoded amplicon or spot in fixed material | Adds sequence discrimination while retaining location | Probe design, amplification bias, base-calling error, optical density limits | Targeted variants, selected isoforms, mutation-aware spatial mapping |
| **Spatial deconvolution with scRNA-seq reference** | Model-inferred cell-type or state mixture from spatial counts | Spot, region, pixel, or segmented cell group | Interprets mixed spatial units using dissociated reference profiles | Reference incompleteness, platform mismatch, missing disease states, model assumptions | Cell-type composition in mixed spots, spatial redistribution, reference-guided annotation |

The main boundary case is molecular breadth. Imaging panels detect what they are designed to detect. A beautifully resolved panel cannot discover genes that were not included. Conversely, sequencing-based spatial capture can survey broad transcriptomes but may not provide single-molecule localization. The strongest designs often combine discovery and validation: broad scRNA-seq or spatial capture to nominate states and genes, followed by targeted imaging to place selected molecules in intact tissue with higher spatial confidence. Readers designing or interpreting the targeted detection assay itself should continue to [Chapter 123](chapter1155.md); readers scaling those measurements into spatial cell atlases, comparing platform families, or integrating them with sequencing-based profiles remain in this chapter.

## 130.5. Multiome, CITE-seq, perturb-seq, and lineage-coupled technologies

Multimodal single-cell assays measure RNA together with at least one additional feature in the same cell or nucleus. The reason is straightforward: RNA state alone is informative but incomplete. Cell-surface proteins define many immune phenotypes better than transcripts alone. Chromatin accessibility can indicate regulatory potential. A guide RNA can identify a perturbation. A genotype can identify donor or clone. A lineage barcode can record ancestry. A spatial coordinate can place a cell in tissue. If these features share the same cell barcode, the analysis can ask how molecular state, identity, regulation, history, and context relate within the same captured unit.

Table 130.3 should compare the major multimodal readouts by molecule measured, biological question, and dominant artifact.

**Table 130.3. Multimodal readouts, best-fit questions, and dominant artifacts.** Help readers choose and interpret multimodal assays by modality rather than by platform brand.

| Modality paired with RNA | Molecule or feature measured | Best-fit biological question | Dominant artifact or confounder | Validation strategy |
| --- | --- | --- | --- | --- |
| **Antibody-derived protein tags** | Oligonucleotide-tagged antibodies for selected proteins | Immune phenotyping and RNA-protein state alignment | Antibody specificity, epitope accessibility, Fc binding, background counts | Panel titration, isotype or negative controls, flow or staining concordance |
| **Chromatin accessibility** | ATAC-like accessible DNA fragments with RNA from same nucleus | Candidate regulatory programs linked to expression states | Accessibility is regulatory potential, not enhancer activity or factor binding | Motif and peak-gene support plus perturbation, reporter, or protein evidence |
| **Sample hashing** | Oligonucleotide sample tags assigned to cells | Multiplex donors or conditions and detect cross-sample doublets | Tag background, threshold ambiguity, failed or unequal labeling | Genotype concordance, known sample markers, singlet and doublet threshold review |
| **Guide or perturbation barcode** | CRISPR guide RNA or linked perturbation identity | Transcriptomic phenotype of knockout, repression, activation, or drug-like perturbation | Guide dropout, incomplete target modulation, off-target effects, viability selection | Multi-guide agreement, target readout, non-targeting controls, rescue or time course |
| **Immune receptor sequence** | TCR or BCR variable-region sequence coupled to RNA state | Clonal expansion, antigen-response state, lineage of lymphocytes | Capture bias, pairing errors, incomplete receptor recovery | Orthogonal repertoire sequencing, clonotype replication, protein marker support |
| **Lineage barcode** | Heritable engineered barcode, scar, viral tag, or natural variant | Ancestry, fate choice, clonal state, tumor evolution | Barcode collision, incomplete sampling, selection, mutation timing | Multiple marks, time course, clone replication, known developmental or tumor context |
| **Genotype or clone marker** | Donor genotype, somatic variant, copy-number pattern, or mitochondrial variant | Donor demultiplexing, malignant clone assignment, allele-linked state | Sparse variant reads, allelic dropout, mapping bias, ambient RNA | Bulk genotype, matched DNA, copy-number consistency, replicate clone recovery |
| **Spatial coordinate** | Tissue position linked to cell or expression profile | Niche, neighborhood, anatomical layer, or spatial redistribution | Registration, segmentation, spot mixing, proximity overinterpretation | Imaging, histology, protein staining, replicated sections, perturbational support when causal |

CITE-seq and related antibody-derived tag methods use antibodies conjugated to oligonucleotides. Each antibody carries a sequence tag that identifies the protein target and can be counted by sequencing alongside RNA. In an immune-cell sample, RNA counts can separate broad transcriptional states while antibody tags for CD3, CD4, CD8, CD19, CD14, CD16, or activation markers can clarify cell identity and state. The method is targeted: it measures the selected protein panel, not the proteome. Interpretation depends on antibody specificity, epitope accessibility, staining conditions, background binding, Fc receptor effects, and normalization of antibody-derived counts.

Cell hashing and sample multiplexing use oligonucleotide tags to label samples before pooling. These tags can identify which donor or condition a cell came from and help detect doublets made from cells with different sample tags. Multiplexing can reduce batch effects by processing several samples together, but it cannot fix confounded experimental design if all cases were processed separately from controls or if sample tags fail. Hashing also adds its own background and thresholding problems. A demultiplexed label is an inferred assignment, not a physical guarantee.

Single-cell multiome assays commonly pair RNA with chromatin accessibility, often in isolated nuclei. Chromatin accessibility measured by ATAC-like chemistry indicates DNA regions accessible to transposase insertion, which can suggest promoters, enhancers, or regulatory elements. Joint RNA and accessibility can connect a gene-expression state to candidate regulatory regions in the same nucleus. The caveat is that accessibility is not direct proof of transcription-factor binding, enhancer activity, or causality. A region can be accessible without driving the observed transcript, and a regulatory effect can occur without a strong accessibility change.

Perturb-seq and related designs connect a perturbation identity to a transcriptomic phenotype. In a pooled CRISPR knockout, CRISPR interference, or CRISPR activation experiment, cells receive guide RNAs or linked barcodes. The single-cell library then captures both the RNA profile and the perturbation identity. The output can reveal how loss or activation of a gene changes cell states, pathways, differentiation, viral response, drug response, or synthetic circuits. This design is more informative than a pooled growth screen when the phenotype is a molecular state rather than survival alone.

Perturbational interpretation requires several controls. Guide identity must be recovered reliably. The perturbation must change the intended target, and the level of target modulation should be measured when possible. Multiple guides per target help separate on-target effects from guide-specific artifacts. Non-targeting guides, safe-targeting controls, rescue experiments, time courses, and orthogonal perturbation modalities strengthen causal claims. Viability selection can remove strongly affected cells before measurement. Cell-cycle redistribution can dominate expression changes. Multiplicity of infection can create multiple perturbations per cell. Off-target effects and incomplete knockdown can mislead target assignment.

Lineage-coupled assays recover ancestry or clonal identity together with RNA state. The lineage signal can come from engineered barcode libraries, CRISPR-edited scars, viral integration tags, immune receptor rearrangements, mitochondrial variants, somatic mutations, or naturally occurring clone markers. In development, lineage information can test whether two differentiated states share progenitors. In cancer, it can connect malignant clone, transcriptomic state, and microenvironment. In immune biology, receptor sequence can connect clonal expansion to activation or exhaustion. The main caveats are barcode collision, incomplete sampling, selection, mutation timing, and state changes after the lineage mark was created.

Figure 130.3 should illustrate how one cell barcode can link RNA, selected protein, chromatin accessibility, perturbation identity, and lineage history.

![Figure 130.3. Multimodal and perturbational barcode logic in one cell](../assets/figures/chapter1118_figure3.png)

**Figure 130.3. Multimodal and perturbational barcode logic in one cell.** Explain how the same cell barcode can connect RNA, selected protein, chromatin accessibility, guide identity, and lineage history.

Multimodal data are not self-validating. If RNA and protein agree, confidence can increase, but both measurements can be biased by sample handling or cell composition. If RNA and protein disagree, the difference may be biologically meaningful because mRNA and protein have different kinetics, localization, and stability; it may also reflect antibody background, dropout, delayed translation, or poor annotation. The point of multimodality is not to declare one layer correct. The point is to formulate claims that specify which molecule, feature, time scale, and evidence layer support the interpretation.

![Figure 130.6. From heritable lineage barcode reads to an ancestry tree linked with RNA state](../assets/figures/chapter1118_figure6.png)

**Figure 130.6. From heritable lineage barcode reads to an ancestry tree linked with RNA state.** Heritable lineage marks accumulate through cell divisions and can be recovered with RNA state, but convergent edits, unsampled intermediates, and capture dropout can distort the candidate ancestry tree. RNA state is overlaid after lineage reconstruction, and known sampling time provides an independent temporal check rather than defining the tree.

## 130.6. Cell-state annotation, batch correction, and integration

Cell-state annotation is the process of assigning biological meaning to a profile, cluster, neighborhood, trajectory position, spatial location, or perturbation response. Marker genes are often the first evidence. In human immune data, TRAC and CD3D support a T-cell label, MS4A1 supports a B-cell label, LST1 and S100A8 support myeloid states, NKG7 supports cytotoxic lymphocyte programs, PECAM1 supports endothelial identity, and EPCAM supports epithelial identity in many tissues. Marker genes are useful because they connect expression patterns to biological knowledge, but markers are not absolute. They can be shared, induced, suppressed, lost by dropout, species-specific, platform-sensitive, or misassigned by ambient RNA.

Reference mapping compares new profiles to a curated atlas. A reference can stabilize annotation, especially in large tissues with many related states. It can also bias interpretation. A healthy reference may not contain diseased states, fetal programs, treated cells, infected cells, tumor-specific states, non-model species, or rare intermediates. Mapping algorithms may force novel biology into the nearest known label. Good annotation therefore includes uncertainty, evidence markers, reference version, and a category for unknown or ambiguous states. A cell label should be a testable summary of evidence, not a mandatory output of software.

Clustering is not equivalent to cell-type discovery. Clustering partitions cells based on distances in a transformed feature space. The result depends on normalization, gene selection, dimensionality reduction, number of neighbors, distance metric, batch correction, and resolution parameters. A cluster can represent a cell type, a transient activation state, cell cycle, donor effect, dissociation stress, doublets, sequencing depth, mitochondrial RNA, or ambient contamination. A true biological continuum can be split into artificial clusters. A discrete cell type can be split by batch. The safest language distinguishes computational clusters from annotated cell types and from cell states.

> **Box 130.2. Questions Before Accepting a Rare Single-Cell Cluster**
>
> Before naming a rare cluster, ask whether the cluster is present in more than one biological replicate, donor, animal, section, or perturbation replicate. Check whether the marker genes are specific, plausible for the tissue, and not dominated by mitochondrial RNA, ribosomal RNA, stress genes, hemoglobin, dissociation markers, or ambient transcripts from abundant neighboring cell types. Inspect UMI counts, detected genes, doublet scores, sample hashes, genotype labels, and incompatible marker combinations. Test whether the cluster survives reasonable changes in filtering, normalization, feature selection, clustering resolution, and integration strategy. Ask whether the cluster is enriched in one batch, one lane, one tissue-processing condition, or one donor. Stronger support comes from protein staining, RNA imaging, index sorting, lineage or genotype evidence, perturbation response, or replicated spatial location. A rare cluster can be important, but rarity raises the burden of artifact control.

Batch correction and integration address unwanted or structured variation across samples, donors, lanes, protocols, tissues, species, or modalities. Methods may remove variation associated with batch labels, align mutual nearest neighbors, find shared latent variables, use anchors, apply matrix factorization, or train nonlinear embeddings. These tools are essential for atlas construction and cross-study comparison. The danger is overcorrection. If all disease samples were processed in one batch and all controls in another, an integration method cannot know which differences are technical and which are biological. It may erase true disease biology or preserve technical artifacts.

Undercorrection is the opposite problem. If technical batches remain separated, analysts may call batch-specific clusters new cell types. Donor identity, sequencing lane, dissociation enzyme, time from tissue removal to fixation, nuclei isolation buffer, freezing history, spatial slide, imaging round, antibody lot, and guide library batch can all create structure. Good design distributes biological conditions across batches, includes replicates, randomizes processing, records metadata, and uses integration as a tool rather than a substitute for design.

Differential expression requires special care because cells are not always independent biological replicates. If a study compares disease and control donors, thousands of cells from one donor should not be treated as thousands of independent donors. The experimental unit is often the donor, animal, organoid line, culture, tissue region, perturbation replicate, or spatial section. Pseudobulk analysis aggregates counts by sample and cell type or state, then uses bulk-like statistical models that respect replicate structure. Cell-level models can be useful, but they must model donor, batch, sample, and overdispersion. Ignoring replicate structure inflates confidence and can turn compositional sampling differences into false molecular claims.

Integration across modalities also carries assumptions. Mapping scRNA-seq profiles onto spatial data can assign likely cell types to spots, but the answer depends on reference completeness and platform compatibility. Mapping RNA onto chromatin accessibility can suggest gene-regulatory relationships, but accessibility and expression are different molecular layers. Mapping perturbation responses onto an atlas can name states, but perturbation may create states absent from the reference. A model can be useful and still be wrong for a specific tissue, disease, species, or experimental condition.

Figure 130.4 should show where confounders enter from study design through sample handling, library construction, normalization, integration, annotation, and reported analysis output.

![Figure 130.4. Where confounders enter a single-cell or spatial RNA study](../assets/figures/chapter1118_figure4.png)

**Figure 130.4. Where confounders enter a single-cell or spatial RNA study.** Make clear that computational interpretation begins with experimental design and sample handling.

## 130.7. Benchmarking, confounders, failure modes, and analysis limits

Benchmarking asks whether a method or analysis recovers known truth, reproducible signal, or biologically plausible structure under defined conditions. For scRNA-seq and snRNA-seq, benchmarks may compare molecule recovery, gene detection, cell-type representation, doublet rates, ambient RNA effects, replicate consistency, differential expression, and agreement with bulk RNA-seq or imaging. For spatial methods, benchmarks may compare resolution, sensitivity, cell segmentation, tissue preservation, spatially variable gene detection, deconvolution accuracy, and agreement with histology. For perturbational assays, benchmarks may compare guide capture, target modulation, false-positive rates, reproducibility across guides, and known pathway responses.

Benchmark results are context-dependent. A method that works well in peripheral blood may perform poorly in fibrotic tissue, adipose tissue, plant tissue, infected lung, necrotic tumor, or formalin-fixed material. A doublet detection method calibrated for diverse cell types may miss doublets among similar subtypes. An ambient RNA correction method may remove true extracellular or phagocytosed RNA in a biological setting where such RNA is part of the question. A spatial segmentation model trained on one tissue morphology may fail in another. Benchmarking should therefore be read as evidence under specified sample, tissue, platform, and analysis conditions.

Dissociation is a major biological confounder. Enzymatic digestion, temperature, mechanical stress, hypoxia, delays, and cell sorting can induce immediate-early genes, heat-shock genes, interferon programs, stress responses, apoptosis, or selective cell loss. Fragile neurons, epithelial cells, granulocytes, adipocytes, and large cells can be underrepresented in whole-cell suspensions. Nuclei isolation reduces some dissociation problems but introduces nuclear-compartment bias and can lose cytoplasmic programs. A study that compares conditions must ask whether one condition dissociated more easily, died faster, or yielded different proportions of damaged cells.

Downstream analyses such as differential abundance, ligand-receptor scoring, trajectory inference, RNA velocity, clone-state association, and perturbation-response modeling require benchmarks matched to their estimands and assumptions. In this chapter, these outputs are evaluated for replicate structure, calibration, sensitivity to preprocessing, reference dependence, negative controls, simulated or experimental ground truth, and robustness to alternative models. [Chapter 106](chapter1101.md) owns what validated outputs reveal about tissues, development, disease, lineage decisions, or causal cell-state mechanisms.

## Recent Consensus

The current consensus is that single-cell and single-nucleus RNA technologies are sparse, biased, high-dimensional sampling methods rather than complete molecular censuses. They support atlas construction and comparative analysis, but remain weaker for absolute RNA quantification, complete isoform reconstruction, and causal inference without additional designs and validation.

snRNA-seq is often the preferred method for frozen, archived, fragile, or complex tissues, not merely a fallback for failed cell dissociation. It still measures a different RNA compartment than scRNA-seq, so combined analyses must address nuclear, intronic, cytoplasmic, mitochondrial, and localized RNA differences.

Spatial barcode capture, high-definition capture, region-of-interest sequencing, multiplexed imaging, and in situ sequencing are complementary. Multimodal and perturbational assays are also complementary, but joint measurement strengthens interpretation only when each modality is controlled. Across all of these technologies, balanced experimental design, biological replication, metadata preservation, orthogonal validation, and replicate-aware statistics remain more important than post hoc correction.

## Open Questions, Controversies, Deprecated Models, and Common Misconceptions

Open questions:

- How can researchers define a cell state across platforms?
- How can researchers benchmark spatial methods across tissues?
- How can researchers model total RNA content?
- How can researchers evaluate foundation-model predictions for denoising, annotation, imputation, and perturbation response?
- Perturbational screens require special caution because guide-associated expression changes can reflect direct target action, compensation, growth selection, death, off-target activity, cell-cycle redistribution, incomplete knockdown, or timing. Multi-guide agreement, rescue, orthogonal perturbation, time course, and pathway validation help separate these possibilities.

Common misconceptions:

- "A single-cell zero means the gene is absent." Dropout, sampling depth, capture efficiency, and expression bursts can create zeros even when transcripts are present.
- "A computational cluster is automatically a cell type." Clusters need marker, replicate, spatial, developmental, or functional context before being named as cell types.
- "Single-nucleus RNA-seq is simply shallow single-cell RNA-seq." Nuclear RNA and whole-cell RNA sample different compartments and can emphasize different genes, isoforms, and processing states.
- "Chromatin accessibility proves enhancer activity." Accessibility marks candidate regulatory DNA; enhancer activity requires transcriptional, perturbation, or reporter evidence in context.
- "Batch correction is neutral." Batch correction can remove real biological signal when biology and technical variables are confounded.
- "Multimodal data validate themselves." Multimodal measurements still need controls, calibration, and consistency checks across assays.
