This chapter explains how cell-resolved, spatial, temporal, lineage, and perturbational RNA measurements change biological understanding. It owns biological discoveries, cross-modal synthesis, and the causal questions that these data can support: which RNA programs define a state, where states occur, how states change, which histories constrain them, and which regulators or environmental signals are necessary or sufficient. It does not teach platform chemistry, count normalization, integration algorithms, or benchmarking workflows; those methods and their comparative evaluation belong to Chapter 130. Technology names appear only where the evidence type determines the biological claim.
Cell-resolved RNA data changed the unit of transcriptomic reasoning from an averaged tissue to a distribution of cells. The gain is not simply finer resolution. It allows investigators to separate stable cell identities from reversible states, identify rare populations, quantify state mixtures, and ask whether apparently uniform tissues contain distinct response programs. The central biological object is therefore not a cluster but a reproducible RNA program linked to lineage, anatomy, function, or response. A cluster is a provisional partition of observations; a cell type or state is a biological interpretation that needs independent support [Cheng 2023; Li and Wang 2021].
Spatial RNA data add anatomical constraint. Cell states that look similar after dissociation can occupy different tissue layers, interfaces, inflammatory foci, or stromal niches. Conversely, one anatomical region can contain several interacting cell states. Spatial adjacency can nominate communication, but RNA expression and proximity alone do not establish secretion, receptor activation, or signaling. The biological advance comes when spatial patterns are connected to morphology, protein localization, pathway response, perturbation, or a reproducible tissue outcome [Longo 2021; Moncada 2020].
Temporal and lineage evidence answer questions that a static atlas cannot. Repeated sampling reveals population-level response waves; nascent-RNA measurements distinguish new transcription from persistence of older RNA; trajectory and RNA-velocity models propose directions through state space; lineage records reveal shared ancestry. These dimensions are complementary, not interchangeable. Similar RNA states can arise from different lineages, one lineage can diversify into several states, and a smooth trajectory can reflect cell cycle or stress rather than differentiation. Strong dynamic claims state whether time, direction, and ancestry were observed or inferred.
Perturbational data move from association toward causality by changing a regulator or environmental input and reading the resulting cell-state distribution. A perturbation can reveal a direct response, a downstream compensatory program, altered survival, or redistribution among states. Causal interpretation therefore requires a chain from perturbation to target engagement, molecular response, state transition, and function. Multiple guides, dose or timing information, rescue, and orthogonal readouts distinguish regulator-specific effects from toxicity or selection [Dixit 2016; Replogle 2022].
Cross-study synthesis is strongest when independent evidence types constrain the same explanation. Cell-resolved profiles define a program; spatial data place it; temporal or lineage evidence order it; perturbation tests a regulator; and functional assays show consequence. Agreement raises confidence, while disagreement is diagnostically useful. The goal is not to fuse every dataset into one latent map, but to build a biological model whose claims remain traceable to measurements and whose alternatives are explicit.
RNA abundance reflects transcription, processing, localization, and decay. It is not a direct measurement of protein abundance, biochemical activity, morphology, ancestry, or future fate. A cell-resolved RNA profile is also an incomplete sample of the cell’s molecules. Chapter 130 explains capture, normalization, integration, and platform-specific artifacts; this chapter asks what biological inferences remain justified after those issues have been addressed.
Observation must be separated from interpretation. RNA counts and tissue coordinates are observations. A label such as “exhausted T cell,” a proposed developmental branch, or a predicted ligand-receptor interaction is an inference. A perturbation identity may be assigned experimentally, but the causal mechanism linking that perturbation to a cell-state change is still an inference. The evidence standard rises with claim strength.
Four recurring examples organize the discussion. Development illustrates trajectories, lineage restriction, and spatial patterning. Immune responses illustrate reversible states and rapidly changing programs. Tumors illustrate clone-state relationships, microenvironmental niches, and selection. Nervous tissue illustrates stable identities embedded in strong anatomical organization. The examples are not platform tutorials; they show how biological conclusions change when cell, place, time, ancestry, and intervention are jointly considered.
Bulk tissue expression can change because each cell changes its RNA program, because the abundance of cell populations changes, or both. Cell-resolved data separate these possibilities. An inflammatory tissue may show higher cytokine RNA because resident macrophages induce a response, infiltrating cells enter the tissue, a rare producer population expands, or damaged cells alter the sampled composition. The first biological task is to decompose a bulk difference into state changes and compositional changes without assuming that either one dominates.
Cell identity is hierarchical. Broad classes such as epithelial, immune, stromal, and neural cells divide into lineages, differentiated types, local subtypes, and transient states. The useful level depends on the question. A coarse macrophage label may be adequate for tissue composition but inadequate for distinguishing inflammatory, reparative, resident, and recruited programs. Excessive subdivision creates fragile labels that reproduce neither across individuals nor across conditions. A defensible taxonomy therefore combines stable core programs with context-dependent state modules.
The distinction between type and state is mechanistic. A cell type retains lineage-linked regulatory architecture and characteristic functions across several conditions. A state can be induced by cytokines, hypoxia, stress, cell cycle, injury, or therapy and may recur in multiple types. Interferon-response RNAs, for example, can rise across epithelial cells, lymphocytes, and fibroblasts. Treating the shared program as a new cell type confuses the inducing signal with lineage identity. Conversely, suppressing state-associated genes during annotation can erase biologically meaningful responses.
Box 106.1. A Cluster Is a Hypothesis, Not a Cell Type
A cluster becomes a credible biological identity through reproducibility, a coherent multi-gene program, expected negative markers, stability across reasonable analyses, anatomical or lineage context, and orthogonal morphology, protein, or function. Rare-state claims additionally require exclusion of mixtures and damage responses. The box contrasts a computational partition with stable type, reversible state, and artifact explanations.
Marker evidence should be compositional. A useful identity is supported by several concordant genes, expected absences, regulatory factors, receptors, effector molecules, and tissue context. A newly proposed type should reproduce across individuals or independent cohorts and should not depend on one computational partition. Protein staining, morphology, anatomical position, lineage evidence, and functional assays strengthen the interpretation. Villani and colleagues’ human dendritic-cell study exemplifies the discovery value of cell-resolved profiling, while also illustrating why new immune categories require validation beyond the initial map [Villani 2017].
Rare states demand special care because rarity magnifies both biological value and artifact risk. A small population may be a transient precursor, treatment-resistant clone, activated immune state, doublet-like mixture, or damaged-cell response. A credible rare-state claim should show coherent genes, replicate recovery, plausible abundance, stable placement under reasonable analysis choices, and an orthogonal feature that can be measured outside the discovery dataset.
Cross-tissue comparison asks whether similarly named states are homologous, convergent, or merely correlated. Fibroblasts from lung, liver, bowel, and tumor may share extracellular-matrix programs yet differ in developmental origin and local signaling. Tissue-resident macrophages can share phagocytic functions while maintaining organ-specific regulatory programs. Synthesis should preserve both layers: a reusable state module and the tissue-specific program that gives it biological meaning.
Position converts a cell-state catalog into tissue biology. A program may be concentrated at an epithelial boundary, around a vessel, within a germinal center, along a developmental axis, or next to an invasive tumor front. The same RNA program can have different implications in different locations because available ligands, extracellular matrix, oxygen, nutrients, and neighboring cells differ. A spatial niche is therefore more than co-occurrence: it is a local environment hypothesized to maintain or constrain a state.
Spatial synthesis begins by distinguishing objects. A coordinate may refer to a single molecule, part of a cell, one cell, or a multicellular region depending on the evidence. Biological conclusions should match that scale. A mixed region can reveal a reproducible neighborhood but cannot by itself assign every transcript to a specific cell. A subcellular signal can reveal polarized RNA localization but should not automatically be interpreted as a tissue-level cell state. Technical resolution and segmentation belong to Chapter 130; the biological rule here is to name the unit represented by each spatial claim.
The strongest niches recur across specimens and align with independent anatomy. In pancreatic cancer, combined single-cell and spatial analyses have related malignant, stromal, and immune programs to tissue architecture [Moncada 2020]. In inflammatory bowel disease, spatially resolved myeloid heterogeneity illustrates how inflammatory programs can be organized around local tissue damage and cellular neighborhoods [Garrido-Trigo 2023]. These maps nominate relationships; they do not by themselves show which neighbor maintains which state.
Gradients require a mechanism. A gradual change in RNA abundance across tissue may reflect diffusion of a signal, varying cell composition, developmental age, oxygen tension, mechanical stress, or section geometry. A morphogen-like interpretation becomes stronger when the candidate signal has a plausible source, receptor distribution, downstream target program, appropriate length scale, and perturbational response. Otherwise “gradient” is a spatial description rather than an explanation.
Cell-cell communication is an evidence ladder. Ligand RNA in a candidate source and receptor RNA in a candidate target establish compatibility. Spatial proximity establishes opportunity. Protein localization, ligand release or presentation, receptor activation, and downstream signaling establish increasingly specific steps. Removing the ligand, receptor, or source population and observing loss of the target-state program supplies causal evidence. The model should specify source, target, mediator, distance, timing, and measured consequence.
Spatial absence is also informative but ambiguous. A state found after dissociation but absent in tissue may be rare, induced during processing, missing from the sampled section, or below spatial detection. A tissue region without a matching dissociated state may contain fragile cells, a mixed neighborhood, or a context-specific program absent from the reference. These disagreements motivate new sampling or imaging rather than forced label transfer.
A static atlas samples many cells once. It can reveal a continuum, but the continuum is not a movie. Temporal reasoning needs at least three dimensions: chronological time, direction of state change, and ancestry. Chronological sampling records when cells were collected. Nascent-RNA information reports recent synthesis. Trajectory or RNA-velocity models infer direction. Lineage records show descent. Each constrains a different part of the history.
Response programs often occur as waves. Immediate-early transcription factors can rise within minutes, effector RNAs follow, feedback inhibitors limit the response, and later remodeling stabilizes or reverses the state. Total RNA abundance mixes new synthesis with persistence and decay. Nascent-RNA measurements can distinguish rapid induction from stable carryover, while dense real-time sampling can reveal transient programs that a single late endpoint misses [Mahat 2024]. This matters whenever an apparent “state” may be a brief phase of a response.
Trajectory inference proposes an ordering through transcriptional states. It is most persuasive when intermediates are densely sampled, branch markers change coherently, measured time agrees with the order, and perturbation shifts occupancy or progression as predicted. A trajectory is weakened when it follows cell cycle, stress, sequencing quality, or uneven sampling. A path through an embedding does not show that individual cells traversed that path.
RNA velocity adds a local directional hypothesis from relationships between nascent or unspliced and mature or spliced RNA. The biological use is not the arrow itself but the testable prediction: cells in one state should later enrich another state, and perturbing a proposed regulator should change that transition. Velocity can fail when transcription, splicing, or decay kinetics depart from the model, especially during abrupt responses or when different RNA compartments are compared [La Manno 2018]. Model details and diagnostics belong to Chapter 130.
Lineage tracing records ancestry. Shared lineage can reveal whether a clone diversifies into several states, whether similar states arose independently, or whether progenitors show fate bias. It cannot by itself identify the regulatory mechanism of commitment. A barcode tree with endpoint RNA states may miss transient intermediates, and incomplete recovery can make branching appear sharper than it was. Tumor lineage studies demonstrate how ancestry and state jointly reveal plasticity and evolutionary history that RNA similarity alone cannot recover [Yang 2022].
Box 106.2. Do Not Confuse Time, Direction, State, and Ancestry
Four questions are paired with evidence: when was the sample observed; what RNA program is present; in which direction might the state change; and which cells share descent. Measured time, nascent RNA, trajectory or velocity models, and lineage records answer different questions. The box shows how combined evidence supports a dynamic model without treating any one dimension as a substitute for the others.
Developmental decisions are often probabilistic. A progenitor program can bias a fate without irreversibly determining it. To claim fate bias, comparable progenitors must be related to later descendants, and alternative explanations such as differential survival or sampling must be addressed. Perturbation then asks whether a regulator changes transition probability, timing, or survival. This distinguishes a marker of impending fate from a driver of fate.
Convergence and divergence are central synthesis patterns. Different lineages can converge on similar injury-response or effector programs, while clonally related cells can diverge under distinct niches. Agreement between lineage, time, space, and RNA state supports a developmental path; disagreement can reveal plasticity, migration, convergent adaptation, or incomplete sampling.
Perturbational RNA data test what changes when a gene, enhancer, RNA, receptor, pathway, or environmental input is altered. The rich readout matters because two interventions with similar effects on growth can act through different programs. One may block ribosome production, another may trigger an interferon response, and a third may selectively remove a developmental state. Perturb-seq established the power of linking pooled genetic interventions to cell-resolved transcriptomic phenotypes [Dixit 2016; Replogle 2022]. Platform implementation belongs to Chapter 130 and Chapter 136.
The first causal distinction is direct response versus population selection. If a perturbation changes RNA within a stable cell population, it may alter regulation. If it kills one state or blocks its recovery, the observed dataset changes composition. Both are biological effects, but they imply different mechanisms. Time courses can separate early target-proximal responses from later compensation and selection. Measuring target engagement prevents a failed intervention from being misread as biological resistance.

Figure 106.3. From Perturbation to a Causal Cell-State Mechanism. The figure follows one intervention through target engagement, early molecular response, state redistribution or transition, and functional consequence. Parallel branches show alternative explanations: failed perturbation, toxicity, altered survival, compensation, and delayed selection. Evidence callouts mark concordant interventions, time course, direct-target assays, rescue, lineage or spatial validation, and functional readout. The visual teaches a causal chain rather than a Perturb-seq platform workflow.
A causal chain should name intermediate steps. For a transcription factor, the chain might be perturbation, altered factor activity, changed target-gene program, shifted state transition, and functional consequence. Chromatin binding or reporter evidence can support the target step; rescue can support specificity; lineage or time-resolved data can support the transition; and a phenotype can establish consequence. A transcriptomic response alone usually identifies downstream association, not direct molecular targets.
Guide consistency, orthogonal interventions, dose response, and rescue address different alternatives. Concordant guides reduce the chance of guide-specific effects. A chemically or genetically distinct intervention tests whether the phenotype follows the intended pathway. Dose response can distinguish thresholded state transitions from toxicity. Rescue with a perturbation-resistant molecule tests specificity, although overexpression can create its own artifacts.
Epistasis turns lists of regulators into pathway structure. If perturbing two genes produces a response predicted from one downstream process rather than the sum of independent effects, the data can order or connect regulators. High-dimensional RNA phenotypes can reveal that two perturbations converge on the same program while differing in secondary effects. Epistasis claims still require adequate perturbation strength and consideration of ceiling effects, survival, and timing.
Context dependence is biological information. A regulator may control a state only in a particular lineage, niche, developmental window, or inflammatory environment. In vivo perturbational studies can expose dependencies absent in culture, while also introducing selection and delivery biases. A robust conclusion states the cells that received the perturbation, the context in which the phenotype appeared, and whether the same mechanism reproduced in an independent system.
Causal cell-state maps differ from atlases. An atlas records observed states and associations. A causal map links regulators and signals to transitions under specified conditions. Edges should be graded: nominated by association, supported by perturbation, placed by temporal evidence, connected to direct targets, and validated by rescue or function. This prevents large network diagrams from implying equal certainty for every edge.
Cross-study synthesis begins with biological comparability, not computational alignment. Studies can be compared when tissue region, developmental stage, treatment, disease definition, sampling time, and cell-state criteria are sufficiently matched. Similar labels may hide different programs, while different labels may describe the same conserved response. The synthesis should therefore compare program components and biological context before accepting names as equivalent.
Replication has several levels. Technical replication asks whether the measurement is stable. Biological replication asks whether a state or relationship recurs across individuals. Contextual replication asks whether it appears across tissues, cohorts, or perturbations where the mechanism predicts it should. Mechanistic replication asks whether the same regulator-state link survives independent perturbation or validation. A ubiquitous program may replicate broadly; a niche-specific mechanism should replicate only in the relevant context.
No single modality is a universal reference. Cell-resolved profiles distinguish programs but lose anatomy. Spatial evidence constrains location but may represent mixtures. Temporal data order responses but usually sample different cells. Lineage records show ancestry but not regulatory mechanism. Perturbations support causality but can alter survival and composition. Strong synthesis assigns each claim to the evidence type best suited to test it.
Evidence convergence is more informative than dataset fusion. Consider a candidate stromal program at a tumor boundary. Cell-resolved RNA data define the program; spatial data place it near invasive epithelium; repeated sampling links it to progression; ligand-receptor analysis nominates a signal; perturbing the signal changes the program; and functional assays alter invasion. Each step removes alternatives. A joint embedding alone cannot replace this chain.

Figure 106.4. Productive Disagreement among Biological Evidence Types. Disagreement among cell-state, spatial, lineage, RNA, protein, and functional evidence can reveal real context dependence or processing, sampling, timing, model, and causal errors; the appropriate endpoint is a revised experiment rather than forced agreement.
Disagreement should be preserved until explained. If a velocity model predicts one branch but lineage data support another, the kinetic model, sampling window, or branch definition may be wrong. If an atlas state lacks spatial support, it may be processing-induced or geographically rare. If a perturbation changes RNA but not protein or function, the response may be compensatory. Reporting such conflicts prevents apparent consensus from being manufactured by smoothing.

Figure 106.5. Four Biological Explanations for One Bulk RNA Change. new ordered marker in Section 106.5 after the cross-modal disagreement paragraph so retired suffixes remain unused and active figure IDs stay in reader order. The same increase in a tissue-level RNA signal can arise from induction within an unchanged population, redistribution among states of the same cell type, infiltration or loss that changes population abundance, or a mixture of state and composition effects. Matched reference-versus-condition miniatures keep cell shape, fill, and RNA-dot abundance as separate visual variables, and each model is paired with a validation appropriate to its causal claim.
Causal language should be graded. “Associated with” is appropriate for co-occurrence. “Predicts” requires out-of-sample performance. “Precedes” requires time. “Required” requires loss-of-function under defined conditions. “Sufficient” requires induction in an appropriate context. “Directly regulates” requires evidence that excludes intermediates, often binding or rapid target-proximal response. “Mediates” requires perturbing the intermediate. The verb should reveal the evidence.

Figure 106.6. From Spatial Co-Occurrence to a Defensible Tissue Niche. new ordered marker in Section 106.5 after the causal-language paragraph so retired suffixes remain unused and active figure IDs stay in reader order. Spatial signal is first assigned to a biological unit, with single-cell and mixed-region interpretations kept separate when segmentation is uncertain. A candidate niche becomes progressively stronger when the neighborhood recurs across specimens, aligns with independent anatomy, presents a plausible local signaling opportunity, responds to perturbation, and has a measured function. The figure preserves the alternative that co-occurrence reflects shared exposure or sampling rather than niche maintenance.
Communication inference illustrates these limits. Ligand and receptor RNAs, spatial proximity, and correlated target programs nominate a pathway. They do not establish protein availability, physical encounter, receptor activation, or necessity. Perturbing the source, mediator, receptor, or target pathway and measuring the predicted response is required for a strong signaling mechanism.

Figure 106.7. Time, Direction, and Ancestry Constrain Different Histories. new ordered marker in Section 106.5 after the communication-inference paragraph so retired suffixes remain unused and active figure IDs stay in reader order. Four aligned evidence lanes distinguish chronological sampling, recent synthesis from nascent RNA, model-based state direction, and heritable lineage records. Candidate histories are then filtered without forcing false a convergent state can remain possible from distinct ancestors, and a similarity-only ancestry claim is rejected when lineage records conflict.
Box 106.3. Evidence Ladder for Cell-Cell Communication
The ladder moves from ligand-receptor RNA compatibility through spatial opportunity, protein availability, receptor or pathway activation, perturbation of the source or mediator, and loss or rescue of the predicted target-cell response. Each rung names the remaining alternative explanations. The teaching goal is to prevent communication diagrams from being presented as signaling mechanisms.
In development, synthesis links progenitor programs to spatial axes, measured time, lineage branches, and perturbation-defined regulators. Early mixed programs may represent competence rather than committed hybrid types. The key question is whether a program predicts and causally influences fate after accounting for position and ancestry.
In immunology, activation programs cross conventional type boundaries. Interferon, stress, cycling, exhaustion, memory, and tissue-residency modules must be separated from lineage identity. Spatial data reveal immune niches, while perturbation tests receptors and transcriptional regulators that maintain them.
In cancer, clone, state, and niche are distinct dimensions. One clone can occupy several phenotypic states, and unrelated clones can converge under hypoxia, therapy, or immune pressure. Spatial and lineage evidence prevent transcriptional similarity from being mistaken for shared ancestry, while perturbation distinguishes vulnerabilities from correlates.
In the nervous system, stable molecular identities interact with region, layer, projection, activity, and disease-associated states. Spatial evidence is especially important because anatomical organization is integral to function. Cell-resolved RNA programs nominate types and responses; physiology, connectivity, morphology, and perturbation establish their meaning.
The biological synthesis in this chapter depends on technology but does not duplicate it. Chapter 130 owns platform chemistry, normalization, integration, benchmarking, segmentation, and modality-specific quality control. Chapter 136 owns pooled-screen implementation. Readers should return here after those workflows to judge whether the resulting evidence supports a cell identity, niche, trajectory, lineage decision, regulator, or causal chain.
Translational use requires compression from discovery programs into robust measurements. A complex state may eventually be recognized by a small RNA panel, protein markers, morphology, or a functional assay. The reduced test must be validated across specimens and contexts rather than assumed to inherit the discovery model’s accuracy. Clinical and engineering decisions should rely on interpretable state definitions and demonstrated links to outcome.
Cell types and states are biological models, not automatic outputs of clustering. Spatial position, measured time, lineage, perturbation, protein, morphology, and function provide orthogonal evidence that can validate or revise RNA-based categories. Tissue niches and trajectories are most useful when they generate explicit, testable mechanisms.
Pseudotime is not chronological time, velocity is not observed fate, lineage is not state, proximity is not signaling, and perturbation response is not necessarily direct mechanism. Cross-modal agreement strengthens inference only when the modalities have independent failure modes and the biological units are comparable.
Perturbational data provide the strongest route from atlas to causal cell-state map, but target engagement, selection, timing, context, and rescue remain essential. Cross-study synthesis should preserve disagreement and grade causal language rather than forcing all data into one universal representation.
Open questions:
Controversies:
Common misconceptions:
Deprecated or weakened claims: