







Cell-state density characterizes the distribution of cells along phenotypic landscapes and is crucial for unraveling the mechanisms that drive diverse biological processes. Here, we present Mellon, an algorithm for estimation of cell-state densities from high-dimensional representations of single-cell data. We demonstrate Mellon’s efficacy by dissecting the density landscape of differentiating systems, revealing a consistent pattern of high-density regions corresponding to major cell types intertwined with low-density, rare transitory states. We present evidence implicating enhancer priming and the activation of master regulators in emergence of these transitory states. Mellon offers the flexibility to perform temporal interpolation of time-series data, providing a detailed view of cell-state dynamics during developmental processes. Mellon facilitates density estimation across various single-cell data modalities, scaling linearly with the number of cells. Our work underscores the importance of cell-state density in understanding the differentiation processes, and the potential of Mellon to provide insights into mechanisms guiding biological trajectories.
VAPOR: A variational autoencoder with transport operators to disentangle cellular gene expression dynamics of co-occurring biological processes in time and space
Abstract Single-cell and spatial transcriptomics enable the analysis of cellular states and dynamics in gene expression, revealing how diverse biological processes relate to these states over time and space. To study these dynamics, trajectory inference methods order cells along computationally inferred paths to reconstruct gradual transitions in cell states. However, by encouraging smooth and continuous trajectories, these approaches tend to conflate co-occurring processes-such as proliferation, maturation, and spatial organization-that are jointly reflected in gene expression, potentially overlooking process-specific gene expression dynamics. To address this, we developed VAPOR, which integrates a variational autoencoder with transport operators to model and disentangle cellular gene expression dynamics for potentially co-occurring biological processes. VAPOR inputs single-cell (or spatial) gene expression data into a variational autoencoder (VAE) to learn the latent states of cells and then models their latent dynamics as an ordinary differential equation. The latent dynamics are further decomposed into process-specific components parameterized by transport operators (TOs) and their corresponding process weights. Each TO defines a process-specific dynamics, and its weight for each cell quantifies the process's contribution to the cell dynamics. After assessment by simulation studies, we applied VAPOR with benchmarking to real data, including time-course scRNA-seq from postconceptual human brain development, spatial transcriptomics of the mouse hippocampus, and cross-species scRNA-seq spanning human and macaque first-trimester forebrain development. In these applications, VAPOR has identified a variety of temporal and spatial co-occurring processes, such as cell cycle, gliogenesis, neurogenesis, and neuronal migration, along with associated dynamic genes, including those species-specific to human and macaque development. VAPOR is available as an open-source tool for general-purpose use. ### Competing Interest Statement The authors have declared no competing interest.

Comparing phenotypic manifolds with Kompot: Detecting differential abundance and gene expression at single-cell resolution
Single-cell studies are frequently designed to compare across conditions such as health and disease. However, existing computational approaches typically rely on grouping cells into discrete populations before making comparisons, which can limit resolution for detecting state-dependent changes. Here, we introduce Kompot, a statistical framework for comparative analysis of multi-condition single-cell data. Kompot quantifies both differential abundance, capturing how cells redistribute across the phenotypic space, and differential expression, identifying condition-specific transcriptional changes that may be localized, heterogeneous, or oppositely regulated across states. By modeling cell density and gene expression as continuous functions over a shared cell-state representation, Kompot enables single-cell–resolution inference with principled uncertainty estimates, without requiring predefined clusters or cell types. Applying Kompot to aging murine bone marrow, we identified a continuum of shifts in hematopoietic stem cell and mature cell states, transcriptional remodeling of monocytes independent of compositional changes, and divergent regulation of oxidative stress response genes across cell types. We demonstrate the utility of Kompot in disease settings by identifying cell-state and gene expression changes associated with improved efficacy of combinatorial immunotherapy in melanoma. Additionally, Kompot enables multi-sample comparative analysis by accounting for sample-to-sample heterogeneity. By capturing both global and cell-state–specific effects of perturbation, the Kompot framework is broadly applicable to dissecting condition-specific effects in complex single-cell landscapes. ### Competing Interest Statement The authors have declared no competing interest. National Institutes of Health, R35GM147125, R01CA292932, T32GM136534, S10OD028685 The Mark Foundation for Cancer Research, https://ror.org/00v7th354, Endeavor Award Edward P. Evans Foundation, https://ror.org/03h22gm35, Discovery Research Grant Brotman Baty Institute, https://ror.org/03jxvbk42, Pilot Award

Temporal tissue dynamics from a spatial snapshot
Physiological and pathological processes such as inflammation and cancer emerge from interactions between cells over time1. However, methods to follow cell populations over time within the native context of a human tissue are lacking because a biopsy offers only a single snapshot. Here we present one-shot tissue dynamics reconstruction (OSDR), an approach to estimate a dynamical model of cell populations based on a single tissue sample. OSDR uses spatial proteomics to learn how the composition of cellular neighbourhoods influences division rate, providing a dynamical model of cell population change over time. We apply OSDR to human breast cancer data2–4, and reconstruct two fixed points of fibroblasts and macrophage interactions5,6. These fixed points correspond to hot and cold fibrosis7, in agreement with co-culture experiments that measured these dynamics directly8. We then use OSDR to discover a pulse-generating excitable circuit of T and B cells in the tumour microenvironment, suggesting temporal flares of anticancer immune responses. Finally, we study longitudinal biopsies from a triple-negative breast cancer clinical trial3, in which OSDR predicts the collapse of the tumour cell population in responders but not in non-responders, based on early-treatment biopsies. OSDR can be applied to a wide range of spatial proteomics assays to enable analysis of tissue dynamics based on patient biopsies.

Benchmarking gene expression reconstruction from single-cell latent representations
Single-cell transcriptomics is typically modeled in low-dimensional latent representations that improve the signal-to-noise ratio of the data. Such representations underpin data integration, cell state discovery, and perturbation prediction, with applications ranging from large-scale organ atlases to latent trajectory modeling. Recent virtual cell approaches further leverage these representations to predict cellular responses as distributional shifts in latent space. Each of these applications ultimately requires faithful gene expression reconstruction from latent spaces for biological interpretation, enabling gene-level analysis of predicted perturbed or batch-corrected cells. Yet representation choice is typically treated as an implementation detail rather than a primary modeling decision, with no systematic evaluation of how well latent representations support gene expression reconstruction. Here, we introduce ReconEval, a benchmark for evaluating gene expression reconstruction from single-cell latent spaces. We benchmark two classes of latent representations: end-to-end trained models such as PCA, autoencoders, and variational autoencoders, and pretrained single-cell foundation model embeddings coupled to newly trained decoders. Reconstruction is evaluated both directly and after latent-space perturbation prediction. Across perturbational and observational datasets totaling over 100 million cells, our metric suite quantifies statistical fidelity; biological signal preservation, including differential expression, coexpression, cell-cycle structure, cytokine response and pathway activity; and perturbation-specific effects. We find that autoencoders achieve the strongest stand-alone reconstruction at low dimensionality, while variational regularization does not improve generalization in reconstruction. Frozen foundation model embeddings retain recoverable gene-level information, with reconstruction quality depending strongly on decoder architecture and pretraining objective. In latent perturbation modeling, high-dimensional PCA matches foundation model embeddings, while low-dimensional AE embeddings are optimal for flow-based generative models. Overall, reconstruction depends critically on the interplay between representation and downstream model, and simpler representations can outperform complex alternatives given appropriate capacity. Our benchmark establishes reconstruction as a critical evaluation axis for single-cell foundation models. We envision it improving the biological interpretability of latent-space modeling, a prerequisite for future virtual cell models to be validated by domain experts and grounded in biology. ### Competing Interest Statement F.J.T. consults for Immunai, CytoReason, Valinor Industries, Bioturing and Phylo Inc., and has ownership interest in RN.AI Therapeutics, Dermagnostix, and Cellarity. The remaining authors declare no competing interests.

The evolving concept of cell identity in the single cell era
Summary: This Spotlight explores emerging technologies that are enabling the systematic and unbiased quantification of cell identity, and how these efforts will enable the construction of high-resolution, dynamic cell atlases.

Embryo-scale reverse genetics at single-cell resolution
The maturation of single-cell transcriptomic technologies has facilitated the generation of comprehensive cellular atlases from whole embryos1–4. A majority of these data, however, has been collected from wild-type embryos without an appreciation for the latent variation that is present in development. Here we present the ‘zebrafish single-cell atlas of perturbed embryos’: single-cell transcriptomic data from 1,812 individually resolved developing zebrafish embryos, encompassing 19 timepoints, 23 genetic perturbations and a total of 3.2 million cells. The high degree of replication in our study (eight or more embryos per condition) enables us to estimate the variance in cell type abundance organism-wide and to detect perturbation-dependent deviance in cell type composition relative to wild-type embryos. Our approach is sensitive to rare cell types, resolving developmental trajectories and genetic dependencies in the cranial ganglia neurons, a cell population that comprises less than 1% of the embryo. Additionally, time-series profiling of individual mutants identified a group of brachyury-independent cells with strikingly similar transcriptomes to notochord sheath cells, leading to new hypotheses about early origins of the skull. We anticipate that standardized collection of high-resolution, organism-scale single-cell data from large numbers of individual embryos will enable mapping of the genetic dependencies of zebrafish cell types, while also addressing longstanding challenges in developmental genetics, including the cellular and transcriptional plasticity underlying phenotypic diversity across individuals.

A unified network systems approach uncovers a core program underlying T follicular helper cell differentiation
Characterizing multi-scale processes underlying immune-state progression is critical for defining their function. This becomes pertinent for functionally diverse and plastic immune cells, such as T follicular helper (Tfh) cells. Here, we adopt a multi-scale network-systems approach that incorporates both regulatory and protein-protein interactions. This approach integrates diverse data types, captures regulation across levels of immune system organization, and recapitulates known Tfh differentiation drivers. Further, we present CoreNet, a core Tfh gene set that is conserved between humans and mice, across tissue types and disease contexts, and is consistent across data modalities. Using CoreNet, we implicate NR3C1 and interleukin (IL)-12 in the regulation of Tfh differentiation. Notably, IL-12 is permissive for differentiation of Tfh precursors but blocks differentiation into germinal center Tfh cells. Overall, this work elucidates networks with unexplored roles governing Tfh differentiation across species and tissues, while providing a generalizable framework. CoreNet is accessible through an interactive web server: https://pitt-csi.shinyapps.io/tfhcorenet/.
Mathematical Discovery of Potential Therapeutic Targets: Application to Rare Melanomas
Patients with rare types of melanoma such as acral, mucosal, or uveal melanoma, have lower survival rates than patients with cutaneous melanoma; these lower survival rates reflect the lower objective response rates to immunotherapy compared to cutaneous melanoma. Understanding tumor-immune dynamics in rare melanomas is critical for the development of new therapies and for improving response rates to current cancer therapies. Progress has been hindered by the lack of clinical data and the need for better preclinical models of rare melanomas. Canine melanoma provides a valuable comparative oncology model for rare types of human melanomas. We analyzed RNA sequencing data from canine melanoma patients and combined this with literature information to create a novel mechanistic mathematical model of melanoma-immune dynamics. Sensitivity analysis of the mathematical model indicated influential pathways in the dynamics, providing support for potential new therapeutic targets and future combinations of therapies. We share our learnings from this work, to help enable the application of this proof-of-concept workflow to other rare disease settings with sparse available data.

Gene regulatory networks: from correlative models to causal explanations
Gene regulatory networks (GRNs) explain how the genome controls cellular behaviour and tissue morphogenesis, serving to connect molecular mechanism to functional output. Single-cell technologies now provide descriptions of these networks with unprecedented detail, but this advance has also revealed gene regulatory systems that are too complex for our existing conceptual frameworks. GRNs, which should provide mechanistic explanations, are increasingly reduced to statistical correlations — ‘hairballs’ that fail to capture molecular causation. Here, we explore why this dilemma exists and propose a path forward. We argue that methods in ‘representation learning’ can be used to model GRNs, without needing to capture every molecular detail. For this framework, we advocate three linked principles: models must be inherently mechanistic, with structures grounded in cellular and evolutionary biology; molecular principles and constraints must be used to reduce the solution space for learning GRN models; and more sophisticated forms of experimental perturbation and synthetic biological engineering are needed to train models and test predictions. By reimagining GRNs through these principles, we can bridge the gap from data abundance to new conceptual understanding.

Dynamical independence: Discovering emergent macroscopic processes in complex dynamical systems
We introduce a notion of emergence for macroscopic variables associated with highly multivariate microscopic dynamical processes. Dynamical independence instantiates the intuition of an emergent macroscopic process as one possessing the characteristics of a dynamical system ``in its own right,'' with its own dynamical laws distinct from those of the underlying microscopic dynamics. We quantify (departure from) dynamical independence by a transformation-invariant Shannon information-based measure of dynamical dependence. We emphasize the data-driven discovery of dynamically independent macroscopic variables, and introduce the idea of a multiscale ``emergence portrait'' for complex systems. We show how dynamical dependence may be computed explicitly for linear systems in both time and frequency domains, facilitating discovery of emergent phenomena across spatiotemporal scales, and outline application of the linear operationalization to inference of emergence portraits for neural systems from neurophysiological time-series data. We discuss dynamical independence for discrete- and continuous-time deterministic dynamics, with potential application to Hamiltonian mechanics and classical complex systems such as flocking and cellular automata.
A multimodal perturbation atlas defines the phenotypic resolution of cellular morphology
Because cells are complex dynamical systems, modeling cellular behaviors requires methods that capture how cells evolve across time, environments, and interventions. Microscopy is uniquely suited to this goal in that it can be applied to living cells in their native context. However, the phenotypic resolving power of live-cell microscopy remains incompletely characterized, particularly relative to molecular assays. Here, we present a multimodal perturbation atlas of 1,000 pooled CRISPR knockouts in A549 cells, profiled by fluorescence microscopy (39 live, 13 fixed markers), label-free phase imaging of the same live cells, and single-cell RNA sequencing (scRNA-seq). Totaling ∼57 million single-cell profiles, our data yield rich cell-biological signatures that map individual gene function. We find that phase imaging matches — and, with sufficient cell coverage, exceeds — the phenotypic resolution of fluorescence imaging and scRNA-seq, while capturing higher-order pathway organization that scRNA-seq does not resolve. These results establish intrinsic morphology as a high-precision readout of cellular state, and lay a foundation for live-cell profiling of phenotypic trajectories. ### Competing Interest Statement The authors have declared no competing interest. Biohub, Redwoood City, CA, USA

Rectangle: robust and scalable multiscale deconvolution informed by single-cell RNA sequencing data
Bulk RNA-seq enables effective profiling of large cohorts and complex experimental designs, but current single-cell-informed deconvolution methods incompletely resolve closely related cell phenotypes, do not scale efficiently to large single-cell datasets, or fail to account for cellular content not represented in the reference. Here, we present Rectangle, an scverse Python framework for single-cell-informed deconvolution of bulk RNA-seq data.
On cell types and cell states
The advent of single-cell genomics has brought about new efforts to characterize and catalog all of the cell types in the human body. Despite these efforts, the very definition of a “cell type” is under debate. In this post, I will discuss a conceptual framework for defining cell types as subsets of states in an underlying cellular state space. Moreover, I will link the cellular state space to biomedical ontologies that attempt to capture biological knowledge regarding cell types.
Integrated computational and in vivo models reveal Key Insights into macrophage behavior during bone healing
Myeloid-derived monocyte and macrophages are key cells in the bone that contribute to remodeling and injury repair. However, their temporal polarization status and control of bone-resorbing osteoclasts and bone-forming osteoblasts responses is largely unknown. In this study, we focused on two aspects of monocyte/macrophage dynamics and polarization states over time: 1) the injury-triggered pro- and anti-inflammatory monocytes/macrophages temporal profiles, 2) the contributions of pro- versus anti-inflammatory monocytes/macrophages in coordinating healing response. Bone healing is a complex multicellular dynamic process. While traditional in vitro and in vivo experimentation may capture the behavior of select populations with high resolution, they cannot simultaneously track the behavior of multiple populations. To address this, we have used an integrated coupled ordinary differential equations (ODEs)-based framework describing multiple cellular species to in vivo bone injury data in order to identify and test various hypotheses regarding bone cell populations dynamics. Our approach allowed us to infer several biological insights including, but not limited to,: 1) anti-inflammatory macrophages are key for early osteoclast inhibition and pro-inflammatory macrophage suppression, 2) pro-inflammatory macrophages are involved in osteoclast bone resorptive activity, whereas osteoblasts promote osteoclast differentiation, 3) Pro-inflammatory monocytes/macrophages rise during two expansion waves, which can be explained by the anti-inflammatory macrophages-mediated inhibition phase between the two waves. In addition, we further tested the robustness of the mathematical model by comparing simulation results to an independent experimental dataset. Taken together, this novel comprehensive mathematical framework allowed us to identify biological mechanisms that best recapitulate bone injury data and that explain the coupled cellular population dynamics involved in the process. Furthermore, our hypothesis testing methodology could be used in other contexts to decipher mechanisms in complex multicellular processes.
Are biological systems poised at criticality?
Many of life's most fascinating phenomena emerge from interactions among many elements--many amino acids determine the structure of a single protein, many genes determine the fate of a cell, many neurons are involved in shaping our thoughts and memories. Physicists have long hoped that these collective behaviors could be described using the ideas and methods of statistical mechanics. In the past few years, new, larger scale experiments have made it possible to construct statistical mechanics models of biological systems directly from real data. We review the surprising successes of this "inverse" approach, using examples form families of proteins, networks of neurons, and flocks of birds. Remarkably, in all these cases the models that emerge from the data are poised at a very special point in their parameter space--a critical point. This suggests there may be some deeper theoretical principle behind the behavior of these diverse systems.

Hierarchical classification of immune cell transcriptomes at population-scale
Accurate immune cell classification is essential for interpreting single-cell RNA sequencing (scRNA-seq) data. However, progress is constrained by the lack of independent, high-resolution benchmarks, as the routine integration of datasets introduces statistical dependencies that artificially inflate model generalizability. Here, we present the single-cell universal classification omnibus (Suco), a resource of independent, uniform expert annotations, and Compocyte, a modular hierarchical classifier. Together, they establish a framework designed for the scale of human population immunology. This approach substantially outperforms existing classifiers while facilitating expert review of ambiguous annotations. Applying Compocyte across 50 studies, including three newly generated datasets, we classified 15.6 million leukocytes from 3,965 patients. Within this expansive cohort, we identified a new tumor-associated resorptive macrophage phenotype, a non-canonical monocyte subtype in subclinical cytokine release syndrome, and the programmatic erosion of T cell memory stemness across metastatic sites. Suco and Compocyte thus provide a generalizable architecture and benchmark capable of sustaining high-resolution annotation across massive clinical cohorts. ### Competing Interest Statement CMR has consulted regarding oncology drug development with Amgen, AstraZeneca, Daiichi Sankyo, Genentech, Merck, and Novartis, and has received licensing and royalty payments for DLL3-directed therapeutics. T.W. reports stock ownership for Roche, Astra Zeneca, Bayer, Innate Pharma, Kyntra, Illumina, 10x Genomics, and Merck KGaA as well as research funding from Atrandi Biosciences, Vilnius, Lithuania; CanVirex AG, Basel Switzerland; and Institut fuer Klinische Krebsforschung GmbH, Frankfurt, Germany, and travel funding from Roche, Basel, Switzerland. S.Z. reports advisory board membership and honoraria from Amgen, Astellas, AstraZeneca, Bayer, Bristol-Myers Squibb, Daiichi Sankyo, Eisai, EUSA, Gilead, Ipsen, Johnson&Johnson, Lilly, MedSir, Medtoday, Merck, MSD, Novartis, Pfizer, Roche, Sanofi Aventis, StreamedUp, Urotrials, Urotube, Zentiva and resarch funding from Eisai. S.Z. reports clinical trial support from Amgen, AstraZeneca, AVEO, Bayer, Biontech, Bristol-Myers Squibb, Calithera, Exelixis, Gilead, Lilly, MSD, Novartis, Pfizer, Roche, Seagen/Astellas, Urotrials and travels & conference support from Amgen, Astellas, AstraZeneca, Bayer, EISAI, Ipsen, Johnson&Johnson, Merck, MSD, Pfizer. All remaining authors declare no relevant competing interests. Spanish Association Against Cancer, PI049999 Federal Ministry of Research, Technology and Space, 001001KT2322 National Cancer Institute, R35 CA263816 National Cancer Institute, U24 CA213274 National Cancer Institute, P30 CA008748 Research Council of Lithuania, P-MIP-24-93
