







The maturation of single-cell transcriptomic technologies has facilitated the generation of comprehensive cellular atlases from whole embryos1–4. A majority of these data, however, has been collected from wild-type embryos without an appreciation for the latent variation that is present in development. Here we present the ‘zebrafish single-cell atlas of perturbed embryos’: single-cell transcriptomic data from 1,812 individually resolved developing zebrafish embryos, encompassing 19 timepoints, 23 genetic perturbations and a total of 3.2 million cells. The high degree of replication in our study (eight or more embryos per condition) enables us to estimate the variance in cell type abundance organism-wide and to detect perturbation-dependent deviance in cell type composition relative to wild-type embryos. Our approach is sensitive to rare cell types, resolving developmental trajectories and genetic dependencies in the cranial ganglia neurons, a cell population that comprises less than 1% of the embryo. Additionally, time-series profiling of individual mutants identified a group of brachyury-independent cells with strikingly similar transcriptomes to notochord sheath cells, leading to new hypotheses about early origins of the skull. We anticipate that standardized collection of high-resolution, organism-scale single-cell data from large numbers of individual embryos will enable mapping of the genetic dependencies of zebrafish cell types, while also addressing longstanding challenges in developmental genetics, including the cellular and transcriptional plasticity underlying phenotypic diversity across individuals.
Quantifying cell-state densities in single-cell phenotypic landscapes using Mellon
Cell-state density characterizes the distribution of cells along phenotypic landscapes and is crucial for unraveling the mechanisms that drive diverse biological processes. Here, we present Mellon, an algorithm for estimation of cell-state densities from high-dimensional representations of single-cell data. We demonstrate Mellon’s efficacy by dissecting the density landscape of differentiating systems, revealing a consistent pattern of high-density regions corresponding to major cell types intertwined with low-density, rare transitory states. We present evidence implicating enhancer priming and the activation of master regulators in emergence of these transitory states. Mellon offers the flexibility to perform temporal interpolation of time-series data, providing a detailed view of cell-state dynamics during developmental processes. Mellon facilitates density estimation across various single-cell data modalities, scaling linearly with the number of cells. Our work underscores the importance of cell-state density in understanding the differentiation processes, and the potential of Mellon to provide insights into mechanisms guiding biological trajectories.

Rectangle: robust and scalable multiscale deconvolution informed by single-cell RNA sequencing data
Bulk RNA-seq enables effective profiling of large cohorts and complex experimental designs, but current single-cell-informed deconvolution methods incompletely resolve closely related cell phenotypes, do not scale efficiently to large single-cell datasets, or fail to account for cellular content not represented in the reference. Here, we present Rectangle, an scverse Python framework for single-cell-informed deconvolution of bulk RNA-seq data.
Benchmarking gene expression reconstruction from single-cell latent representations
Single-cell transcriptomics is typically modeled in low-dimensional latent representations that improve the signal-to-noise ratio of the data. Such representations underpin data integration, cell state discovery, and perturbation prediction, with applications ranging from large-scale organ atlases to latent trajectory modeling. Recent virtual cell approaches further leverage these representations to predict cellular responses as distributional shifts in latent space. Each of these applications ultimately requires faithful gene expression reconstruction from latent spaces for biological interpretation, enabling gene-level analysis of predicted perturbed or batch-corrected cells. Yet representation choice is typically treated as an implementation detail rather than a primary modeling decision, with no systematic evaluation of how well latent representations support gene expression reconstruction. Here, we introduce ReconEval, a benchmark for evaluating gene expression reconstruction from single-cell latent spaces. We benchmark two classes of latent representations: end-to-end trained models such as PCA, autoencoders, and variational autoencoders, and pretrained single-cell foundation model embeddings coupled to newly trained decoders. Reconstruction is evaluated both directly and after latent-space perturbation prediction. Across perturbational and observational datasets totaling over 100 million cells, our metric suite quantifies statistical fidelity; biological signal preservation, including differential expression, coexpression, cell-cycle structure, cytokine response and pathway activity; and perturbation-specific effects. We find that autoencoders achieve the strongest stand-alone reconstruction at low dimensionality, while variational regularization does not improve generalization in reconstruction. Frozen foundation model embeddings retain recoverable gene-level information, with reconstruction quality depending strongly on decoder architecture and pretraining objective. In latent perturbation modeling, high-dimensional PCA matches foundation model embeddings, while low-dimensional AE embeddings are optimal for flow-based generative models. Overall, reconstruction depends critically on the interplay between representation and downstream model, and simpler representations can outperform complex alternatives given appropriate capacity. Our benchmark establishes reconstruction as a critical evaluation axis for single-cell foundation models. We envision it improving the biological interpretability of latent-space modeling, a prerequisite for future virtual cell models to be validated by domain experts and grounded in biology. ### Competing Interest Statement F.J.T. consults for Immunai, CytoReason, Valinor Industries, Bioturing and Phylo Inc., and has ownership interest in RN.AI Therapeutics, Dermagnostix, and Cellarity. The remaining authors declare no competing interests.

Realistic in silico generation and augmentation of single cell RNA-seq data using Generative Adversarial Neural Networks
A fundamental problem in biomedical research is the low number of observations available, mostly due to a lack of available biosamples, prohibitive costs, or ethical reasons. Augmenting few real observations with generated in silico samples could lead to more robust analysis results and a higher reproducibility rate. Here we propose the use of conditional single cell Generative Adversarial Neural Networks (cscGANs) for the realistic generation of single cell RNA-seq data. cscGANs learn non-linear gene-gene dependencies from complex, multi cell type samples and use this information to generate realistic cells of defined types. Augmenting sparse cell populations with cscGAN generated cells improves downstream analyses such as the detection of marker genes, the robustness and reliability of classifiers, the assessment of novel analysis algorithms, and might reduce the number of animal experiments and costs in consequence. cscGANs outperform existing methods for single cell RNA-seq data generation in quality and hold great promise for the realistic generation and augmentation of other biomedical data types. *

Rapid generation of maternal mutants via oocyte transgenic expression of CRISPR-Cas9 and sgRNAs in zebrafish
A time-saving and deletion-prone method facilitates functional study of maternal factors in zebrafish.

Comparing phenotypic manifolds with Kompot: Detecting differential abundance and gene expression at single-cell resolution
Single-cell studies are frequently designed to compare across conditions such as health and disease. However, existing computational approaches typically rely on grouping cells into discrete populations before making comparisons, which can limit resolution for detecting state-dependent changes. Here, we introduce Kompot, a statistical framework for comparative analysis of multi-condition single-cell data. Kompot quantifies both differential abundance, capturing how cells redistribute across the phenotypic space, and differential expression, identifying condition-specific transcriptional changes that may be localized, heterogeneous, or oppositely regulated across states. By modeling cell density and gene expression as continuous functions over a shared cell-state representation, Kompot enables single-cell–resolution inference with principled uncertainty estimates, without requiring predefined clusters or cell types. Applying Kompot to aging murine bone marrow, we identified a continuum of shifts in hematopoietic stem cell and mature cell states, transcriptional remodeling of monocytes independent of compositional changes, and divergent regulation of oxidative stress response genes across cell types. We demonstrate the utility of Kompot in disease settings by identifying cell-state and gene expression changes associated with improved efficacy of combinatorial immunotherapy in melanoma. Additionally, Kompot enables multi-sample comparative analysis by accounting for sample-to-sample heterogeneity. By capturing both global and cell-state–specific effects of perturbation, the Kompot framework is broadly applicable to dissecting condition-specific effects in complex single-cell landscapes. ### Competing Interest Statement The authors have declared no competing interest. National Institutes of Health, R35GM147125, R01CA292932, T32GM136534, S10OD028685 The Mark Foundation for Cancer Research, https://ror.org/00v7th354, Endeavor Award Edward P. Evans Foundation, https://ror.org/03h22gm35, Discovery Research Grant Brotman Baty Institute, https://ror.org/03jxvbk42, Pilot Award

Realistic in silico generation and augmentation of single-cell RNA-seq data using generative adversarial networks
A fundamental problem in biomedical research is the low number of observations available, mostly due to a lack of available biosamples, prohibitive costs, or ethical reasons. Augmenting few real observations with generated in silico samples could lead to more robust analysis results and a higher reproducibility rate. Here, we propose the use of conditional single-cell generative adversarial neural networks (cscGAN) for the realistic generation of single-cell RNA-seq data. cscGAN learns non-linear gene–gene dependencies from complex, multiple cell type samples and uses this information to generate realistic cells of defined types. Augmenting sparse cell populations with cscGAN generated cells improves downstream analyses such as the detection of marker genes, the robustness and reliability of classifiers, the assessment of novel analysis algorithms, and might reduce the number of animal experiments and costs in consequence. cscGAN outperforms existing methods for single-cell RNA-seq data generation in quality and hold great promise for the realistic generation and augmentation of other biomedical data types.

Single-cell multiregion dissection of Alzheimer’s disease
Alzheimer’s disease is the leading cause of dementia worldwide, but the cellular pathways that underlie its pathological progression across brain regions remain poorly understood1–3. Here we report a single-cell transcriptomic atlas of six different brain regions in the aged human brain, covering 1.3 million cells from 283 post-mortem human brain samples across 48 individuals with and without Alzheimer’s disease. We identify 76 cell types, including region-specific subtypes of astrocytes and excitatory neurons and an inhibitory interneuron population unique to the thalamus and distinct from canonical inhibitory subclasses. We identify vulnerable populations of excitatory and inhibitory neurons that are depleted in specific brain regions in Alzheimer’s disease, and provide evidence that the Reelin signalling pathway is involved in modulating the vulnerability of these neurons. We develop a scalable method for discovering gene modules, which we use to identify cell-type-specific and region-specific modules that are altered in Alzheimer’s disease and to annotate transcriptomic differences associated with diverse pathological variables. We identify an astrocyte program that is associated with cognitive resilience to Alzheimer’s disease pathology, tying choline metabolism and polyamine biosynthesis in astrocytes to preserved cognitive function late in life. Together, our study develops a regional atlas of the ageing human brain and provides insights into cellular vulnerability, response and resilience to Alzheimer’s disease pathology.

VAPOR: A variational autoencoder with transport operators to disentangle cellular gene expression dynamics of co-occurring biological processes in time and space
Abstract Single-cell and spatial transcriptomics enable the analysis of cellular states and dynamics in gene expression, revealing how diverse biological processes relate to these states over time and space. To study these dynamics, trajectory inference methods order cells along computationally inferred paths to reconstruct gradual transitions in cell states. However, by encouraging smooth and continuous trajectories, these approaches tend to conflate co-occurring processes-such as proliferation, maturation, and spatial organization-that are jointly reflected in gene expression, potentially overlooking process-specific gene expression dynamics. To address this, we developed VAPOR, which integrates a variational autoencoder with transport operators to model and disentangle cellular gene expression dynamics for potentially co-occurring biological processes. VAPOR inputs single-cell (or spatial) gene expression data into a variational autoencoder (VAE) to learn the latent states of cells and then models their latent dynamics as an ordinary differential equation. The latent dynamics are further decomposed into process-specific components parameterized by transport operators (TOs) and their corresponding process weights. Each TO defines a process-specific dynamics, and its weight for each cell quantifies the process's contribution to the cell dynamics. After assessment by simulation studies, we applied VAPOR with benchmarking to real data, including time-course scRNA-seq from postconceptual human brain development, spatial transcriptomics of the mouse hippocampus, and cross-species scRNA-seq spanning human and macaque first-trimester forebrain development. In these applications, VAPOR has identified a variety of temporal and spatial co-occurring processes, such as cell cycle, gliogenesis, neurogenesis, and neuronal migration, along with associated dynamic genes, including those species-specific to human and macaque development. VAPOR is available as an open-source tool for general-purpose use. ### Competing Interest Statement The authors have declared no competing interest.

Conserved enhancers control notochord expression of vertebrate Brachyury
The cell type-specific expression of key transcription factors is central to development and disease. Brachyury/T/TBXT is a major transcription factor for gastrulation, tailbud patterning, and notochord formation; however, how its expression is controlled in the mammalian notochord has remained elusive. Here, we identify the complement of notochord-specific enhancers in the mammalian Brachyury/T/TBXT gene. Using transgenic assays in zebrafish, axolotl, and mouse, we discover three conserved Brachyury-controlling notochord enhancers, T3, C, and I, in human, mouse, and marsupial genomes. Acting as Brachyury-responsive, auto-regulatory shadow enhancers, in cis deletion of all three enhancers in mouse abolishes Brachyury/T/Tbxt expression selectively in the notochord, causing specific trunk and neural tube defects without gastrulation or tailbud defects. The three Brachyury-driving notochord enhancers are conserved beyond mammals in the brachyury/tbxtb loci of fishes, dating their origin to the last common ancestor of jawed vertebrates. Our data define the vertebrate enhancers for Brachyury/T/TBXTB notochord expression through an auto-regulatory mechanism that conveys robustness and adaptability as ancient basis for axis development.

Targeted mutagenesis of specific genomic DNA sequences in animals for the in vivo generation of variant libraries
Understanding how the number, placement and affinity of transcription factor binding sites dictates gene regulatory programs remains a major unsolved challenge in biology, particularly in the context of multicellular organisms. To uncover these rules, it is first necessary to find the binding sites within a regulatory region with high precision, and then to systematically modulate this binding site arrangement while simultaneously measuring the effect of this modulation on output gene expression. Massively parallel reporter assays (MPRAs), where the gene expression stemming from 10,000s of in vitro-generated regulatory sequences is measured, have made this feat possible in high-throughput in single cells in culture. However, because of lack of technologies to incorporate DNA libraries, MPRAs are limited in whole organisms. To enable MPRAs in multicellular organisms, we generated tools to create a high degree of mutagenesis in specific genomic loci in vivo using base editing. Targeting GFP integrated in the genome of Drosophila cell culture and whole animals as a case study, we show that the base editor AIDevoCDA1 stemming from sea lamprey fused to nCas9 is highly mutagenic. Surprisingly, longer gRNAs increase mutation efficiency and expand the mutating window, which can allow the introduction of mutations in previously untargetable sequences. Finally, we demonstrate arrays of >20 gRNAs that can efficiently introduce mutations along a 200bp sequence, making it a promising tool to test enhancer function in vivo in a high throughput manner.

BE4max and AncBE4max Are Efficient in Germline Conversion of C:G to T:A Base Pairs in Zebrafish
The ease of use and robustness of genome editing by CRISPR/Cas9 has led to successful use of gene knockout zebrafish for disease modeling. However, it still remains a challenge to precisely edit the zebrafish genome to create single-nucleotide substitutions, which account for ~60% of human disease-causing mutations. Recently developed base editing nucleases provide an excellent alternate to CRISPR/Cas9-mediated homology dependent repair for generation of zebrafish with point mutations. A new set of cytosine base editors, termed BE4max and AncBE4max, demonstrated improved base editing efficiency in mammalian cells but have not been evaluated in zebrafish. Therefore, we undertook this study to evaluate their efficiency in converting C:G to T:A base pairs in zebrafish by somatic and germline analysis using highly active sgRNAs to twist and ntl genes. Our data demonstrated that these improved BE4max set of plasmids provide desired base substitutions at similar efficiency and without any indels compared to the previously reported BE3 and Target-AID plasmids in zebrafish. Our data also showed that AncBE4max produces fewer incorrect and bystander edits, suggesting that it can be further improved by codon optimization of its components for use in zebrafish.

An atlas-scale generative model for unified representation learning of bulk RNA-seq data
Public bulk RNA-seq repositories contain hundreds of thousands of samples, creating opportunities for large-scale representation learning, but integration across studies remains challenging because of heterogeneous annotations, experimental protocols, and technical variation. While pre-trained foundation models are now widely available for single-cell RNA-seq, comparable resources for bulk RNA-seq remain scarce, motivating a model that learns a unified, tissue-aware representation directly from bulk data. We trained a supervised variational autoencoder (VAE) on a compendium of 118,263 bulk RNA-seq samples that we assembled from TCGA, GTEx, and ARCHS4 and mapped to 42 tissue categories. The model classifies tissue of origin at 94.9% balanced accuracy (weighted F1 96.2%) and compresses 16,115 genes into a 121-dimensional latent space. Tissue identity is the primary organizing axis of the latent space, while source effects remain secondary. To assess the impact of data volume, we constructed training sets at three different scales (38K, 75K, and 118K samples). Our results demonstrated that reconstruction fidelity improved incrementally with each expansion of the dataset, but with diminishing returns. We validated the model on an independent cohort of 734 paediatric tumour samples from TARGET, achieving 84.6% agreement with the expected tissue of origin. The trained model and code are available at GitHub ([https://github.com/BIMSBbioinfo/flexynesis\_tissue\_vae_manuscript][1]) with an interactive web application. ### Competing Interest Statement The authors have declared no competing interest. [1]: https://github.com/BIMSBbioinfo/flexynesis_tissue_vae_manuscript

Cytosine base editors with increased PAM and deaminase motif flexibility for gene editing in zebrafish
Cytosine base editing is a powerful tool for making precise single nucleotide changes in cells and model organisms like zebrafish, which are valuable for studying human diseases. However, current base editors struggle to edit cytosines in certain DNA contexts, particularly those with GC and CC pairs, limiting their use in modelling disease-related mutations. Here we show the development of zevoCDA1, an optimized cytosine base editor for zebrafish that improves editing efficiency across various DNA contexts and reduces restrictions imposed by the protospacer adjacent motif. We also create zevoCDA1-198, a more precise editor with a narrower editing window of five nucleotides, minimizing off-target effects. Using these advanced tools, we successfully generate zebrafish models of diseases that were previously challenging to create due to sequence limitations. This work enhances the ability to introduce human pathogenic mutations in zebrafish, broadening the scope for genomic research with improved precision and efficiency.

On cell types and cell states
The advent of single-cell genomics has brought about new efforts to characterize and catalog all of the cell types in the human body. Despite these efforts, the very definition of a “cell type” is under debate. In this post, I will discuss a conceptual framework for defining cell types as subsets of states in an underlying cellular state space. Moreover, I will link the cellular state space to biomedical ontologies that attempt to capture biological knowledge regarding cell types.
A multimodal perturbation atlas defines the phenotypic resolution of cellular morphology
Because cells are complex dynamical systems, modeling cellular behaviors requires methods that capture how cells evolve across time, environments, and interventions. Microscopy is uniquely suited to this goal in that it can be applied to living cells in their native context. However, the phenotypic resolving power of live-cell microscopy remains incompletely characterized, particularly relative to molecular assays. Here, we present a multimodal perturbation atlas of 1,000 pooled CRISPR knockouts in A549 cells, profiled by fluorescence microscopy (39 live, 13 fixed markers), label-free phase imaging of the same live cells, and single-cell RNA sequencing (scRNA-seq). Totaling ∼57 million single-cell profiles, our data yield rich cell-biological signatures that map individual gene function. We find that phase imaging matches — and, with sufficient cell coverage, exceeds — the phenotypic resolution of fluorescence imaging and scRNA-seq, while capturing higher-order pathway organization that scRNA-seq does not resolve. These results establish intrinsic morphology as a high-precision readout of cellular state, and lay a foundation for live-cell profiling of phenotypic trajectories. ### Competing Interest Statement The authors have declared no competing interest. Biohub, Redwoood City, CA, USA
