







Gene regulatory networks (GRNs) explain how the genome controls cellular behaviour and tissue morphogenesis, serving to connect molecular mechanism to functional output. Single-cell technologies now provide descriptions of these networks with unprecedented detail, but this advance has also revealed gene regulatory systems that are too complex for our existing conceptual frameworks. GRNs, which should provide mechanistic explanations, are increasingly reduced to statistical correlations — ‘hairballs’ that fail to capture molecular causation. Here, we explore why this dilemma exists and propose a path forward. We argue that methods in ‘representation learning’ can be used to model GRNs, without needing to capture every molecular detail. For this framework, we advocate three linked principles: models must be inherently mechanistic, with structures grounded in cellular and evolutionary biology; molecular principles and constraints must be used to reduce the solution space for learning GRN models; and more sophisticated forms of experimental perturbation and synthetic biological engineering are needed to train models and test predictions. By reimagining GRNs through these principles, we can bridge the gap from data abundance to new conceptual understanding.
VAPOR: A variational autoencoder with transport operators to disentangle cellular gene expression dynamics of co-occurring biological processes in time and space
Abstract Single-cell and spatial transcriptomics enable the analysis of cellular states and dynamics in gene expression, revealing how diverse biological processes relate to these states over time and space. To study these dynamics, trajectory inference methods order cells along computationally inferred paths to reconstruct gradual transitions in cell states. However, by encouraging smooth and continuous trajectories, these approaches tend to conflate co-occurring processes-such as proliferation, maturation, and spatial organization-that are jointly reflected in gene expression, potentially overlooking process-specific gene expression dynamics. To address this, we developed VAPOR, which integrates a variational autoencoder with transport operators to model and disentangle cellular gene expression dynamics for potentially co-occurring biological processes. VAPOR inputs single-cell (or spatial) gene expression data into a variational autoencoder (VAE) to learn the latent states of cells and then models their latent dynamics as an ordinary differential equation. The latent dynamics are further decomposed into process-specific components parameterized by transport operators (TOs) and their corresponding process weights. Each TO defines a process-specific dynamics, and its weight for each cell quantifies the process's contribution to the cell dynamics. After assessment by simulation studies, we applied VAPOR with benchmarking to real data, including time-course scRNA-seq from postconceptual human brain development, spatial transcriptomics of the mouse hippocampus, and cross-species scRNA-seq spanning human and macaque first-trimester forebrain development. In these applications, VAPOR has identified a variety of temporal and spatial co-occurring processes, such as cell cycle, gliogenesis, neurogenesis, and neuronal migration, along with associated dynamic genes, including those species-specific to human and macaque development. VAPOR is available as an open-source tool for general-purpose use. ### Competing Interest Statement The authors have declared no competing interest.

Benchmarking gene expression reconstruction from single-cell latent representations
Single-cell transcriptomics is typically modeled in low-dimensional latent representations that improve the signal-to-noise ratio of the data. Such representations underpin data integration, cell state discovery, and perturbation prediction, with applications ranging from large-scale organ atlases to latent trajectory modeling. Recent virtual cell approaches further leverage these representations to predict cellular responses as distributional shifts in latent space. Each of these applications ultimately requires faithful gene expression reconstruction from latent spaces for biological interpretation, enabling gene-level analysis of predicted perturbed or batch-corrected cells. Yet representation choice is typically treated as an implementation detail rather than a primary modeling decision, with no systematic evaluation of how well latent representations support gene expression reconstruction. Here, we introduce ReconEval, a benchmark for evaluating gene expression reconstruction from single-cell latent spaces. We benchmark two classes of latent representations: end-to-end trained models such as PCA, autoencoders, and variational autoencoders, and pretrained single-cell foundation model embeddings coupled to newly trained decoders. Reconstruction is evaluated both directly and after latent-space perturbation prediction. Across perturbational and observational datasets totaling over 100 million cells, our metric suite quantifies statistical fidelity; biological signal preservation, including differential expression, coexpression, cell-cycle structure, cytokine response and pathway activity; and perturbation-specific effects. We find that autoencoders achieve the strongest stand-alone reconstruction at low dimensionality, while variational regularization does not improve generalization in reconstruction. Frozen foundation model embeddings retain recoverable gene-level information, with reconstruction quality depending strongly on decoder architecture and pretraining objective. In latent perturbation modeling, high-dimensional PCA matches foundation model embeddings, while low-dimensional AE embeddings are optimal for flow-based generative models. Overall, reconstruction depends critically on the interplay between representation and downstream model, and simpler representations can outperform complex alternatives given appropriate capacity. Our benchmark establishes reconstruction as a critical evaluation axis for single-cell foundation models. We envision it improving the biological interpretability of latent-space modeling, a prerequisite for future virtual cell models to be validated by domain experts and grounded in biology. ### Competing Interest Statement F.J.T. consults for Immunai, CytoReason, Valinor Industries, Bioturing and Phylo Inc., and has ownership interest in RN.AI Therapeutics, Dermagnostix, and Cellarity. The remaining authors declare no competing interests.

Targeted mutagenesis of specific genomic DNA sequences in animals for the in vivo generation of variant libraries
Understanding how the number, placement and affinity of transcription factor binding sites dictates gene regulatory programs remains a major unsolved challenge in biology, particularly in the context of multicellular organisms. To uncover these rules, it is first necessary to find the binding sites within a regulatory region with high precision, and then to systematically modulate this binding site arrangement while simultaneously measuring the effect of this modulation on output gene expression. Massively parallel reporter assays (MPRAs), where the gene expression stemming from 10,000s of in vitro-generated regulatory sequences is measured, have made this feat possible in high-throughput in single cells in culture. However, because of lack of technologies to incorporate DNA libraries, MPRAs are limited in whole organisms. To enable MPRAs in multicellular organisms, we generated tools to create a high degree of mutagenesis in specific genomic loci in vivo using base editing. Targeting GFP integrated in the genome of Drosophila cell culture and whole animals as a case study, we show that the base editor AIDevoCDA1 stemming from sea lamprey fused to nCas9 is highly mutagenic. Surprisingly, longer gRNAs increase mutation efficiency and expand the mutating window, which can allow the introduction of mutations in previously untargetable sequences. Finally, we demonstrate arrays of >20 gRNAs that can efficiently introduce mutations along a 200bp sequence, making it a promising tool to test enhancer function in vivo in a high throughput manner.

A unified network systems approach uncovers a core program underlying T follicular helper cell differentiation
Characterizing multi-scale processes underlying immune-state progression is critical for defining their function. This becomes pertinent for functionally diverse and plastic immune cells, such as T follicular helper (Tfh) cells. Here, we adopt a multi-scale network-systems approach that incorporates both regulatory and protein-protein interactions. This approach integrates diverse data types, captures regulation across levels of immune system organization, and recapitulates known Tfh differentiation drivers. Further, we present CoreNet, a core Tfh gene set that is conserved between humans and mice, across tissue types and disease contexts, and is consistent across data modalities. Using CoreNet, we implicate NR3C1 and interleukin (IL)-12 in the regulation of Tfh differentiation. Notably, IL-12 is permissive for differentiation of Tfh precursors but blocks differentiation into germinal center Tfh cells. Overall, this work elucidates networks with unexplored roles governing Tfh differentiation across species and tissues, while providing a generalizable framework. CoreNet is accessible through an interactive web server: https://pitt-csi.shinyapps.io/tfhcorenet/.
Higher-Order Knowledge Representations for Agentic Scientific Reasoning
Scientific inquiry requires systems-level reasoning that integrates heterogeneous experimental data, cross-domain knowledge, and mechanistic evidence into coherent explanations. While Large Language Models (LLMs) offer inferential capabilities, they often depend on retrieval-augmented contexts that lack structural depth. Traditional Knowledge Graphs (KGs) attempt to bridge this gap, yet their pairwise constraints fail to capture the irreducible higher-order interactions that govern emergent physical behavior. To address this, we introduce a methodology for constructing hypergraph-based knowledge representations that faithfully encode multi-entity relationships. Applied to a corpus of ≈\approx 1,100 manuscripts on biocomposite scaffolds, our framework constructs a global hypergraph of 161,172 nodes and 320,201 hyperedges, revealing a scale-free topology (power law exponent ≈\approx 1.23) organized around highly connected conceptual hubs. This representation prevents the combinatorial explosion typical of pairwise expansions and explicitly preserves the co-occurrence context of scientific formulations. We further demonstrate that equipping agentic systems with hypergraph traversal tools, specifically using node-intersection constraints, enables them to bridge semantically distant concepts. By exploiting these higher-order pathways, the system successfully generates grounded mechanistic hypotheses for novel composite materials, such as linking cerium oxide to PCL scaffolds via chitosan intermediates. This work establishes a “teacherless” agentic reasoning system where hypergraph topology acts as a verifiable guardrail, accelerating scientific discovery by uncovering relationships obscured by traditional graph methods.
An atlas-scale generative model for unified representation learning of bulk RNA-seq data
Public bulk RNA-seq repositories contain hundreds of thousands of samples, creating opportunities for large-scale representation learning, but integration across studies remains challenging because of heterogeneous annotations, experimental protocols, and technical variation. While pre-trained foundation models are now widely available for single-cell RNA-seq, comparable resources for bulk RNA-seq remain scarce, motivating a model that learns a unified, tissue-aware representation directly from bulk data. We trained a supervised variational autoencoder (VAE) on a compendium of 118,263 bulk RNA-seq samples that we assembled from TCGA, GTEx, and ARCHS4 and mapped to 42 tissue categories. The model classifies tissue of origin at 94.9% balanced accuracy (weighted F1 96.2%) and compresses 16,115 genes into a 121-dimensional latent space. Tissue identity is the primary organizing axis of the latent space, while source effects remain secondary. To assess the impact of data volume, we constructed training sets at three different scales (38K, 75K, and 118K samples). Our results demonstrated that reconstruction fidelity improved incrementally with each expansion of the dataset, but with diminishing returns. We validated the model on an independent cohort of 734 paediatric tumour samples from TARGET, achieving 84.6% agreement with the expected tissue of origin. The trained model and code are available at GitHub ([https://github.com/BIMSBbioinfo/flexynesis\_tissue\_vae_manuscript][1]) with an interactive web application. ### Competing Interest Statement The authors have declared no competing interest. [1]: https://github.com/BIMSBbioinfo/flexynesis_tissue_vae_manuscript

Are biological systems poised at criticality?
Many of life's most fascinating phenomena emerge from interactions among many elements--many amino acids determine the structure of a single protein, many genes determine the fate of a cell, many neurons are involved in shaping our thoughts and memories. Physicists have long hoped that these collective behaviors could be described using the ideas and methods of statistical mechanics. In the past few years, new, larger scale experiments have made it possible to construct statistical mechanics models of biological systems directly from real data. We review the surprising successes of this "inverse" approach, using examples form families of proteins, networks of neurons, and flocks of birds. Remarkably, in all these cases the models that emerge from the data are poised at a very special point in their parameter space--a critical point. This suggests there may be some deeper theoretical principle behind the behavior of these diverse systems.

Quantifying cell-state densities in single-cell phenotypic landscapes using Mellon
Cell-state density characterizes the distribution of cells along phenotypic landscapes and is crucial for unraveling the mechanisms that drive diverse biological processes. Here, we present Mellon, an algorithm for estimation of cell-state densities from high-dimensional representations of single-cell data. We demonstrate Mellon’s efficacy by dissecting the density landscape of differentiating systems, revealing a consistent pattern of high-density regions corresponding to major cell types intertwined with low-density, rare transitory states. We present evidence implicating enhancer priming and the activation of master regulators in emergence of these transitory states. Mellon offers the flexibility to perform temporal interpolation of time-series data, providing a detailed view of cell-state dynamics during developmental processes. Mellon facilitates density estimation across various single-cell data modalities, scaling linearly with the number of cells. Our work underscores the importance of cell-state density in understanding the differentiation processes, and the potential of Mellon to provide insights into mechanisms guiding biological trajectories.

Comparing phenotypic manifolds with Kompot: Detecting differential abundance and gene expression at single-cell resolution
Single-cell studies are frequently designed to compare across conditions such as health and disease. However, existing computational approaches typically rely on grouping cells into discrete populations before making comparisons, which can limit resolution for detecting state-dependent changes. Here, we introduce Kompot, a statistical framework for comparative analysis of multi-condition single-cell data. Kompot quantifies both differential abundance, capturing how cells redistribute across the phenotypic space, and differential expression, identifying condition-specific transcriptional changes that may be localized, heterogeneous, or oppositely regulated across states. By modeling cell density and gene expression as continuous functions over a shared cell-state representation, Kompot enables single-cell–resolution inference with principled uncertainty estimates, without requiring predefined clusters or cell types. Applying Kompot to aging murine bone marrow, we identified a continuum of shifts in hematopoietic stem cell and mature cell states, transcriptional remodeling of monocytes independent of compositional changes, and divergent regulation of oxidative stress response genes across cell types. We demonstrate the utility of Kompot in disease settings by identifying cell-state and gene expression changes associated with improved efficacy of combinatorial immunotherapy in melanoma. Additionally, Kompot enables multi-sample comparative analysis by accounting for sample-to-sample heterogeneity. By capturing both global and cell-state–specific effects of perturbation, the Kompot framework is broadly applicable to dissecting condition-specific effects in complex single-cell landscapes. ### Competing Interest Statement The authors have declared no competing interest. National Institutes of Health, R35GM147125, R01CA292932, T32GM136534, S10OD028685 The Mark Foundation for Cancer Research, https://ror.org/00v7th354, Endeavor Award Edward P. Evans Foundation, https://ror.org/03h22gm35, Discovery Research Grant Brotman Baty Institute, https://ror.org/03jxvbk42, Pilot Award

Designing RNA sequencing experiments: A practical guide to reproducible gene expression analysis
RNA sequencing (RNA-seq) has become a cornerstone of modern biotechnology, offering a comprehensive and high-resolution view of gene expression that enables the discovery of novel transcripts across diverse biological systems. Its applications extend beyond basic transcriptomics, providing powerful tools for uncovering molecular mechanisms underlying disease, environmental responses, and chemical toxicity. In biotechnology and biomedical research, RNA-seq facilitates the identification of regulatory networks and biomarkers that inform therapeutic development, risk assessment, and precision medicine.

Realistic in silico generation and augmentation of single-cell RNA-seq data using generative adversarial networks
A fundamental problem in biomedical research is the low number of observations available, mostly due to a lack of available biosamples, prohibitive costs, or ethical reasons. Augmenting few real observations with generated in silico samples could lead to more robust analysis results and a higher reproducibility rate. Here, we propose the use of conditional single-cell generative adversarial neural networks (cscGAN) for the realistic generation of single-cell RNA-seq data. cscGAN learns non-linear gene–gene dependencies from complex, multiple cell type samples and uses this information to generate realistic cells of defined types. Augmenting sparse cell populations with cscGAN generated cells improves downstream analyses such as the detection of marker genes, the robustness and reliability of classifiers, the assessment of novel analysis algorithms, and might reduce the number of animal experiments and costs in consequence. cscGAN outperforms existing methods for single-cell RNA-seq data generation in quality and hold great promise for the realistic generation and augmentation of other biomedical data types.

Representer Point Selection for Explaining Deep Neural Networks
We propose to explain the predictions of a deep neural network, by pointing to the set of what we call representer points in the training set, for a given test point prediction. Specifically, we show that we can decompose the pre-activation prediction of a neural network into a linear combination of activations of training points, with the weights corresponding to what we call representer values, which thus capture the importance of that training point on the learned parameters of the network. But it provides a deeper understanding of the network than simply training point influence: with positive representer values corresponding to excitatory training points, and negative values corresponding to inhibitory points, which as we show provides considerably more insight. Our method is also much more scalable, allowing for real-time feedback in a manner not feasible with influence functions.
TRAK: Attributing Model Behavior at Scale
The goal of data attribution is to trace model predictions back to training data. Despite a long line of work towards this goal, existing approaches to data attribution tend to force users to choose between computational tractability and efficacy. That is, computationally tractable methods can struggle with accurately attributing model predictions in non-convex settings (e.g., in the context of deep neural networks), while methods that are effective in such regimes require training thousands of models, which makes them impractical for large models or datasets. In this work, we introduce TRAK (Tracing with the Randomly-projected After Kernel), a data attribution method that is both effective and computationally tractable for large-scale, differentiable models. In particular, by leveraging only a handful of trained models, TRAK can match the performance of attribution methods that require training thousands of models. We demonstrate the utility of TRAK across various modalities and scales: image classifiers trained on ImageNet, vision-language models (CLIP), and language models (BERT and mT5). We provide code for using TRAK (and reproducing our work) at https://github.com/MadryLab/trak .
Software in the natural world: A computational approach to hierarchical emergence
Understanding the functional architecture of complex systems is crucial to illuminate their inner workings and enable effective methods for their prediction and control. Recent advances have introduced tools to characterise emergent macroscopic levels; however, while these approaches are successful in identifying when emergence takes place, they are limited in the extent they can determine how it does. Here we address this limitation by developing a computational approach to emergence, which characterises macroscopic processes in terms of their computational capabilities. Concretely, we articulate a view on emergence based on how software works, which is rooted on a mathematical formalism that articulates how macroscopic processes can express self-contained informational, interventional, and computational properties. This framework establishes a hierarchy of nested self-contained processes that determines what computations take place at what level, which in turn delineates the functional architecture of a complex system. This approach is illustrated on paradigmatic models from the statistical physics and computational neuroscience literature, which are shown to exhibit macroscopic processes that are akin to software in human-engineered systems. Overall, this framework enables a deeper understanding of the multi-level structure of complex systems, revealing specific ways in which they can be efficiently simulated, predicted, and controlled.

Molecular Biology in the Work of Deleuze and Guattari
This article looks at Deleuze and Guattari's understanding of molecular biology, focusing particularly on their reading of two highly influential works by the eminent French molecular biologists François Jacob and Jacques Monod, La logique du vivant (The Logic of Living Systems) and Le hasard et la nécessité (Chance and Necessity). In these two works, Jacob and Monod present the significance of molecular biology in broadly reductionist terms. What is more, the lac operon model of gene regulation that they propose serves to reinforce the so-called Central Dogma of molecular biology, according to which information passes from DNA to RNA to proteins, with no reverse route. However, Deleuze and Guattari discover intensive potentials within the descriptions of molecular biology offered by both writers. It is argued that Jacob's work in particular, as it has developed in the years since the publication of La logique du vivant in 1970, has itself developed these intensive potentials.
The lost art of mathematical modelling
We provide a critique of mathematical biology in light of rapid developments in modern machine learning. We argue that out of the three modelling activities – (1) formulating models; (2) analysing models; and (3) fitting or comparing models to data – inherent to mathematical biology, researchers currently focus too much on activity (2) at the cost of (1). This trend, we propose, can be reversed by realising that any given biological phenomenon can be modelled in an infinite number of different ways, through the adoption of a pluralistic approach, where we view a system from multiple, different points of view. We explain this pluralistic approach using fish locomotion as a case study and illustrate some of the pitfalls – universalism, creating models of models, etc. – that hinder mathematical biology. We then ask how we might rediscover a lost art: that of creative mathematical modelling.