Stochastic Flow Map for Count Data
arXiv:2609.23290v1 Announce Type: cross Abstract: High-dimensional count data are common in scientific applications, but most diffusion and flow models are designed for continuous or categorical data...
CRNDiff is a new count‑native diffusion framework that uses stochastic chemical reaction networks to model nonnegative integer data such as single‑cell RNA sequencing. It provides a closed‑form forward‑noising kernel, enabling efficient reverse sampling via forward‑filtering backward‑sampling and data‑driven selection of the terminal noising time. The method also introduces tilted Feynman–Kac steering to sample rare subpopulations without retraining, and demonstrates superior conditional fidelity and marker‑level preservation on human heart scRNA‑seq data.
arXiv:2609.23290v1 Announce Type: cross Abstract: High-dimensional count data are common in scientific applications, but most diffusion and flow models are designed for continuous or categorical data...
Flow Matching for Count Data introduces count‑FM, a flow‑matching framework tailored to high‑dimensional count data such as single‑cell RNA sequencing and neural spike trains. The method models transitions with a continuous‑time birth‑death process that uses local unit jumps, enabling efficient, simulation‑free learning of conditional transition rates directly in count space. Experiments show that count‑FM variants achieve strong sample quality with fewer parameters and provide interpretable transport paths for tasks including unconditional generation, transport, and conditional generation on real biological datasets.
arXiv:2604. 03779v2 Announce Type: replace-cross Abstract: Diffusion models have excelled at generative tasks for both continuous and token-based domains, but their application to discrete ordinal data remains underdeveloped.
arXiv:2605. 00545v2 Announce Type: replace-cross Abstract: Inferring cellular trajectories from destructive snapshots is complicated by the challenges of stochasticity and non-conservative mass dynamics such as cell proliferation and apoptosis.
arXiv:2602. 24201v2 Announce Type: replace Abstract: Estimating density ratios between pairs of intractable data distributions is a core problem in probabilistic modeling, enabling principled comparisons of sample likelihoods under different data-generating processes across conditions.
arXiv:2607. 29043v1 Announce Type: cross Abstract: Single-cell RNA sequencing (scRNA-seq) has become an essential tool in modern cellular biology, and generating accurate synthetic scRNA-seq data is becoming increasingly important.
arXiv:2609.35947v1 Announce Type: new Abstract: Many inference-time tasks for pretrained discrete diffusion models and diffusion language models reduce to drawing samples from a tilted version of the...
arXiv:2608.25631v1 Announce Type: cross Abstract: Continuous-time Markov chains (CTMCs) provide the backbone for modeling discrete stochastic dynamics across applied, physical, and biological science...
arXiv:2607. 04780v1 Announce Type: cross Abstract: Sequential Monte Carlo (SMC) methods are a natural tool for post-hoc conditioning of pretrained generative models, but in many applications the mutation kernels used by the particle system are biased approximations of an ideal Feynman--Kac flow.
PopPert is a framework that models population-level joint gene expression distributions to predict transcriptional responses to perturbations in single-cell RNA sequencing data. By using a low‑rank Gaussian Copula, it captures gene co‑expression patterns and eliminates the need for cell‑to‑cell correspondence, thereby reducing sensitivity to single‑cell noise. Across multiple benchmarks, PopPert outperforms existing methods in differential expression recovery, perturbation effect estimation, and distribution matching, demonstrating the effectiveness of population‑level joint distribution learning for unpaired single‑cell data.
Sequential Monte Carlo (SMC) methods are a natural tool for post-hoc conditioning of pretrained generative models, but in many applications the mutation kernels used by the particle system are biased approximations of an ideal Feynman--Kac flow. This paper develops a non-asymptotic error analysis for such SMC samplers.
arXiv:2606. 11286v1 Announce Type: cross Abstract: High-content imaging assays quantify cellular responses to chemical and genetic perturbations, yet continuous trajectories of individual cells are unobservable because cells are chemically fixed at acquisition.