arXiv Machine Learning

Stochastic Flow Map for Count Data

arXiv Machine Learning
Sep 23

Flow Matching for Count Data

Flow Matching for Count Data introduces count‑FM, a flow‑matching framework tailored to high‑dimensional count data such as single‑cell RNA sequencing and neural spike trains. The method models transitions with a continuous‑time birth‑death process that uses local unit jumps, enabling efficient, simulation‑free learning of conditional transition rates directly in count space. Experiments show that count‑FM variants achieve strong sample quality with fewer parameters and provide interpretable transport paths for tasks including unconditional generation, transport, and conditional generation on real biological datasets.

By Ganchao Wei, John Pearson
arXiv Machine Learning
5d ago

CRNDiff: Count-Native Diffusion Framework via Chemical Reaction Networks

CRNDiff is a new count‑native diffusion framework that uses stochastic chemical reaction networks to model nonnegative integer data such as single‑cell RNA sequencing. It provides a closed‑form forward‑noising kernel, enabling efficient reverse sampling via forward‑filtering backward‑sampling and data‑driven selection of the terminal noising time. The method also introduces tilted Feynman–Kac steering to sample rare subpopulations without retraining, and demonstrates superior conditional fidelity and marker‑level preservation on human heart scRNA‑seq data.

By Yuxuan Qiu, Praful Gagrani, Tetsuya J Kobayashi
arXiv AI
Jun 8

CountsDiff: A Diffusion Model on the Natural Numbers for Generation and Imputation of Count-Based Data

arXiv:2604. 03779v2 Announce Type: replace-cross Abstract: Diffusion models have excelled at generative tasks for both continuous and token-based domains, but their application to discrete ordinal data remains underdeveloped.

By Renzo G. Soatto, Anders Hoel, Greycen Ren, Shorna Alam, Stephen Bates, Nikolaos P. Daskalakis, Caroline Uhler, Maria Skoularidou
arXiv AI
Jun 2

Variational Learning for Insertion-based Generation

arXiv:2606. 02133v1 Announce Type: cross Abstract: Non-monotonic sequence generation methods, such as masked diffusion models, provide a flexible alternative to left-to-right autoregressive modeling by allowing tokens to be generated in non-fixed and prescribed orders.

By Yangtian Zhang, Zhe Wang, Arthur Gretton, Rex Ying, David van Dijk, Michalis K. Titsias, Jiaxin Shi
arXiv Machine Learning
Sep 23

One-Step Generative Surrogate Models via Block-Triangular Joint Drifting

The paper introduces block‑triangular joint drifting, a method that applies a projected drift field to the joint distribution of consecutive states, enabling one‑step generative surrogate models for stochastic transition dynamics. This architecture preserves the current‑state marginal while directly sampling the conditional distribution of next states, allowing stochastic trajectories to be generated with a single model evaluation per time step. Experiments show that the approach achieves accurate marginal and trajectory‑dependent statistics with favorable accuracy‑cost tradeoffs compared to deterministic, diffusion, flow, and distillation‑based generative surrogates.

By Nicholas Geissler, Shreya Jha, Ricardo Baptista, Benjamin Peherstorfer
arXiv Machine Learning
Jun 2

Flow-Based Density Ratio Estimation for Intractable Distributions with Applications in Genomics

arXiv:2602. 24201v2 Announce Type: replace Abstract: Estimating density ratios between pairs of intractable data distributions is a core problem in probabilistic modeling, enabling principled comparisons of sample likelihoods under different data-generating processes across conditions.

By Egor Antipov, Alessandro Palma, Lorenzo Consoli, Stephan G\"unnemann, Andrea Dittadi, Fabian J. Theis