arXiv AI By Renzo G. Soatto, Anders Hoel, Greycen Ren, Shorna Alam, Stephen Bates, Nikolaos P. Daskalakis, Caroline Uhler, Maria Skoularidou

CountsDiff: A Diffusion Model on the Natural Numbers for Generation and Imputation of Count-Based Data

Read the original on arXiv AI →

arXiv:2604. 03779v2 Announce Type: replace-cross Abstract: Diffusion models have excelled at generative tasks for both continuous and token-based domains, but their application to discrete ordinal data remains underdeveloped.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 23

Flow Matching for Count Data

Flow Matching for Count Data introduces count‑FM, a flow‑matching framework tailored to high‑dimensional count data such as single‑cell RNA sequencing and neural spike trains. The method models transitions with a continuous‑time birth‑death process that uses local unit jumps, enabling efficient, simulation‑free learning of conditional transition rates directly in count space. Experiments show that count‑FM variants achieve strong sample quality with fewer parameters and provide interpretable transport paths for tasks including unconditional generation, transport, and conditional generation on real biological datasets.

By Ganchao Wei, John Pearson
arXiv Machine Learning
5d ago

CRNDiff: Count-Native Diffusion Framework via Chemical Reaction Networks

CRNDiff is a new count‑native diffusion framework that uses stochastic chemical reaction networks to model nonnegative integer data such as single‑cell RNA sequencing. It provides a closed‑form forward‑noising kernel, enabling efficient reverse sampling via forward‑filtering backward‑sampling and data‑driven selection of the terminal noising time. The method also introduces tilted Feynman–Kac steering to sample rare subpopulations without retraining, and demonstrates superior conditional fidelity and marker‑level preservation on human heart scRNA‑seq data.

By Yuxuan Qiu, Praful Gagrani, Tetsuya J Kobayashi
arXiv Machine Learning
Sep 22

Stochastic Flow Map for Count Data

arXiv:2609.23290v1 Announce Type: cross Abstract: High-dimensional count data are common in scientific applications, but most diffusion and flow models are designed for continuous or categorical data...

By Ganchao Wei
arXiv AI
Aug 20

Discretizing Continuous Time Series for Imputation with Masked Diffusion Training

The paper introduces the Masked Diffusion Time-series Imputation Model (MDTIM), which uses a masked diffusion training paradigm to directly predict original values for time series imputation. It separates missing and observed data via a MASK token and employs Stochastic Discretization to convert continuous values into ordinal-aware tokens, preserving temporal dynamics. Experiments on multiple benchmarks show that MDTIM outperforms existing deterministic and generative baselines in robustness and scalability across various missing data scenarios.

By Dongbin Kim, Seungyun Lee, Geonwoo Shin, Jaewook Lee