Flow Matching for Count Data introduces count‑FM, a flow‑matching framework tailored to high‑dimensional count data such as single‑cell RNA sequencing and neural spike trains. The method models transitions with a continuous‑time birth‑death process that uses local unit jumps, enabling efficient, simulation‑free learning of conditional transition rates directly in count space. Experiments show that count‑FM variants achieve strong sample quality with fewer parameters and provide interpretable transport paths for tasks including unconditional generation, transport, and conditional generation on real biological datasets.
By Ganchao Wei, John Pearson
CRNDiff is a new count‑native diffusion framework that uses stochastic chemical reaction networks to model nonnegative integer data such as single‑cell RNA sequencing. It provides a closed‑form forward‑noising kernel, enabling efficient reverse sampling via forward‑filtering backward‑sampling and data‑driven selection of the terminal noising time. The method also introduces tilted Feynman–Kac steering to sample rare subpopulations without retraining, and demonstrates superior conditional fidelity and marker‑level preservation on human heart scRNA‑seq data.
By Yuxuan Qiu, Praful Gagrani, Tetsuya J Kobayashi
arXiv:2509. 22352v3 Announce Type: replace Abstract: Survival analysis is a cornerstone of clinical research by modeling time-to-event outcomes such as metastasis, disease relapse, or patient death.
By Marie Brockschmidt, Maresa Schr\"oder, Stefan Feuerriegel
arXiv:2609.23290v1 Announce Type: cross
Abstract: High-dimensional count data are common in scientific applications, but most diffusion and flow models are designed for continuous or categorical data...
By Ganchao Wei
arXiv:2606. 17106v1 Announce Type: new Abstract: Laboratory tests in electronic health records are collected irregularly, and the absence of a test order can be as informative as the measurement itself.
By Hadi Mehdizavareh, Gabriele Santangelo, Giovanna Nicora, Simon Lebech Cichosz, Arianna Dagliati, Arijit Khan, Riccardo Bellazzi
The paper introduces the Masked Diffusion Time-series Imputation Model (MDTIM), which uses a masked diffusion training paradigm to directly predict original values for time series imputation. It separates missing and observed data via a MASK token and employs Stochastic Discretization to convert continuous values into ordinal-aware tokens, preserving temporal dynamics. Experiments on multiple benchmarks show that MDTIM outperforms existing deterministic and generative baselines in robustness and scalability across various missing data scenarios.
By Dongbin Kim, Seungyun Lee, Geonwoo Shin, Jaewook Lee
RDDMPI introduces a residual denoising diffusion model for multivariate time series imputation. By decomposing the missing signal into a baseline reconstruction and a residual uncertainty component, the method conditions the diffusion process on both the completed signal and its latent representation, using a reliability-aware mechanism to balance baseline influence. Experiments on benchmark datasets show that this approach improves reconstruction accuracy and uncertainty quantification compared to prior diffusion-based methods.
By Ramiro Valdes Jara, David Chapman, Adam Meyers
arXiv:2609.15284v1 Announce Type: new
Abstract: Missing values are ubiquitous in heterogeneous data mining, where numerical, categorical, and binary variables often coexist. Many imputation methods,...
By Sergei Kholkin, Kirill Sokolov, Dmitry Baranchuk, Evgeny Burnaev, Alexander Korotin
arXiv:2606. 05361v1 Announce Type: cross Abstract: Missing data imputation in large-scale surveys faces two challenges that are not well handled by current tabular diffusion methods.
By Yuyu Chen, Taehyo Kim, Hai Shu, Yang Feng
arXiv:2607. 06583v1 Announce Type: cross Abstract: DNA methylation (DNAm) serves as one of the most robust molecular biomarkers of biological aging.
By Chandan Gupta, Syed Haider, Pietro Li\`o
arXiv:2607. 29043v1 Announce Type: cross Abstract: Single-cell RNA sequencing (scRNA-seq) has become an essential tool in modern cellular biology, and generating accurate synthetic scRNA-seq data is becoming increasingly important.
By Yu Song, Hao Sun, Ikuko Nishikawa, Yen-Wei Chen
arXiv:2607. 03103v1 Announce Type: cross Abstract: Clinical cardiac imaging pipelines currently deploy separate models for each dataset and modality, incurring redundant training costs and precluding knowledge sharing across anatomically related tasks.
By Jiahao Liu, Hang Wei, Shuai Wu