Missing value imputation is a fundamental task in machine learning, with most existing methods assuming that all missing entries correspond to unobserved regular values. In many real-world datasets, however, missingness may arise from two distinct sources: some entries are meaningfully missing (intrinsically absent and semantically valid), while others are missing due to the observation process and should be imputed.
arXiv:2609.15284v1 Announce Type: new
Abstract: Missing values are ubiquitous in heterogeneous data mining, where numerical, categorical, and binary variables often coexist. Many imputation methods,...
By Sergei Kholkin, Kirill Sokolov, Dmitry Baranchuk, Evgeny Burnaev, Alexander Korotin
arXiv:2609.37632v1 Announce Type: cross
Abstract: Time series imputation has progressed from statistical and deep learning approaches to diffusion-based models, which have shown strong recent perform...
By Fariza Rashid, Duc Van Le, Rahat Masood, Gustavo Batista, Aruna Seneviratne, Suranga Seneviratne
The paper introduces the Masked Diffusion Time-series Imputation Model (MDTIM), which uses a masked diffusion training paradigm to directly predict original values for time series imputation. It separates missing and observed data via a MASK token and employs Stochastic Discretization to convert continuous values into ordinal-aware tokens, preserving temporal dynamics. Experiments on multiple benchmarks show that MDTIM outperforms existing deterministic and generative baselines in robustness and scalability across various missing data scenarios.
By Dongbin Kim, Seungyun Lee, Geonwoo Shin, Jaewook Lee
arXiv:2606. 03347v1 Announce Type: cross Abstract: Score-based diffusion models have emerged as prominent deep generative models; however, their application to tabular data remains challenging because their backbones assume fully specified inputs, whereas real-world tabular data often contain missing values.
By Jungkyu Kim, Taeyoung Park, Kibok Lee
RDDMPI introduces a residual denoising diffusion model for multivariate time series imputation. By decomposing the missing signal into a baseline reconstruction and a residual uncertainty component, the method conditions the diffusion process on both the completed signal and its latent representation, using a reliability-aware mechanism to balance baseline influence. Experiments on benchmark datasets show that this approach improves reconstruction accuracy and uncertainty quantification compared to prior diffusion-based methods.
By Ramiro Valdes Jara, David Chapman, Adam Meyers
The paper introduces a pre‑training pipeline that creates transformer‑based imputation specialists for tabular data with specific missingness patterns. By featurizing entries, generating synthetic data with configurable missingness modules, and fitting on millions of synthetic tables, the pipeline produces pattern‑specific models that outperform dedicated methods for each missingness pattern. A default model trained only on MCAR data, TabImpute, remains robust across all tested patterns, and the authors release the pipeline, models, and a new benchmark of 42 datasets and 11 missingness patterns.
By Jacob Feitelberg, Dwaipayan Saha, Kyuseong Choi, Zaid Ahmad, Anish Agarwal, Raaz Dwivedi
arXiv:2606. 06328v1 Announce Type: new Abstract: In healthcare, multimodal time series tasks often operate on incomplete observations in practice, for example when ECG segments are lost because electrodes detach or an entire respiratory channel is unavailable during overnight monitoring.
By Ziwen Kan, Wugeng Zheng, Tianlong Chen, Song Wang
arXiv:2607. 29177v1 Announce Type: cross Abstract: Utility data (e.
By Rongchao Xu, Lin Jiang, Dahai Yu, Ximiao Li, Guang Wang
arXiv:2606. 17106v1 Announce Type: new Abstract: Laboratory tests in electronic health records are collected irregularly, and the absence of a test order can be as informative as the measurement itself.
By Hadi Mehdizavareh, Gabriele Santangelo, Giovanna Nicora, Simon Lebech Cichosz, Arianna Dagliati, Arijit Khan, Riccardo Bellazzi
arXiv:2607. 28698v1 Announce Type: new Abstract: Flow matching assumes fully observed training data, which many real-world applications rarely provide.
By Fairoz Nower Khan, Nabuat Zaman Nahim, Peizhong Ju
arXiv:2609.39613v1 Announce Type: new
Abstract: Missing data are a fundamental challenge in statistical analysis and machine learning, as the choice of imputation method substantially impacts downstr...
By Jinwei Li, Michelle Bruch, Daniel Tenbrinck