arXiv:2609.37632v1 Announce Type: cross
Abstract: Time series imputation has progressed from statistical and deep learning approaches to diffusion-based models, which have shown strong recent perform...
By Fariza Rashid, Duc Van Le, Rahat Masood, Gustavo Batista, Aruna Seneviratne, Suranga Seneviratne
arXiv:2607. 07767v1 Announce Type: cross Abstract: Missing values undermine statistical inference and machine learning pipelines, yet most imputation methods rely on heuristics or restrictive parametric assumptions that ignore the joint data distribution.
By Andrea Basteri, Carlo Ciliberto, Alessandro Rudi
The paper introduces a pre‑training pipeline that creates transformer‑based imputation specialists for tabular data with specific missingness patterns. By featurizing entries, generating synthetic data with configurable missingness modules, and fitting on millions of synthetic tables, the pipeline produces pattern‑specific models that outperform dedicated methods for each missingness pattern. A default model trained only on MCAR data, TabImpute, remains robust across all tested patterns, and the authors release the pipeline, models, and a new benchmark of 42 datasets and 11 missingness patterns.
By Jacob Feitelberg, Dwaipayan Saha, Kyuseong Choi, Zaid Ahmad, Anish Agarwal, Raaz Dwivedi
arXiv:2602.01437v2 Announce Type: replace-cross
Abstract: The problem of corrupted data, missing features, or missing modalities continues to plague the modern machine learning landscape. To address...
By Yinsong Wang, Shahin Shahrampour
arXiv:2506. 01544v2 Announce Type: replace Abstract: We introduce Temporal Variational Implicit Neural Representations (TV-INRs), a probabilistic framework for modeling irregular multivariate time series that enables efficient and accurate individualized imputation and forecasting.
By Batuhan Koyuncu, Rachael DeVries, Ole Winther, Isabel Valera
arXiv:2608.30040v1 Announce Type: cross
Abstract: Missing data, measurement error, and population heterogeneity are pervasive challenges in analyzing data arising from modern observational studies an...
By Yasin Khadem Charvadeh, Grace Y. Yi, Mithat G\"onen, Pouya Faroughi
arXiv:2605.04469v2 Announce Type: replace-cross
Abstract: Large-scale population-level datasets, such as the UK Biobank and the All of Us Research Program, often lack covariates needed for a specific...
By Huali Zhao, Tianying Wang
arXiv:2607. 06930v1 Announce Type: cross Abstract: Missing data is prevalent in practical applications, making effective imputation an essential preprocessing step for downstream analysis.
By Chuyao Zhang, E Li, Taochen Chen, Yiqun Zhang, Yuzhu Ji, Shuping Zhao, Peng Liu, Yiu-ming Cheung
arXiv:2607. 08915v1 Announce Type: new Abstract: Missing data is ubiquitous in real-world datasets.
By Minett Tran, Taehee Jeong
arXiv:2606. 06328v1 Announce Type: new Abstract: In healthcare, multimodal time series tasks often operate on incomplete observations in practice, for example when ECG segments are lost because electrodes detach or an entire respiratory channel is unavailable during overnight monitoring.
By Ziwen Kan, Wugeng Zheng, Tianlong Chen, Song Wang
arXiv:2607. 28698v1 Announce Type: new Abstract: Flow matching assumes fully observed training data, which many real-world applications rarely provide.
By Fairoz Nower Khan, Nabuat Zaman Nahim, Peizhong Ju
The paper introduces the Masked Diffusion Time-series Imputation Model (MDTIM), which uses a masked diffusion training paradigm to directly predict original values for time series imputation. It separates missing and observed data via a MASK token and employs Stochastic Discretization to convert continuous values into ordinal-aware tokens, preserving temporal dynamics. Experiments on multiple benchmarks show that MDTIM outperforms existing deterministic and generative baselines in robustness and scalability across various missing data scenarios.
By Dongbin Kim, Seungyun Lee, Geonwoo Shin, Jaewook Lee