The paper introduces a pre‑training pipeline that creates transformer‑based imputation specialists for tabular data with specific missingness patterns. By featurizing entries, generating synthetic data with configurable missingness modules, and fitting on millions of synthetic tables, the pipeline produces pattern‑specific models that outperform dedicated methods for each missingness pattern. A default model trained only on MCAR data, TabImpute, remains robust across all tested patterns, and the authors release the pipeline, models, and a new benchmark of 42 datasets and 11 missingness patterns.
By Jacob Feitelberg, Dwaipayan Saha, Kyuseong Choi, Zaid Ahmad, Anish Agarwal, Raaz Dwivedi
The paper introduces SCR-MF, a two‑stage workflow for single‑cell RNA sequencing imputation that first detects dropout events with scRecover and then imputes missing values using the non‑parametric missForest algorithm. Benchmarking on public and simulated datasets shows that SCR‑MF delivers robust, interpretable results that match or surpass existing methods while maintaining biological fidelity. Runtime analysis indicates that SCR‑MF balances accuracy with computational efficiency, making it well suited for mid‑scale single‑cell studies.
By Ali Anaissi, Deshao Liu, Yuanzhe Jia, Weidong Huang, Widad Alyassine, Junaid Akram
arXiv:2607. 07767v1 Announce Type: cross Abstract: Missing values undermine statistical inference and machine learning pipelines, yet most imputation methods rely on heuristics or restrictive parametric assumptions that ignore the joint data distribution.
By Andrea Basteri, Carlo Ciliberto, Alessandro Rudi
arXiv:2601. 14653v3 Announce Type: replace Abstract: Missing data in single-cell sequencing datasets poses significant challenges for extracting meaningful biological insights.
By Yuyu Liu, Jiannan Yang, Ziyang Yu, Weishen Pan, Fei Wang, Tengfei Ma
arXiv:2606. 05073v1 Announce Type: new Abstract: Missing value imputation is a fundamental task in machine learning, with most existing methods assuming that all missing entries correspond to unobserved regular values.
By Lixing Zhang, Yidong Ouyang, Weifu Li, Shixiang Zhu, Guang Cheng, Liyan Xie
arXiv:2606. 06328v1 Announce Type: new Abstract: In healthcare, multimodal time series tasks often operate on incomplete observations in practice, for example when ECG segments are lost because electrodes detach or an entire respiratory channel is unavailable during overnight monitoring.
By Ziwen Kan, Wugeng Zheng, Tianlong Chen, Song Wang
Missing value imputation is a fundamental task in machine learning, with most existing methods assuming that all missing entries correspond to unobserved regular values. In many real-world datasets, however, missingness may arise from two distinct sources: some entries are meaningfully missing (intrinsically absent and semantically valid), while others are missing due to the observation process and should be imputed.
The paper introduces the Masked Diffusion Time-series Imputation Model (MDTIM), which uses a masked diffusion training paradigm to directly predict original values for time series imputation. It separates missing and observed data via a MASK token and employs Stochastic Discretization to convert continuous values into ordinal-aware tokens, preserving temporal dynamics. Experiments on multiple benchmarks show that MDTIM outperforms existing deterministic and generative baselines in robustness and scalability across various missing data scenarios.
By Dongbin Kim, Seungyun Lee, Geonwoo Shin, Jaewook Lee
Fed-ReMasker is a federated learning approach that adapts the ReMasker masked autoencoder for tabular data imputation, specifically addressing feature-level missingness where entire features are absent at some centers. The method enables centers to impute unobserved features by leveraging knowledge from collaborating institutions. In benchmark tests on synthetic and real-world datasets, Fed-ReMasker achieves the lowest imputation error in the majority of scenarios and remains robust to client heterogeneity, closely matching the performance of a centralized model.
By Ioannis Papathanail, Rooholla Poursoleymani, Lubnaa Abdur Rahman, Stavroula Georgia Mougiakakou
arXiv:2609.37632v1 Announce Type: cross
Abstract: Time series imputation has progressed from statistical and deep learning approaches to diffusion-based models, which have shown strong recent perform...
By Fariza Rashid, Duc Van Le, Rahat Masood, Gustavo Batista, Aruna Seneviratne, Suranga Seneviratne
arXiv:2609.37664v1 Announce Type: new
Abstract: Causal Normalizing Flows (CNFs) enable causal inference from observational data given the causal structure, but they assume fully observed training dat...
By Trung-Dung Hoang, Alceu Bissoto, Tim Fl\"uhmann, David Herzig, Christos Nakas, Lia Bally, Lisa M. Koch
arXiv:2607. 06930v1 Announce Type: cross Abstract: Missing data is prevalent in practical applications, making effective imputation an essential preprocessing step for downstream analysis.
By Chuyao Zhang, E Li, Taochen Chen, Yiqun Zhang, Yuzhu Ji, Shuping Zhao, Peng Liu, Yiu-ming Cheung