Missing data is a persistent obstacle in scientific, social science, and public health research, often biasing analyses and placing accountability on analysts for how they handle missing values. We introduce ImputeViz, an integrated visual analytics dashboard that supports diagnosing missingness, configuring imputation models, and evaluating results.
arXiv:2609.39613v1 Announce Type: new
Abstract: Missing data are a fundamental challenge in statistical analysis and machine learning, as the choice of imputation method substantially impacts downstr...
By Jinwei Li, Michelle Bruch, Daniel Tenbrinck
The paper introduces a pre‑training pipeline that creates transformer‑based imputation specialists for tabular data with specific missingness patterns. By featurizing entries, generating synthetic data with configurable missingness modules, and fitting on millions of synthetic tables, the pipeline produces pattern‑specific models that outperform dedicated methods for each missingness pattern. A default model trained only on MCAR data, TabImpute, remains robust across all tested patterns, and the authors release the pipeline, models, and a new benchmark of 42 datasets and 11 missingness patterns.
By Jacob Feitelberg, Dwaipayan Saha, Kyuseong Choi, Zaid Ahmad, Anish Agarwal, Raaz Dwivedi
arXiv:2607. 29177v1 Announce Type: cross Abstract: Utility data (e.
By Rongchao Xu, Lin Jiang, Dahai Yu, Ximiao Li, Guang Wang
arXiv:2606. 17106v1 Announce Type: new Abstract: Laboratory tests in electronic health records are collected irregularly, and the absence of a test order can be as informative as the measurement itself.
By Hadi Mehdizavareh, Gabriele Santangelo, Giovanna Nicora, Simon Lebech Cichosz, Arianna Dagliati, Arijit Khan, Riccardo Bellazzi
The paper introduces AmazonSWE, a dataset covering over 19,000 river sections in the Amazon basin for 10 years (2016‑2026) that integrates satellite altimetry, including SWOT, to enable large‑scale spatiotemporal graph imputation. The dataset is extremely sparse—fewer than 1% of sections are observed daily—and features a directed acyclic river topology that is larger and structurally distinct from existing benchmarks. The authors demonstrate that conventional imputation methods struggle with this topology, scale, and sparsity, and propose a bidirectional selective state‑space model that outperforms prior approaches, reducing RMSE against in‑situ gauges by 18‑39% and providing predictions for every river section.
whyItMatters":"AmazonSWE offers a novel, real‑world use case that could improve flood forecasting and water resource management by enabling more accurate and comprehensive water surface elevation estimates across a vast, sparsely monitored river network."
By Ruben Cartuyvels, Karim Douch, Gabriele Bertoli, Mounia El Baz, Artemis Vrettou, S\'ebastien Lef\`evre, Diego Fernandez Prieto
arXiv:2606. 06328v1 Announce Type: new Abstract: In healthcare, multimodal time series tasks often operate on incomplete observations in practice, for example when ECG segments are lost because electrodes detach or an entire respiratory channel is unavailable during overnight monitoring.
By Ziwen Kan, Wugeng Zheng, Tianlong Chen, Song Wang
arXiv:2608. 06195v1 Announce Type: cross Abstract: Probabilistic Regression Trees (PRTrees) are a smooth and consistent alternative to classical regression trees, producing continuous predictions through probabilistic split assignments.
By Taiane Schaedler Prass, Alisson Silva Neimaier, Guilherme Pumi
arXiv:2606. 05073v1 Announce Type: new Abstract: Missing value imputation is a fundamental task in machine learning, with most existing methods assuming that all missing entries correspond to unobserved regular values.
By Lixing Zhang, Yidong Ouyang, Weifu Li, Shixiang Zhu, Guang Cheng, Liyan Xie
arXiv:2609.14902v1 Announce Type: cross
Abstract: Shapley value (SV)-based methods are the prevailing framework for feature attribution in machine learning, yet existing population-level Shapley esti...
By Siqi Li, Wangxuan Fan, Yiming Li, Doudou Zhou, Molei Liu
arXiv:2506. 08725v3 Announce Type: replace-cross Abstract: Misuses of t-SNE and UMAP in visual analytics have become increasingly common.
By Hyeon Jeon, Jeongin Park, Sungbok Shin, Jinwook Seo
Probabilistic Regression Trees (PRTrees) are a smooth and consistent alternative to classical regression trees, producing continuous predictions through probabilistic split assignments. This paper extends the PRTree framework to accommodate missing predictor values directly during tree construction, eliminating the need for prior imputation.