arXiv Machine Learning

ImputeViz: A Visual Analytics Dashboard for Diagnosing Missing Data and Comparing Imputation Methods

arXiv:2607. 08579v1 Announce Type: cross Abstract: Missing data is a persistent obstacle in scientific, social science, and public health research, often biasing analyses and placing accountability on analysts for how they handle missing values.

Hugging Face Trending Papers
Jul 9

ImputeViz: A Visual Analytics Dashboard for Diagnosing Missing Data and Comparing Imputation Methods

Missing data is a persistent obstacle in scientific, social science, and public health research, often biasing analyses and placing accountability on analysts for how they handle missing values. We introduce ImputeViz, an integrated visual analytics dashboard that supports diagnosing missingness, configuring imputation models, and evaluating results.

arXiv Machine Learning
Aug 19

One Pipeline, Many Transformers: Pattern-Specific Imputation Specialists for Tabular Missing Data

The paper introduces a pre‑training pipeline that creates transformer‑based imputation specialists for tabular data with specific missingness patterns. By featurizing entries, generating synthetic data with configurable missingness modules, and fitting on millions of synthetic tables, the pipeline produces pattern‑specific models that outperform dedicated methods for each missingness pattern. A default model trained only on MCAR data, TabImpute, remains robust across all tested patterns, and the authors release the pipeline, models, and a new benchmark of 42 datasets and 11 missingness patterns.

By Jacob Feitelberg, Dwaipayan Saha, Kyuseong Choi, Zaid Ahmad, Anish Agarwal, Raaz Dwivedi
arXiv Machine Learning
Sep 11

A Dataset and Model for Imputing Water Surface Elevation on a Large and Extremely Sparse Spatiotemporal Graph

The paper introduces AmazonSWE, a dataset covering over 19,000 river sections in the Amazon basin for 10 years (2016‑2026) that integrates satellite altimetry, including SWOT, to enable large‑scale spatiotemporal graph imputation. The dataset is extremely sparse—fewer than 1% of sections are observed daily—and features a directed acyclic river topology that is larger and structurally distinct from existing benchmarks. The authors demonstrate that conventional imputation methods struggle with this topology, scale, and sparsity, and propose a bidirectional selective state‑space model that outperforms prior approaches, reducing RMSE against in‑situ gauges by 18‑39% and providing predictions for every river section. whyItMatters":"AmazonSWE offers a novel, real‑world use case that could improve flood forecasting and water resource management by enabling more accurate and comprehensive water surface elevation estimates across a vast, sparsely monitored river network."

By Ruben Cartuyvels, Karim Douch, Gabriele Bertoli, Mounia El Baz, Artemis Vrettou, S\'ebastien Lef\`evre, Diego Fernandez Prieto
arXiv Machine Learning
Aug 7

Handling Missing Data in Probabilistic Regression Trees

arXiv:2608. 06195v1 Announce Type: cross Abstract: Probabilistic Regression Trees (PRTrees) are a smooth and consistent alternative to classical regression trees, producing continuous predictions through probabilistic split assignments.

By Taiane Schaedler Prass, Alisson Silva Neimaier, Guilherme Pumi
Hugging Face Trending Papers
Aug 6

Handling Missing Data in Probabilistic Regression Trees

Probabilistic Regression Trees (PRTrees) are a smooth and consistent alternative to classical regression trees, producing continuous predictions through probabilistic split assignments. This paper extends the PRTree framework to accommodate missing predictor values directly during tree construction, eliminating the need for prior imputation.