arXiv AI

MBDiff: Multi-view Behavior-aware Diffusion Model for Probabilistic Utility Data Imputation

arXiv:2607. 29177v1 Announce Type: cross Abstract: Utility data (e.

Hugging Face Trending Papers
Jun 3

Learning What Not to Impute: An Uncertainty-Aware Diffusion Framework for Meaningful Missingness

Missing value imputation is a fundamental task in machine learning, with most existing methods assuming that all missing entries correspond to unobserved regular values. In many real-world datasets, however, missingness may arise from two distinct sources: some entries are meaningfully missing (intrinsically absent and semantically valid), while others are missing due to the observation process and should be imputed.

arXiv Machine Learning
Sep 11

RDDMPI: Residual Denoising Diffusion Model for Probabilistic Multivariate Time Series Imputation

RDDMPI introduces a residual denoising diffusion model for multivariate time series imputation. By decomposing the missing signal into a baseline reconstruction and a residual uncertainty component, the method conditions the diffusion process on both the completed signal and its latent representation, using a reliability-aware mechanism to balance baseline influence. Experiments on benchmark datasets show that this approach improves reconstruction accuracy and uncertainty quantification compared to prior diffusion-based methods.

By Ramiro Valdes Jara, David Chapman, Adam Meyers
arXiv Machine Learning
Aug 19

One Pipeline, Many Transformers: Pattern-Specific Imputation Specialists for Tabular Missing Data

The paper introduces a pre‑training pipeline that creates transformer‑based imputation specialists for tabular data with specific missingness patterns. By featurizing entries, generating synthetic data with configurable missingness modules, and fitting on millions of synthetic tables, the pipeline produces pattern‑specific models that outperform dedicated methods for each missingness pattern. A default model trained only on MCAR data, TabImpute, remains robust across all tested patterns, and the authors release the pipeline, models, and a new benchmark of 42 datasets and 11 missingness patterns.

By Jacob Feitelberg, Dwaipayan Saha, Kyuseong Choi, Zaid Ahmad, Anish Agarwal, Raaz Dwivedi
arXiv AI
Aug 20

Discretizing Continuous Time Series for Imputation with Masked Diffusion Training

The paper introduces the Masked Diffusion Time-series Imputation Model (MDTIM), which uses a masked diffusion training paradigm to directly predict original values for time series imputation. It separates missing and observed data via a MASK token and employs Stochastic Discretization to convert continuous values into ordinal-aware tokens, preserving temporal dynamics. Experiments on multiple benchmarks show that MDTIM outperforms existing deterministic and generative baselines in robustness and scalability across various missing data scenarios.

By Dongbin Kim, Seungyun Lee, Geonwoo Shin, Jaewook Lee
arXiv Machine Learning
Sep 11

LoaDiff: Conditional Generation of Electricity Consumption Time Series for Energy Analytics

LoaDiff is a diffusion-based generative model that produces year-long, sub-hourly smart‑meter electricity consumption time series. It can be conditioned on static household attributes like appliance ownership and dynamic factors such as calendar dates and outdoor temperature. Evaluations on three residential datasets show that LoaDiff generates realistic, diverse load profiles, limits memorization, retains useful information for downstream tasks, and responds coherently to conditioning changes.

By Mariia Baranova, Adrien Petralia, Etienne Le Naour, Nathan Etourneau, Guillaume Hofmann, Themis Palpanas
Hugging Face Trending Papers
Jul 9

ImputeViz: A Visual Analytics Dashboard for Diagnosing Missing Data and Comparing Imputation Methods

Missing data is a persistent obstacle in scientific, social science, and public health research, often biasing analyses and placing accountability on analysts for how they handle missing values. We introduce ImputeViz, an integrated visual analytics dashboard that supports diagnosing missingness, configuring imputation models, and evaluating results.