arXiv Machine Learning By Jacob Feitelberg, Dwaipayan Saha, Kyuseong Choi, Zaid Ahmad, Anish Agarwal, Raaz Dwivedi

One Pipeline, Many Transformers: Pattern-Specific Imputation Specialists for Tabular Missing Data

Read the original on arXiv Machine Learning →

The paper introduces a pre‑training pipeline that creates transformer‑based imputation specialists for tabular data with specific missingness patterns. By featurizing entries, generating synthetic data with configurable missingness modules, and fitting on millions of synthetic tables, the pipeline produces pattern‑specific models that outperform dedicated methods for each missingness pattern. A default model trained only on MCAR data, TabImpute, remains robust across all tested patterns, and the authors release the pipeline, models, and a new benchmark of 42 datasets and 11 missingness patterns.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Jun 3

Learning What Not to Impute: An Uncertainty-Aware Diffusion Framework for Meaningful Missingness

Missing value imputation is a fundamental task in machine learning, with most existing methods assuming that all missing entries correspond to unobserved regular values. In many real-world datasets, however, missingness may arise from two distinct sources: some entries are meaningfully missing (intrinsically absent and semantically valid), while others are missing due to the observation process and should be imputed.

arXiv AI
Sep 24

Fed-ReMasker: Federated Tabular Imputation under Feature-Level Missingness

Fed-ReMasker is a federated learning approach that adapts the ReMasker masked autoencoder for tabular data imputation, specifically addressing feature-level missingness where entire features are absent at some centers. The method enables centers to impute unobserved features by leveraging knowledge from collaborating institutions. In benchmark tests on synthetic and real-world datasets, Fed-ReMasker achieves the lowest imputation error in the majority of scenarios and remains robust to client heterogeneity, closely matching the performance of a centralized model.

By Ioannis Papathanail, Rooholla Poursoleymani, Lubnaa Abdur Rahman, Stavroula Georgia Mougiakakou