arXiv AI By Chuyao Zhang, E Li, Taochen Chen, Yiqun Zhang, Yuzhu Ji, Shuping Zhao, Peng Liu, Yiu-ming Cheung

Imputation Meets Clustering: Exploiting Latent Subgroup Structure for Missing Data Recovery

Read the original on arXiv AI →

arXiv:2607. 06930v1 Announce Type: cross Abstract: Missing data is prevalent in practical applications, making effective imputation an essential preprocessing step for downstream analysis.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
5d ago

Robust Graph Clustering Network for Multiple Missing Data

The paper introduces the Robust Graph Clustering Network for Multiple Missing Data (RGCN), a method designed to cluster graphs with simultaneous missing node attributes and structural links. RGCN employs a view‑decoupled dual‑branch imputation to reduce cross‑view interference, a multi‑hyperspherical mixture prior to improve cluster compactness and separability on a directional latent manifold, and a boundary‑aware contrastive enhancement objective to counteract cluster blurring caused by imputation bias. Experiments on real‑world datasets show that RGCN consistently outperforms state‑of‑the‑art baselines across various missing data patterns.

By Keyuan Qiu, Renda Han, Zhen Tang, Qiang He, Xingwei Wang, Wenxin Zhang, Guangzhen Yao, Junxin Chen, Qingjian Ni
arXiv Machine Learning
Jun 16

Unsupervised Learning for Missing Modalities in Multimodal Learning

arXiv:2606. 15743v1 Announce Type: new Abstract: This paper addresses the missing-modality challenge in multi-modal learning by introducing Unsupervised Learning for Missing Modalities in Multi-Modal Learning (UL4M4), a flexible framework that imputes missing feature embeddings in a task-independent manner before supervised prediction.

By Hassan Ismkhan, Hamid Bouchahcia
arXiv Machine Learning
Aug 19

One Pipeline, Many Transformers: Pattern-Specific Imputation Specialists for Tabular Missing Data

The paper introduces a pre‑training pipeline that creates transformer‑based imputation specialists for tabular data with specific missingness patterns. By featurizing entries, generating synthetic data with configurable missingness modules, and fitting on millions of synthetic tables, the pipeline produces pattern‑specific models that outperform dedicated methods for each missingness pattern. A default model trained only on MCAR data, TabImpute, remains robust across all tested patterns, and the authors release the pipeline, models, and a new benchmark of 42 datasets and 11 missingness patterns.

By Jacob Feitelberg, Dwaipayan Saha, Kyuseong Choi, Zaid Ahmad, Anish Agarwal, Raaz Dwivedi