arXiv Machine Learning

SafeImpute: Reliable Clinical Data Imputation via Conformal Selection

arXiv:2607. 05613v1 Announce Type: new Abstract: Clinical care often relies on key laboratory indicators, yet real-world patient visits are sparse and tests are ordered irregularly, leading to pervasive missingness.

arXiv Machine Learning
Jul 14

Imputation-free transformer learning enables robust Alzheimer's disease prediction and calibrated uncertainty quantification across heterogeneous clinical cohorts

arXiv:2607. 11656v1 Announce Type: cross Abstract: Accurate diagnostic classification and disease-severity prediction for Alzheimer's disease are hampered by the incompleteness and heterogeneity of real-world clinical data.

By Christelle Schneuwly Diaz, Narmina Baghirova, Duy-Thanh Vu, Duy-Cat Can, Gilles Allali, Philippe Ryvlin, Oliver Y. Ch\'en
arXiv AI
Sep 24

Fed-ReMasker: Federated Tabular Imputation under Feature-Level Missingness

Fed-ReMasker is a federated learning approach that adapts the ReMasker masked autoencoder for tabular data imputation, specifically addressing feature-level missingness where entire features are absent at some centers. The method enables centers to impute unobserved features by leveraging knowledge from collaborating institutions. In benchmark tests on synthetic and real-world datasets, Fed-ReMasker achieves the lowest imputation error in the majority of scenarios and remains robust to client heterogeneity, closely matching the performance of a centralized model.

By Ioannis Papathanail, Rooholla Poursoleymani, Lubnaa Abdur Rahman, Stavroula Georgia Mougiakakou
arXiv AI
Aug 24

Curriculum-Aware Interpolate-then-Refine: Learned Physiological Time-Series Imputation under Realistic Missingness

The paper introduces Curriculum‑Aware Interpolate‑then‑Refine (CAIR), a two‑stage framework for imputing physiological time‑series data. CAIR first learns a coarse base curve with a bidirectional‑GRU interpolator and then refines it through three Transformer passes, trained under a random‑gap curriculum that mimics realistic missingness. Evaluations on continuous glucose monitoring and arterial pressure datasets show CAIR outperforms all baselines across MCAR, MAR, and NMAR mechanisms, especially for long gaps and value‑dependent dropout, while also preserving clinically relevant burden metrics.

By Yu-Chao Huang, Haochen Zhang, Nicholas Konz, Tianlong Chen
arXiv Machine Learning
Aug 19

One Pipeline, Many Transformers: Pattern-Specific Imputation Specialists for Tabular Missing Data

The paper introduces a pre‑training pipeline that creates transformer‑based imputation specialists for tabular data with specific missingness patterns. By featurizing entries, generating synthetic data with configurable missingness modules, and fitting on millions of synthetic tables, the pipeline produces pattern‑specific models that outperform dedicated methods for each missingness pattern. A default model trained only on MCAR data, TabImpute, remains robust across all tested patterns, and the authors release the pipeline, models, and a new benchmark of 42 datasets and 11 missingness patterns.

By Jacob Feitelberg, Dwaipayan Saha, Kyuseong Choi, Zaid Ahmad, Anish Agarwal, Raaz Dwivedi
arXiv AI
Sep 15

GRIN+: Towards Fast Yet Effective Machine Unlearning for Imbalanced Medical Data

GRIN+ is a new machine unlearning framework that targets fast and precise data erasure in imbalanced medical datasets. It separates unlearning‑specific knowledge from general representations by analyzing gradient contributions of forget and retain sets, introduces a class‑adaptive influence scoring to counter gradient dominance, and uses a direction‑constrained update to protect essential clinical knowledge. Benchmarks on skin cancer, brain tumor, and breast ultrasound data show that GRIN+ balances privacy, efficiency, and utility, achieving high diagnostic accuracy and faster runtime than existing methods.

By Minghui Huang, Junxiao Wang