arXiv Machine Learning By Xinrui He, Mengting Ai, Junting Wang, Curtiss B. Cook, Jingrui He

SafeImpute: Reliable Clinical Data Imputation via Conformal Selection

Read the original on arXiv Machine Learning →

arXiv:2607. 05613v1 Announce Type: new Abstract: Clinical care often relies on key laboratory indicators, yet real-world patient visits are sparse and tests are ordered irregularly, leading to pervasive missingness.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 14

Imputation-free transformer learning enables robust Alzheimer's disease prediction and calibrated uncertainty quantification across heterogeneous clinical cohorts

arXiv:2607. 11656v1 Announce Type: cross Abstract: Accurate diagnostic classification and disease-severity prediction for Alzheimer's disease are hampered by the incompleteness and heterogeneity of real-world clinical data.

By Christelle Schneuwly Diaz, Narmina Baghirova, Duy-Thanh Vu, Duy-Cat Can, Gilles Allali, Philippe Ryvlin, Oliver Y. Ch\'en
arXiv AI
Sep 24

Fed-ReMasker: Federated Tabular Imputation under Feature-Level Missingness

Fed-ReMasker is a federated learning approach that adapts the ReMasker masked autoencoder for tabular data imputation, specifically addressing feature-level missingness where entire features are absent at some centers. The method enables centers to impute unobserved features by leveraging knowledge from collaborating institutions. In benchmark tests on synthetic and real-world datasets, Fed-ReMasker achieves the lowest imputation error in the majority of scenarios and remains robust to client heterogeneity, closely matching the performance of a centralized model.

By Ioannis Papathanail, Rooholla Poursoleymani, Lubnaa Abdur Rahman, Stavroula Georgia Mougiakakou
arXiv AI
Aug 24

Curriculum-Aware Interpolate-then-Refine: Learned Physiological Time-Series Imputation under Realistic Missingness

The paper introduces Curriculum‑Aware Interpolate‑then‑Refine (CAIR), a two‑stage framework for imputing physiological time‑series data. CAIR first learns a coarse base curve with a bidirectional‑GRU interpolator and then refines it through three Transformer passes, trained under a random‑gap curriculum that mimics realistic missingness. Evaluations on continuous glucose monitoring and arterial pressure datasets show CAIR outperforms all baselines across MCAR, MAR, and NMAR mechanisms, especially for long gaps and value‑dependent dropout, while also preserving clinically relevant burden metrics.

By Yu-Chao Huang, Haochen Zhang, Nicholas Konz, Tianlong Chen