The paper introduces UdonCare, a hierarchy‑pruning method that iteratively partitions patients into latent domains using medical ontologies, aiming to improve domain generalization in clinical prediction tasks. It addresses challenges of missing domain labels and lack of clinical insight by discovering hierarchy‑grounded patient domains. Experiments on MIMIC‑III, MIMIC‑IV, and eICU datasets show UdonCare outperforms eight baseline methods across four prediction tasks with significant domain gaps.
By Pengfei Hu, Xiaoxue Han, Fei Wang, Yue Ning
arXiv:2607. 15447v1 Announce Type: new Abstract: Recent research in clinical machine learning, focusing on outcome predictions in intensive care unit (ICU), has shifted from bespoke supervised models to foundation models, utilising modern representation learning methods.
By Jingteng Li, Alexander Capstick, Louise Rigny, Iona Biggart, Neil J Sebire, Payam Barnaghi
arXiv:2607. 06163v1 Announce Type: cross Abstract: Foundation Models for Electronic Health Records (FEMRs) are pretrained on large-scale structured patient data, enabling them to convert longitudinal patient trajectories into generalizable representations for diverse clinical prediction tasks.
By Jie Huang, Pengfei Yin, Zihan Xu, Daniel Capurro, Mike Conway, Ting Dang
arXiv:2608. 08182v1 Announce Type: cross Abstract: Machine learning models for MALDI-TOF mass spectrometry have shown considerable promise for clinical microbiology tasks such as microbial identification and antimicrobial resistance prediction.
By Alejandro L. Garc\'ia-Navarro, Carlos Sevilla-Salcedo, Bel\'en Rodr\'iguez-S\'anchez, Vanessa G\'omez-Verdejo
The paper introduces CAST, a concept-guided artifact suppression tuning framework that uses sparse autoencoders to identify and suppress note-specific artifacts in clinical language models. CAST labels latent features with an LLM-assisted pipeline and ICD‑10 constraints, then fine‑tunes the model while providing post‑hoc per‑concept attributions for auditability. In experiments on MIMIC‑IV discharge‑note mortality prediction, CAST outperforms standard fine‑tuned encoders and competes with strong LLM baselines while offering a feature‑level audit trail of clinical concepts and suppressed artifacts.
By Jin Mu, Guanhua Chen
EXPOSE is a framework that applies Sparse Autoencoders to Vision Foundation Model embeddings in computational pathology, aiming to separate biological signals from domain‑specific noise. By training a sparse representation of VFM features and using a linear classifier to flag domain‑specific latent dimensions, the method masks these components before downstream relapse prediction, avoiding the need to retrain the backbone model. Experiments on a large prostate cancer dataset demonstrate that removing domain‑specific features improves cross‑domain performance and raises the Domain Robustness Index (DoRI).
By Anja Witte, Maximilian Lennartz, Jan Baumbach, Guido Sauter, Stefan Bonn, Patrick Fuhlert, Marina Zimmermann
arXiv:2603. 02221v2 Announce Type: replace-cross Abstract: In clinical tabular prediction, classical machine learning models with feature engineering often outperform neural methods.
By Zizheng Zhang, Yiming Li, Justin Xu, Jinyu Wang, Rui Wang, Lei Song, Jiang Bian, David W Eyre, Jingjing Fu
arXiv:2608. 10969v1 Announce Type: new Abstract: Healthcare data, such as Intensive Care Unit (ICU) records, comprise heterogeneous multivariate time series sampled at irregular intervals with pervasive missingness.
By Ruirui Wang, Yanke Li, Manuel G\"unther, Diego Paez-Granados
The paper introduces a deep generative model using conditional variational autoencoders to augment vital sign data from healthy individuals so that it mimics patterns of specific clinical conditions. Trained on a publicly available ICU dataset, the model learns the underlying dynamics of ICU data and reshapes healthy data to align with target clinical labels. A proposed distance metric demonstrates that the generated samples are more aligned with intended clinical labels than baseline methods.
By Rafael Pina, Varuna De Silva, Mindula Illeperuma
arXiv:2601.17228v2 Announce Type: replace
Abstract: Deep learning models in computational pathology often fail to generalize across cohorts and institutions due to domain shift. Existing approaches e...
By Tengyue Zhang, Ruiwen Ding, Luoting Zhuang, Yuxiao Wu, Erika F. Rodriguez, William Hsu
arXiv:2606. 28798v1 Announce Type: new Abstract: Objective: ICD codes are central to reimbursement, research, and population health surveillance, yet automated coding systems often struggle to integrate diagnostic signals from both clinical narratives and structured electronic health record (EHR) variables.
By Chengyuan Liu, Xinyue Zhang, Yao Li, Guanting Chen
arXiv:2608. 20315v1 Announce Type: new Abstract: Predictive models over structured electronic health records (EHRs) remain central to machine learning for healthcare, but few have jointly emphasized quantitative laboratory information and interpretability with respect to input medical events.
By Jun Ni Du, Lukas Adamek, Maxim Kryukov, Flavio Dormont, Ziv Bar-Joseph, Sven Jager, Brandon Rufino