arXiv Machine Learning

Discovering Subgroups with Exceptional Survival Characteristics

arXiv:2602. 22179v2 Announce Type: replace Abstract: In many applications, it is important to identify subpopulations that survive longer or shorter than the rest of the population.

arXiv Statistics ML
Sep 23

Efficient and scalable clustering of survival curves

arXiv:2512.16481v2 Announce Type: replace-cross Abstract: Survival analysis encompasses a broad range of methods for analyzing time-to-event data, with one key objective being the comparison of survi...

By Nora M. Villanueva, Marta Sestelo, Luis Meira-Machado
arXiv Computation and Language
3d ago

Large Language Models are Approximate Survival Estimators

arXiv:2609.38181v1 Announce Type: new Abstract: Survival analysis estimates time-to-event outcomes from patient covariates and is widely used for medical risk assessment. Patients seeking prognostic...

By Juan M Zambrano Chaves, Peniel Argaw, Risa Ueno, Carlo Bifulco, Kristina Young, Rom Leidner, Tristan Naumann, Hoifung Poon
arXiv Machine Learning
Jun 5

Symb-xMIL: Symbolic Explanations for Multiple Instance Learning in Digital Pathology

arXiv:2606. 06224v1 Announce Type: cross Abstract: Explanations of multiple instance learning (MIL) models are widely used for validation and discovery in digital histopathology.

By Yanqing Luo (Berlin Institute for the Foundations of Learning and Data, Berlin, Germany, Machine Learning Group, Technische Universit\"at Berlin, Berlin, Germany), Julius Hense (Berlin Institute for the Foundations of Learning and Data, Berlin, Germany, Machine Learning Group, Technische Universit\"at Berlin, Berlin, Germany), Niklas Preni{\ss}l (Institute of Pathology, Charit\'e Universit\"atsmedizin, Berlin, Germany, Berlin Institute of Health at Charit\'e -- Universit\"atsmedizin Berlin, BIH Biomedical Innovation Academy, BIH Charit\'e Digital Clinician Scientist Program, Berlin, Germany), Andreas Mock (Institute of Pathology, Ludwig Maximilian University of Munich, Munich, Germany, Division of Translational Medical Oncology, DKFZ, Heidelberg, Germany, NCT Heidelberg, Heidelberg, Germany, German Cancer Consortium), Klaus-Robert M\"uller (Berlin Institute for the Foundations of Learning and Data, Berlin, Germany, Machine Learning Group, Technische Universit\"at Berlin, Berlin, Germany, Department of Artificial Intelligence, Korea University, Seoul, Korea, Max-Planck Institute for Informatics, Saarbr\"ucken, Germany), Thomas Schnake (Department of Chemistry, Chemical Physics Theory Group, University of Toronto, Canada, Vector Institute for Artificial Intelligence, Toronto, Canada, Acceleration Consortium, University of Toronto, Canada), Mina Jamshidi Idaji (Berlin Institute for the Foundations of Learning and Data, Berlin, Germany, Machine Learning Group, Technische Universit\"at Berlin, Berlin, Germany)
arXiv Machine Learning
Sep 22

From Latent Biomarkers to Clinical Rules: Embedding-Guided Rule Mining and Attribution-Based Translation for Interpretable Tabular Learning

The paper introduces a four-step pipeline that mines decision rules in the latent space of an FT-Transformer and then translates those rules back into measurable clinical features. By treating embedding dimensions that separate patient groups as latent biomarkers, small decision trees are used to extract rules, which are then mapped to raw features using gradient-input saliency and CLS attention attribution. Across six public clinical datasets, the translated rules generally outperformed raw-feature rules, achieving significant AUROC gains, though some high-performing latent rules could not be fully captured by simple raw-feature conditions.

By Majid Lotfian Delouee, Hamed Ayoobi, Sjors G. J. G. In 't Veld, Martijn C. Schut