arXiv Machine Learning

The C-index illusion: discrimination without calibration in published survival models

arXiv:2607. 19526v1 Announce Type: new Abstract: "Stop Chasing the C-index when Evaluating Survival Analysis Models" (ICML 2026, Spotlight) argued normatively, on synthetic data, that evaluating survival models by discrimination alone, i.

arXiv Machine Learning
Jul 22

When Are Scoring Rules Proper? Bridging Theory and Practice in Survival Model Evaluation

arXiv:2212. 05260v4 Announce Type: replace-cross Abstract: Proper scoring rules encourage probabilistic predictions that match the true underlying distribution and are central to model evaluation, with increasing relevance in automated workflows such as AutoML.

By John Zobolas, Raphael Sonabend, Riccardo De Bin, Johannes Piller, Philipp Kopper, Lukas Burk, Andreas Bender
arXiv AI
Sep 2

Counterfactual Fragility Certificates: Exposing High-Confidence Brittleness under Structured Evidence Failure

The paper introduces Counterfactual Fragility Certificates (CFC), a model‑agnostic audit protocol that maps each prediction to an evidence‑failure trajectory, summarizing it with metrics such as greedy flip budget, margin‑collapse area, degradation thresholds, and fragility dominance score. CFC is shown to identify brittle high‑confidence predictions on seven tabular benchmarks with an AUROC of 0.915, outperforming existing scalar scores by up to +0.405. The method remains effective across various perturbation and review‑budget scenarios, and can also inform fragility‑aware regularization and temperature correction.

By Filippo Cenacchi, Longbing Cao, Runze Yang
arXiv Machine Learning
Aug 3

Incorporating data drift to perform survival analysis on credit risk

arXiv:2601. 20533v2 Announce Type: replace-cross Abstract: Survival analysis has become a standard approach for modelling time to default by time-varying covariates in credit risk.

By Jianwei Peng (Humboldt-Universit\"at zu Berlin), Stefan Lessmann (Humboldt-Universit\"at zu Berlin, Bucharest University of Economic Studies)
arXiv Machine Learning
Aug 31

Timing-Aware Repurchase Prediction for Web-Scale E-Commerce: Survival Models for Multi-Surface Grocery Recommendation

The paper proposes replacing multiple horizon‑specific binary classifiers with a single survival model to predict time‑to‑repurchase in grocery e‑commerce. Empirical analysis shows a slightly decreasing hazard (k≈0.9) and that a Log‑Normal model best fits marginal distributions while Weibull best fits residuals. A single Accelerated Failure Time (AFT) model matches or surpasses per‑horizon classifiers with fewer trees, and a 4‑parameter calibration maps survival CDFs to horizon probabilities without monotonicity violations, revealing a trade‑off between calibration and ranking within the AFT family.

By Akshay Kekuda, Shreeranjani Srirangamsridharan, Ishan Bhatt, Yanan Cao, Sinduja Subramaniam, Evren Korpeoglu, Kaushiki Nag, Kannan Achan