arXiv:2212. 05260v4 Announce Type: replace-cross Abstract: Proper scoring rules encourage probabilistic predictions that match the true underlying distribution and are central to model evaluation, with increasing relevance in automated workflows such as AutoML.
By John Zobolas, Raphael Sonabend, Riccardo De Bin, Johannes Piller, Philipp Kopper, Lukas Burk, Andreas Bender
arXiv:2506. 02075v3 Announce Type: replace-cross Abstract: The current state of evaluation in survival analysis is plagued by the persistent use of evaluation metrics in ways that are misaligned with the stated modeling objective.
By Christian Marius Lillelund, Shi-ang Qi, Russell Greiner, Christian Fischer Pedersen
arXiv:2608. 08202v1 Announce Type: new Abstract: Data-centric curation pipelines frequently rely on model confidence scores to flag and filter noisy or mislabeled training instances.
By Sai Srikar Boddupalli
arXiv:2606. 18479v1 Announce Type: new Abstract: Reject inference methods are widely used to mitigate survival bias in credit scoring, yet their effectiveness remains poorly understood.
By Bruno Scarone, Ricardo Baeza-Yates
arXiv:2607. 10466v1 Announce Type: new Abstract: Survival models can model time-to-event outcomes using partially observed data.
By Yanqi Xu, Hui Dai, Carlos Fernandez-Granda, Krzysztof J. Geras, Yiqiu Shen
arXiv:2608. 02665v1 Announce Type: cross Abstract: A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form.
By Yongxi Zhou, Junwei Yao, Yuanzhe Liu, Zihan Dong, Wenbo Ye, Jiaxi Wen, Lai Yun Choi
The paper introduces Counterfactual Fragility Certificates (CFC), a model‑agnostic audit protocol that maps each prediction to an evidence‑failure trajectory, summarizing it with metrics such as greedy flip budget, margin‑collapse area, degradation thresholds, and fragility dominance score. CFC is shown to identify brittle high‑confidence predictions on seven tabular benchmarks with an AUROC of 0.915, outperforming existing scalar scores by up to +0.405. The method remains effective across various perturbation and review‑budget scenarios, and can also inform fragility‑aware regularization and temperature correction.
By Filippo Cenacchi, Longbing Cao, Runze Yang
arXiv:2605. 12895v2 Announce Type: replace-cross Abstract: Clinical decision-support systems are expert systems whose recommendations clinicians act on directly, yet they are usually cleared on one aggregate accuracy number from a held-out test set.
By Rohith Reddy Bellibatlu, Manpreet Singh, Yash Jajoo, Shyamal Lakhanpal, Abhishek Israni
arXiv:2608. 14617v1 Announce Type: cross Abstract: A recurring proposal in legal AI is to improve case-outcome prediction by fusing uncertainty tools (evidence graphs with belief propagation, sequential Bayesian odds updating, Dempster-Shafer combination, and conformal prediction) into one pipeline.
By Surya Saka
arXiv:2601. 20533v2 Announce Type: replace-cross Abstract: Survival analysis has become a standard approach for modelling time to default by time-varying covariates in credit risk.
By Jianwei Peng (Humboldt-Universit\"at zu Berlin), Stefan Lessmann (Humboldt-Universit\"at zu Berlin, Bucharest University of Economic Studies)
The paper proposes replacing multiple horizon‑specific binary classifiers with a single survival model to predict time‑to‑repurchase in grocery e‑commerce. Empirical analysis shows a slightly decreasing hazard (k≈0.9) and that a Log‑Normal model best fits marginal distributions while Weibull best fits residuals. A single Accelerated Failure Time (AFT) model matches or surpasses per‑horizon classifiers with fewer trees, and a 4‑parameter calibration maps survival CDFs to horizon probabilities without monotonicity violations, revealing a trade‑off between calibration and ranking within the AFT family.
By Akshay Kekuda, Shreeranjani Srirangamsridharan, Ishan Bhatt, Yanan Cao, Sinduja Subramaniam, Evren Korpeoglu, Kaushiki Nag, Kannan Achan
arXiv:2604. 08870v3 Announce Type: replace-cross Abstract: Student dropout is a persistent concern in Learning Analytics, yet comparative studies frequently evaluate predictive models under heterogeneous protocols, prioritizing discrimination over temporal interpretability and calibration.
By Rafael da Silva, Jeff Eicher, Gregory Longo