arXiv:2609.26468v1 Announce Type: new
Abstract: A key factor in deciding whether to trust an automatic prediction is its confidence score, which should be calibrated to match the actual probability o...
By Sophie Henning, Georg Hofmann, Alexander Schulte, Alexander Fraser, Annemarie Friedrich
A key factor in deciding whether to trust an automatic prediction is its confidence score, which should be calibrated to match the actual probability of the prediction being correct. Most confidence c...
arXiv:2607. 05628v1 Announce Type: cross Abstract: Accurate and efficient classification of thoracic diseases in chest X-ray (CXR) images is crucial for timely diagnosis and treatment.
By Mohammad S. Majdi, Jeffrey J. Rodriguez
arXiv:2112. 02353v3 Announce Type: replace-cross Abstract: Hierarchical classification aims to sort the object into a hierarchical structure of categories.
By Renzhen Wang, De cai, Kaiwen Xiao, Xixi Jia, Xiao Han, Deyu Meng
arXiv:2606. 03245v1 Announce Type: cross Abstract: Concepts of calibration formalize the compatibility between probabilistic predictions and the respective outcomes.
By Johannes Resin, Lu Yang, Tilmann Gneiting
arXiv:2602. 08986v2 Announce Type: replace-cross Abstract: In hierarchical multi-label classification, a persistent challenge is enabling model predictions to reach deeper levels of the hierarchy for more detailed or fine-grained classifications.
By Isaac Xu, Martin Gillis, Ayushi Sharma, Benjamin Misiuk, Craig J. Brown, Thomas Trappenberg
The paper introduces the concept of observational multiplicity, where multiple probabilistic classifiers can perform similarly yet produce conflicting predictions, undermining interpretability and safety. It proposes measuring this arbitrariness through a regret metric that captures how predictions could shift with different training labels. The authors present a general method to estimate regret, show it varies across dataset groups, and discuss its use for safety via abstention and targeted data collection.
By Erin George, Deanna Needell, Berk Ustun
arXiv:2607. 20641v1 Announce Type: new Abstract: Federated learning (FL) enables multiple clinical institutions to collaboratively train a shared disease classifier without centralizing patient data.
By Afsaneh Mahanipour, Hana Khamfroush
arXiv:2606. 28598v1 Announce Type: cross Abstract: Prediction sets should have high coverage to be useful, but some coverage notions are more practically relevant than others.
By Aabesh Bhattacharyya, Tiffany Ding, Rina Foygel Barber
arXiv:2607. 11954v1 Announce Type: cross Abstract: Clinical code prediction maps unstructured discharge summaries to ICD-10-CM leaf codes in a large, sparse, and deeply hierarchical label space.
By Amritpal Singh, Sebastian Torres, Khawar Shakeel, Syed Ahmad Chan Bukhari
The paper extends ex‑ante evaluation of Predict‑Then‑Optimize methods from binary to multiclass classification by simulating predictions at specified performance levels and mapping prediction errors to decision regret. It introduces a first‑order approximation that estimates regret from individual misclassifications, reducing computational effort. Experiments show the simulation accurately reproduces target performance and that the approximation is close for some problems, though it falters when simultaneous misclassifications interact significantly.
By Pieter Smet
arXiv:2604. 04241v2 Announce Type: replace Abstract: Risk scoring systems are widely used in high-stakes domains to assist decision-making.
By Wenhao Chi, \c{S}. \.Ilker Birbil