arXiv:2609.26468v1 Announce Type: new
Abstract: A key factor in deciding whether to trust an automatic prediction is its confidence score, which should be calibrated to match the actual probability o...
By Sophie Henning, Georg Hofmann, Alexander Schulte, Alexander Fraser, Annemarie Friedrich
arXiv:2607. 08299v2 Announce Type: replace Abstract: Diagnostic decision making often relies on a sequence of pathology tests that bridge patient symptoms and final disease diagnosis.
By Abu Rafe Md Jamil, Nayan Malakar
arXiv:2609.36532v1 Announce Type: cross
Abstract: In multiclass probabilistic prediction, Utility Calibration (UC), which focuses auditing on specified utilities, has recently received attention as a...
By Futoshi Futami, Jerry Huang, Ichiro Takeuchi
arXiv:2607. 18465v1 Announce Type: new Abstract: Crowdsourced labeling provides valuable labeled data for domains across natural language processing, computer vision, and video.
By Ju Chen, Sijia Xu, Jun Feng, Zhiqiang Gao, Zhengyi Yang
Crowdsourced labeling provides valuable labeled data for domains across natural language processing, computer vision, and video. Label aggregation aims to infer latent true labels from noisy and biased annotations, with the key lying in annotator reliability estimation.
arXiv:2606. 15029v1 Announce Type: new Abstract: LLM judges are used to reduce the need for costly human labor in evaluating open-ended text generation.
By Alyssa Unell, Natalie Dullerud, Naomi Boneh, Meena Jagadeesan, Tatsu Hashimoto, Nigam Shah, Sanmi Koyejo
arXiv:2609.38705v1 Announce Type: new
Abstract: Foundation models for computational pathology are usually evaluated using AUC and accuracy, while calibration is often left untested. This matters beca...
By Wenjun Liu, Saeed Hassanpour
arXiv:2609.24303v1 Announce Type: new
Abstract: Post-training often improves task performance but can degrade confidence calibration, leaving post-trained language models (PoLMs) more overconfident t...
By Linhan Luo, Lequan Lin, Dai Shi, Feng Chen, Jos\'e Miguel Hern\'andez-Lobato, Junbin Gao
The paper introduces the concept of observational multiplicity, where multiple probabilistic classifiers can perform similarly yet produce conflicting predictions, undermining interpretability and safety. It proposes measuring this arbitrariness through a regret metric that captures how predictions could shift with different training labels. The authors present a general method to estimate regret, show it varies across dataset groups, and discuss its use for safety via abstention and targeted data collection.
By Erin George, Deanna Needell, Berk Ustun
The paper introduces a Calibrated Reflection approach to improve confidence estimation in Large Language Models (LLMs). It combines structured reasoning with a distance‑aware calibration technique, featuring a Maximum Confidence Selection method, a reflection‑based prompting mechanism, and an ordinal‑aware calibration strategy. Experiments on datasets such as HelpSteer2, Llama T‑REx, and a proprietary conversational set show the method works for both conversational and fact‑based classification tasks.
By Umesh Bodhwani, Yuan Ling, Shujing Dong, Yarong Feng, Hongfei Li, Ayush Goyal
arXiv:2512. 17788v2 Announce Type: replace Abstract: Multi-instance partial-label learning (MIPL) is a weakly supervised framework that extends the principles of multi-instance learning (MIL) and partial-label learning (PLL) to address the challenges of inexact supervision in both instance and label spaces.
By Wei Tang, Yin-Fang Yang, Weijia Zhang, Min-Ling Zhang
arXiv:2608. 10406v1 Announce Type: cross Abstract: Web search, product search, and question-answering retrieval systems often assign a relevance label and confidence score to each query-candidate pair.
By Inwoo Tae, Yongjae Lee