arXiv:2607. 20537v1 Announce Type: cross Abstract: We introduce ReliableTableQA, a framework for training an LLM to annotate the statistical reliability of tabular QA results, not whether the query is answerable, but whether the computed answer is statistically meaningful.
By Huei-Chung Hu, Hsin-Tai Wu, Koyo Kobayashi
arXiv:2609.13288v1 Announce Type: new
Abstract: Video-language models can answer multiple-choice questions with high confidence yet be wrong. We study whether answer-level reliability scores can be i...
By Guoxiang Ren, Rohitash Chandra
The paper investigates how AI systems that perform best‑of‑n search require different validation strategies as the search width changes. It shows that auditing only small search widths leaves a gap in reliability estimates for larger widths, and proposes retaining candidate ranks and truth labels to estimate reliability across all widths up to N. The authors derive theoretical bounds on the minimax mean‑squared error, design procedures that achieve these bounds, and demonstrate that a shared audit can significantly reduce maximum error across many widths in practical CodeRM pools.
By Ricardo Fitas
The paper introduces Counterfactual Fragility Certificates (CFC), a model‑agnostic audit protocol that maps each prediction to an evidence‑failure trajectory, summarizing it with metrics such as greedy flip budget, margin‑collapse area, degradation thresholds, and fragility dominance score. CFC is shown to identify brittle high‑confidence predictions on seven tabular benchmarks with an AUROC of 0.915, outperforming existing scalar scores by up to +0.405. The method remains effective across various perturbation and review‑budget scenarios, and can also inform fragility‑aware regularization and temperature correction.
By Filippo Cenacchi, Longbing Cao, Runze Yang
arXiv:2608. 09768v1 Announce Type: new Abstract: A prediction that is both confident and wrong is a critical reliability failure because it can bypass abstention and human review precisely when the model is mistaken.
By Ange-Cl\'ement Akazan, Ineza Remy Mugenga, Abebe Geletu, Jean Medard Ngnotchouye, Issa Karambal
arXiv:2606. 03650v1 Announce Type: cross Abstract: Choosing or ranking language models for a specific application is hardest when no task-specific labeled data exists, and standard public benchmarks cannot be trusted, their items having likely leaked into pretraining, so scores reflect memorization rather than fitness.
By Alexander Apartsin, Yehudit Aperstein
arXiv:2609.10333v1 Announce Type: new
Abstract: Uncertainty estimation for medical vision--language models (VLMs) using conformal prediction has gained increasing attention due to its distribution-fr...
By Xuan Cuong Ngo, Ngan Le
arXiv:2609.36900v1 Announce Type: new
Abstract: Reward hacking occurs when policy optimization exploits a brittle reward interface or an overly permissive proxy objective, improving the training scor...
By Wan Tian, Zhongyi Li, Xiang Xu, Minhao Zou, Yijie Peng, Fuzhen Zhuang
arXiv:2604.08974v2 Announce Type: replace
Abstract: Uncertainty quantification techniques measure confidence in language model outputs to support critical applications like hallucination detection an...
By Lorenzo Jaime Yu Flores, Cesare Spinoso di-Piano, Jackie Chi Kit Cheung
arXiv:2609.26468v1 Announce Type: new
Abstract: A key factor in deciding whether to trust an automatic prediction is its confidence score, which should be calibrated to match the actual probability o...
By Sophie Henning, Georg Hofmann, Alexander Schulte, Alexander Fraser, Annemarie Friedrich
Safety-Flag is a unified benchmark that consolidates seven popular safety datasets into a single balanced flag/do‑not‑flag protocol, providing item‑level decisions and confidence scores for multiple large language models and dedicated guards. The benchmark evaluates moderator reliability across three dimensions—error direction, probability calibration, and confidence‑based error ranking—revealing that aggregate accuracy masks significant differences, such as one model flagging 85% of benign content while another misses 54% of harmful content. The study shows that general‑purpose models are overconfident, but temperature tuning can substantially improve calibration, and confidence‑based abstention can reduce selective risk, though performance varies with how well confidence ranks errors.
By Yibo Hu
arXiv:2608. 02455v1 Announce Type: new Abstract: Human-centered assessment tasks, which are essential for systematic decision-making, rely heavily on human judgment and typically lack verifiable ground truth.
By Zejun Xie, Xintong Li, Guang Wang, Desheng Zhang