arXiv:2606. 13221v2 Announce Type: replace Abstract: Evaluating new large language models typically requires costly human annotation campaigns at scale.
By Bora Kargi, David Salinas
The paper argues that calibration—how well a language model’s confidence aligns with its actual correctness—should be a standard evaluation metric for large language models (LLMs). It notes that while calibration metrics exist, they are rarely applied outside specialized NLP subfields, leading to unverified confidence scores in new models, datasets, and benchmarks. The authors highlight the risks of miscalibration both at deployment (overconfident errors causing harm) and in research workflows (affecting LLM-as-a-judge, synthetic data generation, and active learning). They call for every NLP subfield to pair its primary performance metric with a calibration score, treating calibration as an essential property of every model.
By Mario Sanz-Guerrero, Katharina von der Wense
arXiv:2605. 15416v2 Announce Type: replace-cross Abstract: Jung et al.
By Gaojie Jin, Yong Tao, Lijia Yu, Tianjin Huang
arXiv:2608. 19323v1 Announce Type: cross Abstract: Uncertainty quantification (UQ) is essential for the safe deployment of large language models (LLMs).
By Sokhna Diarra Mbacke, Mouloud Belbahri, Gabriel Loaiza-Ganem
arXiv:2504.18346v4 Announce Type: replace-cross
Abstract: Large Language Models (LLMs) have been transformative across many domains. However, hallucination, i.e., confidently outputting incorrect inf...
By Toghrul Abbasli, Kentaroh Toyoda, Yuan Wang, Leon Witt, Muhammad Asif Ali, Yukai Miao, Dan Li, Qingsong Wei
arXiv:2607. 20526v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in settings where fluent but incorrect answers can be costly.
By Matthew ffrench-Constant, Daniel Yang, Xinmeng Huang, Sanyam Kapoor