arXiv AI By Jos\'e Luciano Ver\c{c}osa Marques, Frederico Jorge Heitmann, Daniel Omar Perez, Reinaldo Cesar, Marcelo Vinicius de Paula, T\'arcio Andr\'e dos Santos Barros

Technical Manual for Toolkit for Confidence-Corpus Consistency via Fine-Tuning on a Fabricated Corpus

Read the original on arXiv AI →

The article presents a technical manual for an open toolkit designed to evaluate how a language model’s confidence reflects its factual knowledge. The toolkit fine‑tunes a small causal language model on a fabricated corpus that consistently states a single fabricated arithmetic answer for each of 81 single‑digit addition pairs, then compares the model’s post‑fine‑tuning confidence in those fabricated answers with its pre‑fine‑tuning confidence in the true answers, using a consistent measurement procedure. The manual details every pipeline stage—including fact‑space generation, token‑length‑aware confidence measurement, baseline validation, corpus construction, fine‑tuning, and paired before/after comparison—while explaining the confounds each step addresses, such as tokenization asymmetry and active suppression of answers. "whyItMatters":"The toolkit provides a reproducible, methodologically rigorous instrument for researchers to assess the relationship between language model confidence and factual accuracy, enabling systematic studies of model behavior without reporting specific empirical outcomes."

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 18

An Analysis of Training-Free Self-Reported Confidence in Language Models

The paper investigates whether language models’ self-reported confidence is meaningful without additional training. By evaluating three training‑free signals—direct verbalization, post‑hoc probability estimates, and agreement across multiple generations—on 100 TriviaQA questions, the authors find that direct verbalization alone achieves high AUROC scores (0.956 and 0.937) for correctness prediction, while agreement-based methods perform noticeably worse. Re‑eliciting confidence for the same answers shows modest score shifts and occasional decision flips, and an audit of biography claims reveals only a small confidence gap between supported and contradicted statements.

By Lukas Meyer, Sofia Rossi, Wei Chen, Thomas Laurent, Yiming Li
arXiv Computation and Language
Sep 25

Calibration Is Not Enough: Evaluating Confidence Estimation Under Language Variations

The paper introduces a new evaluation framework for confidence estimation in large language models, focusing on three properties: robustness to prompt changes, stability across semantically equivalent answers, and sensitivity to semantically different answers. It demonstrates that existing confidence estimation methods perform well on robustness and stability but often fail to detect differences in answer meaning, revealing gaps in current evaluation practices. The framework aims to guide the selection of confidence estimators for practical applications.

By Yuxi Xia, Dennis Ulmer, Terra Blevins, Yihong Liu, Hinrich Sch\"utze, Benjamin Roth
arXiv Machine Learning
Sep 21

Semantic Calibration Prevails Where Token Confidence Fails: Benchmarking Long-Form Scientific QA

The paper presents the first large‑scale benchmark for uncertainty quantification (UQ) calibration in long‑form scientific question answering, evaluating four UQ methods on 685,000 responses from up to 20 large language models across seven datasets. It shows that instruction tuning leads to token‑level probability polarization, undermining token‑level uncertainty signals, while reasoning model families differ in how they handle this effect. Only semantic consistency—consistency of the final answer—provides well‑calibrated outputs, demonstrating that semantic calibration remains robust in multi‑step, dependency‑rich reasoning.

By Philip M\"uller, Nicholas Popovi\v{c}, Michael F\"arber, Peter Steinbach
arXiv Machine Learning
Aug 5

M-GATE: Multilingual Grammar, Accuracy in Translation, and Efficiency Benchmark for Large Language Models

arXiv:2608. 03803v1 Announce Type: cross Abstract: Multilingual language models are deployed across a hundred or more languages, yet most benchmarks test whether a model can perform a task _in_ a language rather than whether it commands the language itself, conflating fluency with proficiency.

By Tom\'a\v{s} Burkert, Angelika Peljak-{\L}api\'nska, David Zelen\'y