arXiv AI By Kazem Faghih, Yize Cheng, Shoumik Saha, Mobina Pournemat, Armin Gerami, Soheil Feizi

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy

Read the original on arXiv AI →

arXiv:2607. 22554v1 Announce Type: new Abstract: Large language models (LLMs) often achieve strong accuracy on benchmarks, yet it remains unclear how reliably they apply this knowledge when the same question is phrased in different but equivalent ways.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 15

How Semantically Stable Are LLM Refusals? Measuring Confusion in Local Safety Boundaries

The paper introduces Semantic Confusion to assess how consistently large language models refuse similar prompts. It presents ParaGuard, a 10k‑prompt corpus of controlled paraphrase clusters, and proposes three token‑level metrics—Confusion Index, Confusion Rate, and Confusion Depth—to measure contradictory refusal decisions across meaning‑preserving paraphrases. Experiments show that global false rejection rates can mask local inconsistencies, revealing that refusal evaluation must consider both frequency and consistency across nearby paraphrases.

By Riad Ahmed Anonto, Md Labid Al Nahiyan, Md Tanvir Hassan
arXiv AI
Sep 25

PROOF: Profiling Reliability of Object-Level Facts in Large Language Models

PROOF is a benchmark that profiles the reliability of object-level facts in instruction-tuned language models by converting a frozen Wikidata snapshot into 18,486 English multiple-choice questions grounded in 11,779 semantic facts across 101 classes, 392 properties, and 14 domains. Each question includes an explicit "I don't know" option, a "No correct option" control, and nine controlled formulations, with 1,849 questions designed as no-correct-option traps. The study evaluates 18 open-weight model deployments on 166,374 prompts, revealing wide variability in factual accuracy, sensitivity to wording changes, and the impact of decoder perturbations.

By Andrei Chetvergov, Mikhail Solovev, Timofei Sivoraksha, Stepan Ukolov, Valeriia Kuschenko, Alexander Evseev, Sergey Bolovtsov
arXiv AI
2d ago

Certainty Is Not Just Correctness: Rethinking Token-Level Certainty in LLM Reasoning

The paper investigates token‑level certainty as a proxy for correctness in large language models. It finds that certainty better predicts whether a model will answer a question correctly than it does whether a specific response is correct, and that certainty varies by token type and position. The authors show that using certainty early in generation to allocate responses and later to weight votes improves accuracy while dramatically cutting token cost.

By Yunfan Zhou, Ye Zhu, Zhihai Wang, Jianguo Yao, Haibing Guan, Xijun Li