arXiv:2606. 08471v1 Announce Type: cross Abstract: Recently, language models have made rapid progress across various domains and applications.
By Marina Igitkhanian, Erik Arakelyan
PROOF is a benchmark that profiles the reliability of object-level facts in instruction-tuned language models by converting a frozen Wikidata snapshot into 18,486 English multiple-choice questions grounded in 11,779 semantic facts across 101 classes, 392 properties, and 14 domains. Each question includes an explicit "I don't know" option, a "No correct option" control, and nine controlled formulations, with 1,849 questions designed as no-correct-option traps. The study evaluates 18 open-weight model deployments on 166,374 prompts, revealing wide variability in factual accuracy, sensitivity to wording changes, and the impact of decoder perturbations.
By Andrei Chetvergov, Mikhail Solovev, Timofei Sivoraksha, Stepan Ukolov, Valeriia Kuschenko, Alexander Evseev, Sergey Bolovtsov
arXiv:2608. 11403v1 Announce Type: new Abstract: Self-consistency (SC) via majority vote is a widely used way to spend inference-time compute: sample N chains of thought, return the plurality answer.
By Utkarsh Bahuguna
arXiv:2608.22048v1 Announce Type: new
Abstract: Large language models are increasingly deployed on local hardware for privacy, cost, and accessibility reasons. Yet many evaluations emphasize accuracy...
By Orion Powers, Daniella Seum, Khaled Slhoub
LogicSkills is a benchmark designed to isolate three core logical abilities in large language models: formal symbolization, countermodel construction, and validity assessment. The dataset draws items from the two-variable fragment of first‑order logic without identity, presented in both English and a Carrollian nonce‑word language, and all instances are solver‑verified with Z3. Results show that conventional instruction‑tuned LLMs excel at validity assessment but struggle with symbolization and countermodel construction, whereas recent reasoning‑tuned models perform well across all tasks, indicating a more systematic logical skill profile.
By Brian Rabern, Philipp Mondorf, Barbara Plank
arXiv:2607. 22554v1 Announce Type: new Abstract: Large language models (LLMs) often achieve strong accuracy on benchmarks, yet it remains unclear how reliably they apply this knowledge when the same question is phrased in different but equivalent ways.
By Kazem Faghih, Yize Cheng, Shoumik Saha, Mobina Pournemat, Armin Gerami, Soheil Feizi