arXiv:2607.14109v2 Announce Type: replace
Abstract: Probing the capabilities of Large Language Models (LLMs) and building robust solutions for Multiple-Choice Question Answering (MCQA) remain central...
By Inder Preet, Shuxin Lin, Dhaval Patel
arXiv:2509. 14704v3 Announce Type: replace Abstract: Benchmark saturation and training-data contamination increasingly obscure whether reported gains in large language models (LLMs) reflect genuine advances in reasoning or familiarity with recurring patterns in benchmark problems.
By Masaharu Mizumoto, Dat Nguyen, Zhiheng Han, Xingfu Li, Yo Nakawake, Le Minh Nguyen
arXiv:2608. 16627v1 Announce Type: cross Abstract: Natural language explanations (NLEs) are increasingly used as inputs, for example, as few-shot rationales that influence model behavior in in-context learning (ICL).
By Mahdi Dhaini, Adam Dejl, Juraj Vladika, Volkan \"Ozer, Barbara Plank, Gjergji Kasneci
arXiv:2608. 03550v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting remains the standard baseline for evaluating models' reasoning abilities.
By Denys Pushkin, Albert Q. Jiang, Aryo Lotfi, Colin Sandon, Emmanuel Abb\'e
arXiv:2607. 06327v1 Announce Type: cross Abstract: Uncertainty estimation (UE) enables LLM-powered systems to recognize when to abstain, yet existing research has predominantly focused on English.
By Andrea Alfarano, Andrea Bacciu, Saab Mansour, Amin Mantrach, Marcello Federico
arXiv:2607. 11266v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting has significantly advanced the reasoning capabilities of Large Language Models (LLMs), yet it often incurs substantial computational costs due to over-reasoning: the generation of redundant, verbose, or irrelevant steps.
By Daeyeop Lee, Hwanjo Yu
arXiv:2609.24650v1 Announce Type: new
Abstract: Readability assessment is essential for tailoring texts to intended audiences across educational, healthcare, and information retrieval domains. Howeve...
By Rapha\"el Thieffry, Matej Martinc
arXiv:2501. 11790v5 Announce Type: replace-cross Abstract: Recent studies have raised significant concerns regarding the reliability of current mathematics benchmarks, highlighting issues such as simplistic design and potential data contamination.
By Zijin Hong, Hao Wu, Su Dong, Junnan Dong, Yilin Xiao, Yujing Zhang, Zhu Wang, Feiran Huang, Linyi Li, Hongxia Yang, Xiao Huang
PROOF is a benchmark that profiles the reliability of object-level facts in instruction-tuned language models by converting a frozen Wikidata snapshot into 18,486 English multiple-choice questions grounded in 11,779 semantic facts across 101 classes, 392 properties, and 14 domains. Each question includes an explicit "I don't know" option, a "No correct option" control, and nine controlled formulations, with 1,849 questions designed as no-correct-option traps. The study evaluates 18 open-weight model deployments on 166,374 prompts, revealing wide variability in factual accuracy, sensitivity to wording changes, and the impact of decoder perturbations.
By Andrei Chetvergov, Mikhail Solovev, Timofei Sivoraksha, Stepan Ukolov, Valeriia Kuschenko, Alexander Evseev, Sergey Bolovtsov
arXiv:2602. 23971v4 Announce Type: replace-cross Abstract: Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an alignment failure, particularly in high-stakes advisory and social contexts.
By Magda Dubois, Cozmin Ududec, Christopher Summerfield, Lennart Luettgau
arXiv:2606. 03304v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly evaluated in multilingual settings, yet their inference behavior in low-resource African languages remains underexplored especially under pure prompting without fine-tuning.
By Anuj Tiwari, Terry Oko-odion, Hannah Nwokocha
arXiv:2607.28077v2 Announce Type: replace
Abstract: Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models, but prompt groups with identica...
By Shuang Liang, Haoyang Zhou, Yifan Gong, Guowei Wang, Xiting Wang