arXiv:2608. 11138v1 Announce Type: cross Abstract: We propose that a model's uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a confident prediction is \emph{fragile} under perturbation of its attention pathways.
By Minsoo Kim, Sungyoung Ji, Kisung Moon, Ilyong Yoon
arXiv:2606.27206v2 Announce Type: replace
Abstract: Garden path sentences present a processing difficulty for humans--- the sentence prefix leads the listener towards one interpretation, until the li...
By Alan Zhou, Milo\v{s} Stanojevi\'c, John T. Hale
The paper investigates integer‑sequence benchmarks from the OEIS by applying a two‑part minimum description length (MDL) learner that searches for P‑recursive recurrences. It finds that MDL difficulty correlates with a combinatorial parameter count, that most sequences fit a recurrence on a prefix but not at full length (the “wilderness” regime), and that language models do not hallucinate in the wilderness but instead hedge, showing that memorisation dominates perceived competence. The study provides a cheap, contamination‑free difficulty signal for OEIS‑derived benchmarks.
By Sabilashan Ganeshan
arXiv:2608.22452v1 Announce Type: new
Abstract: Surprisal, the negative log-probability a language model assigns to a word given its preceding context, reliably predicts adult reading times. Does it...
By Francisco Portillo L\'opez
arXiv:2509. 14704v3 Announce Type: replace Abstract: Benchmark saturation and training-data contamination increasingly obscure whether reported gains in large language models (LLMs) reflect genuine advances in reasoning or familiarity with recurring patterns in benchmark problems.
By Masaharu Mizumoto, Dat Nguyen, Zhiheng Han, Xingfu Li, Yo Nakawake, Le Minh Nguyen
The paper evaluates how three large mixture‑of‑experts models (Alibaba, OpenAI, NVIDIA) can be fine‑tuned to reason in a low‑resource language, specifically Greek. Accuracy metrics show little change, but the authors uncover significant qualitative improvements: after supervised fine‑tuning, models reason in Greek on ~98% of items, with better grammaticality and retained general ability. Reinforcement learning with pre‑registered rewards further eliminates reasoning‑channel leaks and format skips, while the Greek‑reasoning habit remains robust to an accuracy‑only gradient.
By Ayoub Kirouane, Christos Petrocheilos
arXiv:2604.27251v3 Announce Type: replace-cross
Abstract: Large Language Models (LLMs) acquire reasoning capabilities through shared inference patterns in pre-training data, which are further elicite...
By Xingwei Tan, Marco Valentino, Mahmud Elahi Akhter, Yuxiang Zhou, Maria Liakata, Nikolaos Aletras
arXiv:2510. 05969v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed on complex reasoning tasks, yet little is known about their ability to internally evaluate problem difficulty, which is an essential capability for adaptive reasoning and efficient resource allocation.
By Sunbowen Lee, Qingyu Yin, Chak Tou Leong, Jialiang Zhang, Yicheng Gong, Shiwen Ni, Min Yang, Xiaoyu Shen
arXiv:2608. 08447v1 Announce Type: cross Abstract: Multilingual reasoning models are commonly evaluated by whether they arrive at the correct answer, but not by whether they preserve the intended language while reasoning and responding.
By Muhammad Ali Shafique, Kelly Marchisio
arXiv:2607. 11266v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting has significantly advanced the reasoning capabilities of Large Language Models (LLMs), yet it often incurs substantial computational costs due to over-reasoning: the generation of redundant, verbose, or irrelevant steps.
By Daeyeop Lee, Hwanjo Yu
arXiv:2603.18007v2 Announce Type: replace-cross
Abstract: The study explores whether current Large Language Models (LLMs) exhibit Theory of Mind (ToM) capabilities -- specifically, the ability to inf...
By Anna Babarczy, Andras Lukacs, Peter Vedres, Zeteny Bujka
arXiv:2607. 06327v1 Announce Type: cross Abstract: Uncertainty estimation (UE) enables LLM-powered systems to recognize when to abstain, yet existing research has predominantly focused on English.
By Andrea Alfarano, Andrea Bacciu, Saab Mansour, Amin Mantrach, Marcello Federico