arXiv:2606. 13607v1 Announce Type: new Abstract: When large language models (LLMs) fail to generalize or make haphazard errors in reasoning, it is often taken as evidence that LLMs are not truly reasoning, but rather performing a kind of pattern matching.
By Zach Studdiford, Gary Lupyan
The paper investigates whether current reasoning models exhibit systematicity—the idea that understanding one concept should extend to closely related variations—by extending rule induction tasks from cognitive science. Using task isomorphisms like recombination and substitution, the authors generate structurally equivalent task variants and test models on them. Results show that while models can solve the original tasks, they frequently fail on these equivalent variants, indicating a lack of systematicity in their reasoning abilities.
By Simon Schug, Brenden M. Lake
arXiv:2512.00729v2 Announce Type: replace
Abstract: Motivated by the observed human-like behaviours in Large Reasoning Models (LRMs), this paper introduces a comprehensive taxonomy to characterise at...
By Yuxiang Chen, Zuohan Wu, Ziwei Wang, Xiangning Yu, Xujia Li, Linyi Yang, Mengyue Yang, Jun Wang, Lei Chen
The paper investigates whether reasoning representations—explanations for large language model outputs—aid humans in evaluating those outputs. A controlled human study tested six reasoning formats across tasks of varying complexity, measuring structural understanding, error detection, and trust calibration. Results revealed a mismatch: participants favored planning- and decomposition-based representations, yet simpler chain-of-thought traces better supported verification, trust, and interpretability, while preferred formats increased calibration risks.
By Jaewoo Lim, Sungbok Shin, Sanghyun Hong
arXiv:2608. 14161v1 Announce Type: new Abstract: LLMs exhibit social biases that can produce inaccurate and discriminatory inferences, posing risks in high-stakes applications.
By Varsha Ramineni, Hossein A. Rahmani, Jerome Ramos, Karin Sevegnani, Emine Yilmaz
arXiv:2506. 21571v3 Announce Type: replace-cross Abstract: Large Reasoning Models (LRMs), which autonomously produce a reasoning Chain of Thought (CoT) before producing final responses, offer a promising approach to interpreting and monitoring model behaviors.
By Jianshuo Dong, Yujia Fu, Chuanrui Hu, Chao Zhang, Han Qiu