arXiv:2512.00729v2 Announce Type: replace
Abstract: Motivated by the observed human-like behaviours in Large Reasoning Models (LRMs), this paper introduces a comprehensive taxonomy to characterise at...
By Yuxiang Chen, Zuohan Wu, Ziwei Wang, Xiangning Yu, Xujia Li, Linyi Yang, Mengyue Yang, Jun Wang, Lei Chen
The paper critiques current systematic generalization benchmarks for oversimplifying the task by relying on linear action composition, productivity-based tests, and action-explicit goals. It introduces TranSGrid, a new testbed that integrates deductive, inductive, and abductive reasoning in a unified setting. Experiments with seven Transformers show a significant performance drop on TranSGrid compared to standard held-out tests, indicating that existing simplifications mask the true difficulty of systematic generalization.
By Chengwen Qi, Deheng Ye, Yatao Bian
The paper introduces a multi-stage rule‑chaining framework for the Abstraction and Reasoning Corpus (ARC), aiming to model cognitive generalization by inferring abstract rules from few examples. It combines three solvers—a deterministic rule discovery module, a pattern‑composition engine, and a structural abstraction layer—executed sequentially in a fallback hierarchy that reuses earlier reasoning traces to improve interpretability and generalization. The system achieved over 95% accuracy on ARC tasks, demonstrating strong performance across deterministic, compositional, and abstract categories.
By Deblina Kar
arXiv:2506. 21571v3 Announce Type: replace-cross Abstract: Large Reasoning Models (LRMs), which autonomously produce a reasoning Chain of Thought (CoT) before producing final responses, offer a promising approach to interpreting and monitoring model behaviors.
By Jianshuo Dong, Yujia Fu, Chuanrui Hu, Chao Zhang, Han Qiu
When large language models (LLMs) fail to generalize or make haphazard errors in reasoning, it is often taken as evidence that LLMs are not truly reasoning, but rather performing a kind of pattern matching. The implication is that people's behavior does not exhibit the same types of failures because human reasoning uses principled and abstract world models.
arXiv:2606. 01462v1 Announce Type: new Abstract: Studies of human reasoning have shown that people are typically stronger at evaluating reasoning than producing it from scratch.
By Mingzhong Sun, Teresa Yeo, Armando Solar-Lezama, Tan Zhi-Xuan