MindTopo is a benchmark that tests foundation models on topological reasoning, covering five cognitive properties—continuity, separation, order, enclosure, and knots—across two cognitive levels: reasoning and planning. It contains 11,030 instances from 13 procedurally generated task types, and evaluates 14 multimodal large language models, including agent configurations with image and video generation. Results show that models perform better on reasoning than planning, and even the best model lags far behind human performance, with fine‑tuning and reinforcement learning improving reasoning more than planning.
By Yunfei Ge, Anbang Liu, Qineng Wang, Johnalbert Garnica, Jianwen Lyu, Zihan Wang, Reuben Tan, Jianfeng Gao, Ruohan Zhang, Yining Hong, Jiajun Wu, Manling Li
arXiv:2606. 29278v1 Announce Type: new Abstract: We introduce the Complexity Ceiling Benchmark (CCB), a controlled evaluation of how language-model reasoning decays as the number of required sequential steps grows.
By Shubh Chapra, Dhruv Kumar, Murari Mandal, Yash Sinha
The paper introduces a five-task diagnostic experiment that separates perceptual and reasoning failures in multimodal large language models on physics and geometry benchmarks. It finds that misinterpreting diagrams hurts performance even on text-only solvable problems, and that accuracy improves when models receive human-authored captions. The study shows that correcting captions can recover many errors, revealing distinct reasoning bottlenecks that differ by domain, while a heavily pretrained model still underperforms and often truncates reasoning traces.
By Raj Jaiswal, Sree Krishna Uppalapati, Dhruvkumar Patel, Ria Khatoniar, Tanuja Ganu, Rajiv Ratn Shah
In this work, we explore an alternative paradigm for spatial reasoning by explicitly disentangling 3D perception from reasoning, rather than jointly acquiring implicit 3D perception and reasoning through large-scale training. Our key observation is that modern perception models excel at estimating continuous 3D geometry, whereas large language models (LLMs) are particularly effective at compositional and symbolic reasoning.
arXiv:2609.09030v1 Announce Type: new
Abstract: Chain-of-thought reasoning provides a structured computation between a model's input and final answer. Yet it is often evaluated through endpoint accur...
By Mar Gonz\`alez I Catal\`a, Haitz S\'aez de Oc\'ariz Borde, Davide Murari, Carola-Bibiane Sch\"onlieb, Pietro Li\`o, George Monta\~nez
arXiv:2608. 05242v1 Announce Type: new Abstract: In this work, we explore an alternative paradigm for spatial reasoning by explicitly disentangling 3D perception from reasoning, rather than jointly acquiring implicit 3D perception and reasoning through large-scale training.
By Haoze Sun, Jiequan Cui, Qingshan Xu, Richang Hong