arXiv AI By Xinyan Jiang, Ninghao Liu, Di Wang, Lijie Hu

Beyond Scalars: Evaluating and Understanding LLM Reasoning via Geometric Progress and Stability

Read the original on arXiv AI →

arXiv:2603. 10384v3 Announce Type: replace Abstract: Evaluating LLM reliability via scalar probabilities often fails to capture the structural dynamics of reasoning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 11

MindTopo: Can Foundation Models Reason in Topological Space?

MindTopo is a benchmark that tests foundation models on topological reasoning, covering five cognitive properties—continuity, separation, order, enclosure, and knots—across two cognitive levels: reasoning and planning. It contains 11,030 instances from 13 procedurally generated task types, and evaluates 14 multimodal large language models, including agent configurations with image and video generation. Results show that models perform better on reasoning than planning, and even the best model lags far behind human performance, with fine‑tuning and reinforcement learning improving reasoning more than planning.

By Yunfei Ge, Anbang Liu, Qineng Wang, Johnalbert Garnica, Jianwen Lyu, Zihan Wang, Reuben Tan, Jianfeng Gao, Ruohan Zhang, Yining Hong, Jiajun Wu, Manling Li
arXiv AI
Sep 17

Lost in Perception: Isolating Perceptual and Reasoning Failures in Multimodal Physics and Geometry Reasoning

The paper introduces a five-task diagnostic experiment that separates perceptual and reasoning failures in multimodal large language models on physics and geometry benchmarks. It finds that misinterpreting diagrams hurts performance even on text-only solvable problems, and that accuracy improves when models receive human-authored captions. The study shows that correcting captions can recover many errors, revealing distinct reasoning bottlenecks that differ by domain, while a heavily pretrained model still underperforms and often truncates reasoning traces.

By Raj Jaiswal, Sree Krishna Uppalapati, Dhruvkumar Patel, Ria Khatoniar, Tanuja Ganu, Rajiv Ratn Shah
Hugging Face Trending Papers
Aug 5

Disentangling 3D Modeling from Spatial Reasoning

In this work, we explore an alternative paradigm for spatial reasoning by explicitly disentangling 3D perception from reasoning, rather than jointly acquiring implicit 3D perception and reasoning through large-scale training. Our key observation is that modern perception models excel at estimating continuous 3D geometry, whereas large language models (LLMs) are particularly effective at compositional and symbolic reasoning.