arXiv AI

Spatial Strategies, Not Actions: Vector-Quantized Geodesics as Tools for LLM-Driven Agents

The paper proposes a new architecture for large language model (LLM) agents that enhances spatial understanding by combining geometrical tools with an LLM orchestrator in grid‑world environments. It first gathers geodesic trajectories, vector‑quantizes them to create a representative subset, and then has the LLM label each trajectory with a natural language description, turning them into reusable tools. During operation, the LLM selects the appropriate tool based on the current state and goal, while low‑level control executes the chosen trajectory, enabling efficient decision‑making in a partially observable 2D grid setting.

arXiv AI
Jun 12

SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning

arXiv:2606. 13673v1 Announce Type: cross Abstract: Spatial reasoning, the ability to determine where objects are, how they relate, and how they move in 3D, remains a fundamental challenge for vision-language models (VLMs).

By Seokju Cho, Ryo Hachiuma, Abhishek Badki, Hang Su, Byung-Kwan Lee, Chan Hee Song, Sifei Liu, Subhashree Radhakrishnan, Seungryong Kim, Yu-Chiang Frank Wang, Min-Hung Chen
arXiv AI
Jun 9

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks

arXiv:2606. 09669v1 Announce Type: new Abstract: Spatial reasoning is a foundational capability for multimodal large language models (MLLMs) to perceive and operate within the physical world.

By Hongcheng Gao, Hailong Qu, Jingyi Tang, Jiahao Wang, Zihao Huang, Hengkang Qiao, Shihong Huang, Junming Yang, Yi Li, Hongyixuan Yuan, Wenjie Li, Bohan Zeng, Wenbo Li, Bo Wang, Jianhui Liu, Olive Huang, Haoyang Huang, Wentao Zhang, Guoqing Huang, Nan Duan, Yinpeng Dong
arXiv AI
Sep 21

AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents

AgentVidBench is a new multi‑hop video question‑answering benchmark designed to evaluate spatial, temporal, and causal reasoning in multimodal large language models (MLLMs). Unlike existing tests that focus on simple scene queries or global summaries, AgentVidBench includes step‑by‑step solution traces to assess whether agents gather the necessary evidence to justify their answers. Experiments with 12 MLLMs show limited single‑turn performance, but integrating these models into agentic workflows improves both accuracy and trajectory scores, establishing AgentVidBench as a comprehensive testbed for future research on agentic video understanding.

By Seoyeon An, Hyeonseo Jang, Minsu Kim, Chanho Lee, Younghan Park, Kangwook Lee
arXiv AI
Jun 2

TrafficClaw: A Generalizable LLM Agent in the Unified Physical Environment for Urban Traffic Control

arXiv:2604. 17456v2 Announce Type: replace Abstract: Large language model (LLM) agents have shown strong capabilities in long-horizon reasoning, tool use, and decision-making in digital environments, yet extending them to physically grounded systems remains challenging.

By Siqi Lai, Pan Zhang, Yuping Zhou, Jindong Han, Yansong Ning, Hao Liu
arXiv AI
Sep 7

MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model

MultihopSpatial is a new benchmark for Vision‑Language Models that focuses on multi‑hop, compositional spatial reasoning with queries ranging from 1 to 3 hops across varied spatial perspectives. It introduces the Acc@50IoU metric, which jointly evaluates answer selection and precise bounding‑box prediction, and provides a large‑scale training corpus, MultihopSpatial‑Train, to improve spatial intelligence. Evaluation of 37 state‑of‑the‑art VLMs shows that compositional spatial reasoning remains a significant challenge, and reinforcement learning fine‑tuning on the corpus boosts both intrinsic spatial reasoning and downstream embodied manipulation performance.

By Youngwan Lee, Soojin Jang, Yoorhim Cho, Seunghwan Lee, Yong-Ju Lee, Sung Ju Hwang