Clearing the Fog: Towards Installing and Refining Proactive Exploration Capabilities in LLM Agents
arXiv:2608. 14339v1 Announce Type: new Abstract: We study proactive exploration in LLM agents, i.
The paper introduces the ACE framework, which uses spatially explicit inspection to guide embodied exploration. By combining evidence‑grounded perception with exposure‑informed movement, ACE provides a spatially resolved decision paradigm that improves cue assessment and movement direction. Experiments show ACE boosts navigation task success by 18.0% and exploration efficiency by 10.3% over previous baselines.
arXiv:2608. 14339v1 Announce Type: new Abstract: We study proactive exploration in LLM agents, i.
arXiv:2609.36906v1 Announce Type: new Abstract: Reliable embodied decisions under partial observability require informative observations and sufficient supporting evidence. However, semantic scores a...
VABench is a benchmark that tests general‑purpose multimodal large language models (MLLMs) on embodied spatial intelligence by requiring them to observe, reason, act, and revise based on visual demonstrations and active perception. The benchmark includes 14 task families, a fixed model‑agnostic controller, and evaluates models on target localization, spatial relations, and long‑horizon composition tracks without providing privileged object poses or learned action heads. Results show that while the best model achieves perfect target localization, overall task success remains modest, and active camera control and geometric transfer significantly influence performance.
arXiv:2608. 08077v1 Announce Type: new Abstract: Theory of Space framework (ToS) assesses the spatial understanding of curiosity-driven Vision-Language Models (VLMs) under partial observability.
The paper introduces GTA‑VLA, an interactive Vision‑Language‑Action framework that lets users guide robot policies with explicit visual cues such as affordance points, boxes, and traces. Unlike traditional direct sense‑to‑act models, GTA‑VLA incorporates a spatial‑visual Chain‑of‑Thought that blends human guidance with internal task planning, and couples this reasoning module with a lightweight reactive action head for efficient execution. Experiments on the SimplerEnv WidowX benchmark show a state‑of‑the‑art 81.2 % success rate, and the framework significantly improves task success under out‑of‑domain visual shifts and spatial ambiguities, demonstrating the benefit of interactive reasoning for failure recovery in embodied control.
SeekVLN is a new framework for Vision‑Language Navigation that addresses the problem of agents acting on insufficient evidence, termed Progress Myopia. It combines semantic progress reasoning with active evidence seeking, trained first with Future‑guided Reverse Generation to augment expert trajectories, and then refined via Counterfactual Contrastive Policy Optimization to reward beneficial seeking actions. Experiments on simulated benchmarks show significant gains, improving success rates by 12.7% on R2R‑CE and 7.5% on RxR‑CE, and real‑world tests demonstrate human‑like evidence‑seeking behavior.
arXiv:2608. 16794v1 Announce Type: cross Abstract: Language and vision-language models generate plausible embodied plans but do not guarantee executability, as their outputs can violate environment dynamics or act on incorrectly grounded entities.
Embodied Reasoning constitutes a fundamental capability of embodied intelligence, serving as the basis for autonomous perception, reasoning, and interaction within physical environments. Recent studie...
arXiv:2607. 01754v1 Announce Type: new Abstract: On-policy exploration is a crucial component for training robust Vision-Language Navigation agents, as it exposes the policy to a broader state distribution.
arXiv:2609.32292v2 Announce Type: replace-cross Abstract: Vision-and-language navigation increasingly relies on general-purpose semantic planners, yet translating correct high-level intent into relia...
arXiv:2608.22971v1 Announce Type: new Abstract: Embodied Reasoning constitutes a fundamental capability of embodied intelligence, serving as the basis for autonomous perception, reasoning, and intera...
Language and vision-language models generate plausible embodied plans but do not guarantee executability, as their outputs can violate environment dynamics or act on incorrectly grounded entities. We present a neurosymbolic agent that factors long-horizon household tasks into task-directed visual exploration and constrained symbolic planning.