arXiv AI

Clearing the Fog: Towards Installing and Refining Proactive Exploration Capabilities in LLM Agents

arXiv:2608. 14339v1 Announce Type: new Abstract: We study proactive exploration in LLM agents, i.

arXiv Computer Vision
Sep 22

Active Spatial Inspection for Effective and Efficient Embodied Exploration

The paper introduces the ACE framework, which uses spatially explicit inspection to guide embodied exploration. By combining evidence‑grounded perception with exposure‑informed movement, ACE provides a spatially resolved decision paradigm that improves cue assessment and movement direction. Experiments show ACE boosts navigation task success by 18.0% and exploration efficiency by 10.3% over previous baselines.

By Wenbin Wang, Xiang Bai, Yizhao Wang, Hang Sun, Dong Ren, Jie Qin, Qingquan Li, Bing Wang
Hugging Face Trending Papers
Jul 2

Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation

On-policy exploration is a crucial component for training robust Vision-Language Navigation agents, as it exposes the policy to a broader state distribution. However, such exploration inevitably leads to trajectories that deviate from expert demonstrations, resulting in a semantic mismatch between the executed visual stream and the original language instruction.

arXiv AI
Sep 10

Efficient Exploration Is Enough

arXiv:2609.07575v1 Announce Type: cross Abstract: This work introduces an alternative view of efficient exploration and studies its theoretical and empirical implications in the absence of extrinsic...

By Mikel Malag\'on, Jon Vadillo, Josu Ceberio, Michael Bowling, Jose A. Lozano
arXiv AI
4d ago

Seek Before You Move: Evidence Seeking for Progress Grounding in Vision-Language Navigation

SeekVLN is a new framework for Vision‑Language Navigation that addresses the problem of agents acting on insufficient evidence, termed Progress Myopia. It combines semantic progress reasoning with active evidence seeking, trained first with Future‑guided Reverse Generation to augment expert trajectories, and then refined via Counterfactual Contrastive Policy Optimization to reward beneficial seeking actions. Experiments on simulated benchmarks show significant gains, improving success rates by 12.7% on R2R‑CE and 7.5% on RxR‑CE, and real‑world tests demonstrate human‑like evidence‑seeking behavior.

By Zhimin Wang, Meiyuan Zhu, Duo Wu, Linjia Kang, Yajun Wang, Yuan Ni, Xiaohang Wang, Tianlu Pan, Jingyan Jiang, Yaowei Wang, Zhi Wang
arXiv AI
Sep 21

Prompt-Driven Exploration: Language as an Exploration Space for VLA Reinforcement Learning

arXiv:2607.08837v4 Announce Type: replace-cross Abstract: Exploration is essential to RL since a policy cannot improve by repeatedly sampling the behaviors it already prefers. Standard methods inject...

By Sunshine Jiang, John Marangola, David Zhang, Raghuram Kowdeed, Ruiyang Luo, Nitish Dashora, Richard Li, Pulkit Agrawal, Zhang-Wei Hong