arXiv:2606. 09749v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have demonstrated impressive end-to-end performance across a variety of robotic manipulation tasks.
By Seongbin Park, Fan Zhang, Baharan Mirzasoleiman, Shahriar Talebi, Nader Sehatbakhsh
arXiv:2606. 12299v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models provide a natural language interface to robot control, but the mapping from language to behavior is often brittle and unintuitive: semantically similar instructions can induce drastically different behaviors, while some capabilities may not be elicitable through prompting alone.
By Hyun Joe Jeong, Gokul Swamy, Andrea Bajcsy
The paper introduces a compact visual navigation system that decomposes the task into three analytically‑computed geometric interfaces and three small learned modules: an egress predictor, a navigation predictor, and an endpoint‑pinned residual diffusion generator. Only 0.58 M of the 23 M parameters are trained on 44 k frames, achieving competitive success rates and the lowest collision rate among evaluated methods across 6 060 point‑goal episodes in 60 environments. The design allows further parameter reduction by replacing the frozen image encoder with a 0.54 M MobileNetV2, supports zero‑shot deployment on a Jetson Orin Nano UGV, and enables transparent failure analysis under sensor corruption.
By Edward Beng Wai Tan, Siew-Kei Lam
arXiv:2609.17628v1 Announce Type: cross
Abstract: Active perception allows autonomous agents to select their viewpoints rather than passively process the viewpoints given to them, enabling them to ta...
By \"U. Bora G\"okbakan (WILLOW), St\'ephane Caron (ISIR), Philippe Sou\`eres (LAAS-GEPETTO)
arXiv:2609.39145v1 Announce Type: cross
Abstract: Unreliable visual inputs can harm task performance and cause potential physical safety risks for vision-language-action (VLA) models. We analyze how...
By Heejae Suh, Jongwook Han, Zahra Gholami, Yohan Jo
arXiv:2607. 18200v1 Announce Type: cross Abstract: Robots in cluttered indoor spaces often fail not because they cannot generate collision-free paths, but because a fixed safety margin is mis-calibrated: conservative margins cause detours and timeouts, while permissive margins lead to near-boundary shortcuts under perception bias.
By Junyi Hu, Shuaihang Yuan, Geeta Chandra Raju Bethala, Anthony Tzes, Yi Fang
arXiv:2609.36520v1 Announce Type: cross
Abstract: RGB-only end-to-end visual navigation policies remain vulnerable to collisions in real-world dynamic environments, motivating a dedicated safety laye...
By Seungyeon Yoo, Gawon Lee, Seungwoo Jung, Inkyu Jang, H. Jin Kim
arXiv:2607. 20785v1 Announce Type: cross Abstract: Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently.
By Arjun Majumdar, Avinash Sooriyarachchi, Benjamin Tibi, Chris Bamford, Elliot Chane-Sane, Guillaume Lample, Khyathi Raghavi Chandu, Ludovic Ho Fuh, Mathieu Poiree, Olivier Duchenne, Rosalie Millner, Srijan Mishra, Theo Cachet, Thomas Chabal
arXiv:2607. 29169v1 Announce Type: cross Abstract: Vision-language-action (VLA) policies achieve strong performance in robotic manipulation but remain vulnerable to runtime disturbances that break the temporal alignment among visual observations, robot states, and executed actions.
By Wenda Yu, Tianshi Wang, Fengling Li, Xin Li, Jingjing Li, Lei Zhu
arXiv:2608.21395v1 Announce Type: cross
Abstract: NoMaD [31] is a learned vision-navigation policy that unifies goal-conditioned navigation and exploration in a single goal-masked diffusion policy. I...
By Blossom Treesa Bastian, Keerthi S. Shetty, Manish Kolachalam, Rani Malhotra, Ashish Dutta
arXiv:2509. 25533v2 Announce Type: replace-cross Abstract: As Vision Language Models (VLMs) are deployed across safety-critical applications, understanding and controlling their behavioral patterns has become increasingly important.
By Ravikumar Balakrishnan, Mansi Phute
arXiv:2608. 15284v1 Announce Type: cross Abstract: Navigation instruction generation from ego-centric RGB video in continuous environments is an important yet challenging task for human-robot interaction and scalable dataset construction.
By Haolin Yang, Yuxing Long, Zihan Yang, Hao Dong