arXiv AI By Anindya Jana, Snehasis Banerjee, Arup Sadhu, Ranjan Dasgupta

A Modular Vision-Language-Action Robotics Framework for Indoor Environments

Read the original on arXiv AI →

arXiv:2606. 31144v1 Announce Type: cross Abstract: This paper presents an integrated system for the CMU Vision-Language-Action (VLA) Challenge, designed to enable an autonomous agent to perform complex tasks based on natural language instructions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 14

A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation

arXiv:2607. 09792v1 Announce Type: cross Abstract: Navigation is a fundamental capability of autonomous systems, yet most existing approaches rely on highly structured models and strong prior assumptions, limiting their robustness in open and uncertain real-world environments.

By Liuyi Wang, Kai Sheng, Zongtao He, Jinlong Li, Yongrui Qin, Haojie Dai, Xiangyi Wang, Jingwei Yang, Qingqing Yan, Chengju Liu, Qijun Chen