arXiv AI

Safe Multi-Robot Coordination via VLM-LLM Reasoning and Reachability Analysis

arXiv Computer Vision
Sep 14

AnchorVLN: Geometry-Anchored Vision-Language Grounding Reasoning for Open-Vocabulary Navigation

AnchorVLN is an open‑vocabulary vision‑language navigation system that separates semantic proposals from geometric metrics. It uses a VLM to generate semantics while a geometry module supplies reliable metric quantities such as range and bearing, all within a Model Context Protocol server. The system achieves 64.4% on instruction following and improves object‑reference accuracy, reducing median center error from 3.37 m to 2.48 m.

By Long Giang Vu, Chengkai Yao, Yuxin Liu, FNU Aryan, Rajath Chandrashekar Aralikatti
arXiv Computer Vision
Sep 21

VeriFuse: Bounded Vision-Language Arbitration and Reason-Guided Refinement for Cooperative 3D Perception

VeriFuse is a bounded arbitration framework that integrates vision‑language models (VLMs) into vehicle‑infrastructure cooperative 3D perception. Each agent first generates independent detections, then VeriFuse creates a unified candidate pool of geometric proposals and cross‑source hypotheses. A frozen VLM selects among three actions—SELECT, REFINE, or REJECT—to resolve ambiguity and produce final 3D detections, achieving strong AP50/AP70 scores on the DAIR‑V2X dataset while keeping vehicle‑side BEV AP50 drop minimal under delay.

By Hongyi Lin, Yiyao Liu, Qi Kang, Heye Huang, Yang Liu, Haris Koutsopoulos, Jinhua Zhao
arXiv AI
Jul 10

Time-to-Collision Based Dynamic Obstacle Avoidance Using Pretrained Vision Models for Robots in Unstructured Environments

arXiv:2607. 07885v1 Announce Type: cross Abstract: Dynamic obstacle avoidance in unstructured outdoor environments remains a critical challenge for autonomous mobile robots, particularly when large-scale robot-specific training data and simulation-based policies are impractical.

By Erik Jagnandan, Mulugeta Haile, Gregory Barber, Pratik Chaudhari
arXiv AI
Aug 11

REMAC: Self-Reflective and Self-Evolving Multi-Agent Collaboration for Long-Horizon Robot Manipulation

arXiv:2503. 22122v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have demonstrated remarkable capabilities in robotic planning, particularly for long-horizon tasks that require a holistic understanding of the environment for task decomposition.

By Puzhen Yuan, Angyuan Ma, Yunchao Yao, Huaxiu Yao, Masayoshi Tomizuka, Mingyu Ding