arXiv Computer Vision

LIBERO-VPro: Benchmarking Closed-Loop Visual Robustness of Robotic Foundation Models

LIBERO-VPro is a benchmark designed to assess the closed‑loop visual robustness of robotic foundation models by systematically perturbing visual inputs during task execution. It spans four dimensions—Visual Evidence Degradation, Camera Staleness, Visual Source Consistency, and Task‑Relevant Scene Variation—across 12 challenge categories, 96 settings, and 3,296 task‑condition cases. Evaluations on six models over 196,000 simulated episodes and 200 real‑world rollouts show that high nominal performance can hide significant weaknesses in visual grounding, adaptation, and sensitivity to stale or missing observations.

arXiv AI
Jun 30

RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis

arXiv:2606. 28385v1 Announce Type: cross Abstract: Recent advances in robot world models enable synthetic video generation for embodied prediction and planning.

By Minh-Loi Nguyen, Nghiem Tuong Diep, Hung Khang Nguyen, Minh Le, Doanh Le Thien, Hoang H. Tran, Dung D. Le, Vu N. Duong, Daniel Sonntag, An Thai Le, Duy Minh Ho Nguyen, Vien Anh Ngo, Tran Van Nhiem
arXiv AI
Sep 7

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

RoboSPA is a large-scale robotic manipulation dataset and benchmark designed to evaluate Vision‑Language‑Action models on fine‑grained spatial reasoning and long‑horizon procedural planning. It contains 10 task categories, 56 base tasks, and 280 variants across five difficulty levels, with 527K trajectories collected from multiple embodiments and scenes. The benchmark introduces diagnostic metrics beyond binary success, revealing that current VLA models struggle with complex spatial relations, precise execution, and memory‑intensive planning.

By Zhenxuan Fan, Bo Zhang, Yutong Lin, Yuqian Yuan, Juekai Lin, Liang Liang, Zhuoyi Huang, Wenqiao Zhang, Juncheng Li, Siliang Tang, Jun Xiao, Yueting Zhuang
arXiv AI
Sep 4

FailBench: How Reliable are VLMs at Judging Robot Task Success?

FailBench is a new benchmark for robot failure detection, containing 2,197 manipulation attempts from 14 public sources, with 75% of failures occurring naturally. The study evaluates 13 vision‑language model (VLM) detectors, finding the best model achieves only 0.77 mean balanced accuracy, and that fine‑tuned failure detectors often underperform general‑purpose VLMs. Performance varies with visual evidence, excelling when object motion is observable but dropping to near chance on contact‑intensive assembly tasks, and input‑level cropping of outcome‑relevant regions improves the top detector by 2.4 percentage points.

By Zaruhi Navasardyan, Tatul Danielyan, Hrant Davtyan
arXiv Computer Vision
5d ago

VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control

VABench is a benchmark that tests general‑purpose multimodal large language models (MLLMs) on embodied spatial intelligence by requiring them to observe, reason, act, and revise based on visual demonstrations and active perception. The benchmark includes 14 task families, a fixed model‑agnostic controller, and evaluates models on target localization, spatial relations, and long‑horizon composition tracks without providing privileged object poses or learned action heads. Results show that while the best model achieves perfect target localization, overall task success remains modest, and active camera control and geometric transfer significantly influence performance.

By Zhongbo Zhang, Jiayi Jin, Yifan Wang, Zaibin Zhang, Haiwen Diao, Lijun Wang, Huchuan Lu
arXiv AI
Jun 9

Benchmarking Vision-Language-Action Models on SO-101: Failure and Recovery Analysis

arXiv:2606. 08881v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have demonstrated strong generalization in robotic manipulation, yet existing evaluations are primarily conducted in simulation or on expensive robotic platforms, leaving their robustness on affordable real-world robots largely unexplored.

By Yi Yu, Xinchuan Qiu
arXiv Machine Learning
Jun 3

SeeTraceAct: Visibility-Aware Latent Planning from Cross-Embodiment Demonstration Videos

arXiv:2606. 02745v1 Announce Type: cross Abstract: Vision-language-action models (VLAs) are promising general-purpose robot policies, but adapting them to new tasks typically requires costly task-specific teleoperation data.

By Jaehyeon Son, Junhyun Kim, Kyle Kam, Jeremiah Coholich, Seok Joon Kim, Jinhoo Kim, Chris Dongjoo Kim, Jaemin Cho, Dieter Fox, Zsolt Kira
Hugging Face Trending Papers
Jul 30

RoboBRIDGE: A Modular Framework for Bridging Policies to Robust Real-World Robotic Agents

Vision-Language-Action (VLA) models have attracted growing interest as a scalable approach to robotic manipulation. While these models are effective action predictors, deploying them as robotic agents exposes critical gaps: no mechanism for failure recovery, inconsistent execution over long horizons, and limited robustness to shifts in observations, tasks, or embodiments.

arXiv AI
Sep 1

RoboPhys-3D: A Comprehensive Embodied World Model Evaluation via 3D Reconstruction

RoboPhys-3D is a 3D‑grounded embodied world model benchmark built on RoboTwin 2.0, featuring 50 manipulation tasks, 5,000 episodes, and 25,000 multi‑view ground‑truth videos. It evaluates video world models by processing both generated and ground‑truth videos through the same 3D reconstruction pipeline, allowing the separation of reconstruction‑induced from generation‑induced errors. The benchmark defines 50 metrics across four sub‑dimensions—pixel fidelity, 3D geometry consistency, state understanding, and task completeness—and introduces the Average Full Score and RoboPhyscore for holistic assessment, with RoboPhyscore showing strong correlation with human judgments.

By Tianyi Wang, Jiazhou Chen, Yiming Xu, Xiangyu Li, Tianyi Zeng, Chih-Hsien Chou, Ning Lu, Liang Peng, Junfeng Jiao, Christian Claudel
arXiv AI
Jul 29

CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model

arXiv:2607. 25487v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models translate natural-language commands into robot action sequences, but leading systems on the LIBERO-Plus robustness benchmark use three- to seven-billion-parameter backbones whose memory demands can exceed embedded robotic budgets.

By Minhyeok Lee, Chiyoung Kim, Chanhoe Gu, Seongrok Kim, Sanghyuk Roy Choi, Donghwan Hwang, Donghun Ryu, Seokhyun Kim