arXiv Computer Vision

ProTracer: Proprioception-Guided Failure Diagnosis in Robot Manipulation

ProTracer is a training‑free framework that uses Vision‑Language Models (VLMs) together with proprioceptive signals to analyze robot manipulation failures. It performs binary failure detection, categorization, explanation generation, and introduces failure onset localization—identifying the earliest moment a robot deviates from a valid trajectory leading to failure. The method leverages proprioceptive dynamics to pinpoint informative action boundaries and converts robot‑state signals into natural‑language descriptions for joint multimodal reasoning, achieving strong performance on both conventional failure diagnosis and the new failure onset localization task.

arXiv Computer Vision
Sep 7

FailureSpot: Label-Efficient Timestamp-Level Failure Detection for Vision-Language-Action Models

The paper introduces FailureSpot, a label‑efficient method for detecting failures at the timestamp level in vision‑language‑action (VLA) policies. It first generates weak supervision from unlabeled VLA action chunks by identifying abnormal patterns, then employs active learning to annotate only the most uncertain trajectories. Experiments on multiple VLA policies demonstrate improved performance for both timestamp‑level and trajectory‑level failure detection.

By Jie Ma, Zongxi Liu, Yi Zhu
arXiv AI
Sep 4

FailBench: How Reliable are VLMs at Judging Robot Task Success?

FailBench is a new benchmark for robot failure detection, containing 2,197 manipulation attempts from 14 public sources, with 75% of failures occurring naturally. The study evaluates 13 vision‑language model (VLM) detectors, finding the best model achieves only 0.77 mean balanced accuracy, and that fine‑tuned failure detectors often underperform general‑purpose VLMs. Performance varies with visual evidence, excelling when object motion is observable but dropping to near chance on contact‑intensive assembly tasks, and input‑level cropping of outcome‑relevant regions improves the top detector by 2.4 percentage points.

By Zaruhi Navasardyan, Tatul Danielyan, Hrant Davtyan
arXiv Computer Vision
Aug 28

TrapVLA: Trapping Vision-Language-Action Models in Configured Failure Modes

The paper introduces Configured Failure Trapping, a new backdoor attack targeting Vision‑Language‑Action (VLA) models that activates through subtle textual triggers and forces the robot to fail in a specific, controlled manner. It presents a data engine for generating high‑quality target trajectories, an automated evaluation suite, and two benchmarks—Trap‑LIBERO and Trap‑RoboTwin—covering four failure modes. The authors propose TrapVLA, a method that learns trigger‑induced action residuals to steer policies toward the desired failure behavior, demonstrating effectiveness in both simulation and real‑world robotic experiments while maintaining performance on clean data.

By Jun-Hui Liu, Kun-Yu Lin, Yi-Lin Wei, Xu-Han Chen, Yinghao Li, Zhuohao Li, Yuan-Ming Li, Qing Zhang, Xiaoyi Fan, Dongmei Jiang, Yan Li, Wei-Shi Zheng
arXiv AI
Jun 9

Benchmarking Vision-Language-Action Models on SO-101: Failure and Recovery Analysis

arXiv:2606. 08881v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have demonstrated strong generalization in robotic manipulation, yet existing evaluations are primarily conducted in simulation or on expensive robotic platforms, leaving their robustness on affordable real-world robots largely unexplored.

By Yi Yu, Xinchuan Qiu
arXiv AI
Jun 8

ViVa: A Video-Generative Value Model for Robot Reinforcement Learning

arXiv:2604. 08168v2 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models have advanced robot manipulation through large-scale pretraining, but real-world deployment remains challenging due to partial observability and delayed feedback.

By Jindi Lv, Hao Li, Jie Li, Fankun Kong, Yang Wang, Pengfei Yi, Yifei Nie, Xiaofeng Wang, Zheng Zhu, Chaojun Ni, Qiuping Deng, Hengtao Li, Jiancheng Lv, Guan Huang
arXiv AI
2d ago

VLA-Scope: Shift-Aware Failure Prediction for Vision-Language-Action Models

VLA-Scope is a two‑stage framework designed to predict failures in vision‑language‑action models under distribution shifts. The first stage detects out‑of‑distribution inputs and classifies their shift categories using pooled image and language representations. For OOD inputs, the second stage updates failure risk during execution by combining shift category, action‑prefix features, and execution progress, achieving a ROC‑AUC of 0.8497 after 60 actions and outperforming baseline methods.

By Kaiwen Zhu, Dongfang Liu, Liangkai Liu
arXiv AI
5d ago

AntiGrounding: Executable Robot Trajectories as Visual Prompts for VLM-Guided Manipulation

AntiGrounding is a visual action-selection framework that turns short robot trajectories into both executable motion plans and rendered prompts for vision‑language model evaluation. After filtering for feasibility, each trajectory is scored on safety, task alignment, efficiency, and physical plausibility using structured multi‑view visual question answering, and the best trajectories are refined and validated by a digital twin before real‑world execution. In eight real‑world manipulation tasks, the system achieved a 71.25% success rate with a single GPT‑6 Astra evaluator, outperforming baseline methods.

By Wenbo Li, Yiteng Chen, Wenhao Li, Qingyao Wu
arXiv Computation and Language
Sep 7

Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models

The paper introduces ROBORMBENCH, a benchmark comprising 2,390 real‑robot trajectories, 21,673 verified paraphrases, and ground‑truth progress labels, to evaluate paraphrase robustness in vision‑language reward models (VLMs). It demonstrates that current VLMs often give different rewards for semantically equivalent goal descriptions, sometimes flipping a robot’s outcome from failure to success. The study finds that this instability is widespread, worsens with more divergent rewrites, and is not mitigated by model scale or explicit reasoning, though dedicated reward models trained with trajectory‑grounded supervision show greater stability.

By Wonje Jeung, Sangyeon Yoon, Hyesoo Hong, Yoonjun Cho, Dongjae Jeon, Bumjun Kim, Jean Oh, Youngjae Yu, Albert No