RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis
arXiv:2606. 28385v1 Announce Type: cross Abstract: Recent advances in robot world models enable synthetic video generation for embodied prediction and planning.
arXiv:2606. 28385v1 Announce Type: cross Abstract: Recent advances in robot world models enable synthetic video generation for embodied prediction and planning.
The paper introduces Configured Failure Trapping, a new backdoor attack targeting Vision‑Language‑Action (VLA) models that activates through subtle textual triggers and forces the robot to fail in a specific, controlled manner. It presents a data engine for generating high‑quality target trajectories, an automated evaluation suite, and two benchmarks—Trap‑LIBERO and Trap‑RoboTwin—covering four failure modes. The authors propose TrapVLA, a method that learns trigger‑induced action residuals to steer policies toward the desired failure behavior, demonstrating effectiveness in both simulation and real‑world robotic experiments while maintaining performance on clean data.
Vision-Language-Action (VLA) models have attracted growing interest as a scalable approach to robotic manipulation. While these models are effective action predictors, deploying them as robotic agents exposes critical gaps: no mechanism for failure recovery, inconsistent execution over long horizons, and limited robustness to shifts in observations, tasks, or embodiments.
arXiv:2607. 23784v1 Announce Type: cross Abstract: While vision-language-action models have demonstrated impressive zero-shot manipulation capabilities, they remain fundamentally black box policies that are difficult to interpret, adapt, or correct when they inevitably fail.
arXiv:2603. 15600v2 Announce Type: replace-cross Abstract: Accurate process supervision remains a critical challenge for long-horizon robotic manipulation.
arXiv:2512. 24125v3 Announce Type: replace-cross Abstract: General-purpose robotic systems operating in open-world environments must achieve both broad generalization and high-precision action execution, a combination that remains challenging for existing Vision-Language-Action (VLA) models.
arXiv:2606. 08881v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have demonstrated strong generalization in robotic manipulation, yet existing evaluations are primarily conducted in simulation or on expensive robotic platforms, leaving their robustness on affordable real-world robots largely unexplored.
arXiv:2512. 19178v2 Announce Type: replace-cross Abstract: Bridging the gap between natural language commands and autonomous execution in unstructured environments remains an open challenge for robotics.
arXiv:2510. 14828v3 Announce Type: replace Abstract: Improving the reasoning capabilities of embodied agents is crucial for robots to complete complex human instructions in long-view manipulation tasks successfully.
arXiv:2510. 17640v4 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have shown strong manipulation capability when trained with large-scale imitation learning datasets.
arXiv:2601. 00969v3 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models provide strong action priors for robotic manipulation, but their reactive behavior can fail under distribution shift and long-horizon task structure.
arXiv:2606. 18043v1 Announce Type: cross Abstract: Vision-language-action models (VLAs) combine vision-language backbones with expressive generative action heads trained via flow matching on large-scale robotic datasets.