arXiv AI By Alexandru-Bogdan Dura, Sebastian Balmus, Radu Tudor Ionescu

RoboVAD: A Large Cross-Domain Evaluation Benchmark for Anomaly Detection in Robotic Arm Manipulation Videos

Read the original on arXiv AI →

RoboVAD is a large-scale benchmark for video anomaly detection in robotic arm manipulation, featuring cross-domain evaluation where both tasks and anomaly types are unseen during training. The dataset challenges existing VAD methods, with state-of-the-art approaches, including a new method tailored for robotic arms, still achieving less than 70% micro-averaged frame-level AUC in the hardest setting. The authors provide the dataset and code publicly for further research.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Aug 25

Frame-Level Evaluation in Weakly Supervised Video Anomaly Detection Mostly Measures Video-Level Ranking

The paper investigates how weakly supervised video anomaly detectors, trained with only video‑level labels, are evaluated using frame‑level metrics such as Micro‑AUROC and AP. It shows that these metrics largely measure a detector’s ability to separate different videos rather than correctly ordering anomalous moments within a single video, a phenomenon termed temporal dilution. Experiments demonstrate that a detector can achieve high pooled scores even when it assigns the same score to every frame in a video, indicating that current evaluation practices may overstate temporal localization performance.

By Inpyo Song, Jangwon Lee
arXiv AI
Jun 30

RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis

arXiv:2606. 28385v1 Announce Type: cross Abstract: Recent advances in robot world models enable synthetic video generation for embodied prediction and planning.

By Minh-Loi Nguyen, Nghiem Tuong Diep, Hung Khang Nguyen, Minh Le, Doanh Le Thien, Hoang H. Tran, Dung D. Le, Vu N. Duong, Daniel Sonntag, An Thai Le, Duy Minh Ho Nguyen, Vien Anh Ngo, Tran Van Nhiem
arXiv Computer Vision
Sep 2

ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training

ZimaBlue is a scalable framework that learns generalizable World Action Models (WAMs) from large-scale egocentric videos. It follows a three-stage curriculum: causal video pre‑training, video‑action mid‑training with a unified action representation, and final specialization to a target robot. The system employs an asynchronous Slow‑Fast architecture to enable real‑time 30 Hz action prediction, achieving a jump in real‑robot zero‑shot success from 36.1% to 77.8% when leveraging over 120,000 hours of embodied video.

By Xionghao Wu, Yijun Yang, Shiyang Zhou, Haoze Sun, Jianhui Liu, Songsong Yu, Jiyao Zhang, Wenbo Li, Bo Wang, Guoqing Ma, Lin Song, Renjie Liao, Shenghe Zheng, Wei Tang, Xiaojuan Qi, Yanwei Li, Yuan Zhang, Zhuotao Tian, Haoyang Huang, Nan Duan