arXiv AI

VOR-Bench: A Human Perception-Driven Benchmark for Video Object Removal

VOR-Bench is a new benchmark for video object removal that addresses shortcomings in current evaluation methods by providing a dataset with paired edited videos and graffiti masks, a realistic motion-capable paired-video acquisition framework (rMPAF), and a perception-driven scoring model (VOR-MDSM). The dataset includes diverse data from model-generated, tool-rendered, and camera-captured sources, ensuring robust real-world assessment. Experiments show that VOR-Bench’s evaluation results correlate strongly (ρ > 0.9) with human subjective judgments, bridging the gap between traditional metrics and human preference.

arXiv AI
Jun 30

MotionAtlas: Detailed Region Captioning for Motion-Centric Videos

arXiv:2606. 29531v1 Announce Type: cross Abstract: We propose MotionAtlas, a system for detailed captioning of motion-centric videos, comprising (1) a dedicated human-annotated benchmark, (2) a scalable, high-quality pipeline to construct training samples, and (3) a family of powerful Video-MLLMs.

By Weisong Liu, Haochen Wang, Kuan Gao, Yuhao Wang, Yikang Zhou, Zhongwei Ren, Jacky Mai, Anna Wang, Yanwei Li, Jason Li, Zhaoxiang Zhang
Hugging Face Trending Papers
Aug 19

CamWorldQA: Perceptual Quality Assessment of Camera-Controlled World Video Generation

CamWorldQA introduces the first benchmark for assessing the perceptual quality of camera‑controlled world video generation, featuring 720 videos generated by six methods from 20 source videos across six camera trajectories, each scored by human raters. The paper also presents CWQA, a no‑reference quality assessment network that combines spatial, temporal motion, and optical flow features to predict quality scores. Experiments show CWQA outperforms existing VQA methods on the CamWorldQA dataset.

arXiv Computer Vision
Sep 4

WorldReward: Reward Modeling for Camera-Conditioned World Models

WorldReward introduces a vision‑language model–based reward system for camera‑conditioned world models, combining action consistency and visual quality evaluation. It processes paired videos by splitting them into action‑aligned chunks, structuring visual evidence, and aggregating decisions through voting. The model is trained on a large, reasoning‑augmented preference dataset and outperforms GPT‑5.5 on a human‑annotated benchmark, improving both action execution and visual quality when applied to RL post‑training.

By Yibin Wang, Zehan Wang, Junshu Tang, Zhimin Li, Yujie Zhou, Jiazi Bu, Pengyang Ling, Feng Han, Zhixiong Zhang, Long Xing, Shengyuan Ding, Ziang Li, Cheng Jin, Yuhang Zang, Jiaqi Wang, Tianyu Pang
arXiv Computer Vision
Sep 16

High-Fidelity Video Quality Assessment with VQA-Specific Saliency

High-Fidelity Video Quality Assessment (HFVQA) is a new framework that uses fixed-size spatio‑temporal patches across multiple scales, including the original resolution, to preserve low‑level quality cues and semantic context. It incorporates a lightweight auxiliary network that learns VQA‑specific saliency directly from quality supervision, enabling the model to focus on the most important spatio‑temporal regions. By combining high‑fidelity cues with task‑specific saliency, HFVQA achieves state‑of‑the‑art performance on standard no‑reference VQA benchmarks while processing only about 12% of the candidate patches, making it computationally efficient.

By Hakan Emre Gedik, Shashank Gupta, Alan Bovik
arXiv Computer Vision
Aug 27

VGA-BenchV2: An Expanded Unified Benchmark and Multi-Model Framework for Evaluating Video Aesthetics and Generation Quality

VGA‑BenchV2 is an expanded, human‑aligned benchmark and optimization framework that jointly evaluates video generation quality and aesthetic value. It builds on the original VGA‑Bench taxonomy, adding 52 sub‑dimensions and 1,016 curated prompts to generate over 60,000 videos from 12 mainstream models. The benchmark significantly enlarges human supervision with 36,000 task‑level annotations and introduces a hybrid evaluator (VAQA‑Net, VTag‑Net, VGQA‑Net) that aligns well with human judgments and can be used as a reward model for reinforcement‑learning fine‑tuning.

By Longteng Jiang, DanDan Zheng, Qianqian Qiao, Heng Huang, Huaye Wang, Yihang Bo, Bao Peng, Jingdong Chen, Jun Zhou, Xin Jin
arXiv Computer Vision
Sep 1

ACE-Ego-Hand: Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Motion Recovery

arXiv:2608.20308v2 Announce Type: replace Abstract: Egocentric video offers scalable manipulation data for embodied AI, yet recovering metric 3D hand trajectories remains challenging due to severe ob...

By Yufei Liu, Xixi Wang, Hao Li, Ganlong Zhao, Kaitong Cai, Chengkai Jin, Chunxiao Liu, Jianbo Liu, Siyuan Huang, Xingang Pan, Hongsheng Li