arXiv AI By Haonan Huang, Tianrui Qiu, Xianghao Zang, Yinan Du, Zhixiang He, Chi Zhang, Hao Sun, Zhongjiang He, Tianwei Cao, Xuchong Zhang, Hongbin Sun, Kongming Liang, Zhanyu Ma

VOR-Bench: A Human Perception-Driven Benchmark for Video Object Removal

Read the original on arXiv AI →

VOR-Bench is a new benchmark for video object removal that addresses shortcomings in current evaluation methods by providing a dataset with paired edited videos and graffiti masks, a realistic motion-capable paired-video acquisition framework (rMPAF), and a perception-driven scoring model (VOR-MDSM). The dataset includes diverse data from model-generated, tool-rendered, and camera-captured sources, ensuring robust real-world assessment. Experiments show that VOR-Bench’s evaluation results correlate strongly (ρ > 0.9) with human subjective judgments, bridging the gap between traditional metrics and human preference.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 30

MotionAtlas: Detailed Region Captioning for Motion-Centric Videos

arXiv:2606. 29531v1 Announce Type: cross Abstract: We propose MotionAtlas, a system for detailed captioning of motion-centric videos, comprising (1) a dedicated human-annotated benchmark, (2) a scalable, high-quality pipeline to construct training samples, and (3) a family of powerful Video-MLLMs.

By Weisong Liu, Haochen Wang, Kuan Gao, Yuhao Wang, Yikang Zhou, Zhongwei Ren, Jacky Mai, Anna Wang, Yanwei Li, Jason Li, Zhaoxiang Zhang
Hugging Face Trending Papers
Aug 19

CamWorldQA: Perceptual Quality Assessment of Camera-Controlled World Video Generation

CamWorldQA introduces the first benchmark for assessing the perceptual quality of camera‑controlled world video generation, featuring 720 videos generated by six methods from 20 source videos across six camera trajectories, each scored by human raters. The paper also presents CWQA, a no‑reference quality assessment network that combines spatial, temporal motion, and optical flow features to predict quality scores. Experiments show CWQA outperforms existing VQA methods on the CamWorldQA dataset.

arXiv Computer Vision
Sep 4

WorldReward: Reward Modeling for Camera-Conditioned World Models

WorldReward introduces a vision‑language model–based reward system for camera‑conditioned world models, combining action consistency and visual quality evaluation. It processes paired videos by splitting them into action‑aligned chunks, structuring visual evidence, and aggregating decisions through voting. The model is trained on a large, reasoning‑augmented preference dataset and outperforms GPT‑5.5 on a human‑annotated benchmark, improving both action execution and visual quality when applied to RL post‑training.

By Yibin Wang, Zehan Wang, Junshu Tang, Zhimin Li, Yujie Zhou, Jiazi Bu, Pengyang Ling, Feng Han, Zhixiong Zhang, Long Xing, Shengyuan Ding, Ziang Li, Cheng Jin, Yuhang Zang, Jiaqi Wang, Tianyu Pang
arXiv Computer Vision
Sep 16

High-Fidelity Video Quality Assessment with VQA-Specific Saliency

High-Fidelity Video Quality Assessment (HFVQA) is a new framework that uses fixed-size spatio‑temporal patches across multiple scales, including the original resolution, to preserve low‑level quality cues and semantic context. It incorporates a lightweight auxiliary network that learns VQA‑specific saliency directly from quality supervision, enabling the model to focus on the most important spatio‑temporal regions. By combining high‑fidelity cues with task‑specific saliency, HFVQA achieves state‑of‑the‑art performance on standard no‑reference VQA benchmarks while processing only about 12% of the candidate patches, making it computationally efficient.

By Hakan Emre Gedik, Shashank Gupta, Alan Bovik