arXiv Computer Vision By Wenxue Li, Peiyan Guan, Haoyang Jiang, Junxian Cai, Hualuo Liu, Chunjie Zhang, Chong Guan, Songlian Li, Taiyi Wu, Yongjian Yu, Xiaotong Zhao, Alan Zhao, Eric Liu, Xi Chen, Yu Liu, Lei Zhu

OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation

Read the original on arXiv Computer Vision →

OmniVBench introduces a comprehensive benchmark and a large-scale dataset for omni reference-to-video (R2V) generation, addressing gaps in existing evaluations that focus only on limited reference types and holistic consistency. The benchmark expands evaluation across 7 task families and 18 fine-grained tasks, covering content, motion, style, structure, narrative, and multi-reference settings, and employs a factor‑grounded evaluation with 12,172 checklist items to assess preservation, disentanglement, and routing of reference factors. The accompanying Omni‑R2V Dataset provides 340K training samples derived from professional video footage, along with task‑specific pipelines for scalable data construction, enabling broader research and revealing performance gaps in current R2V models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Aug 27

RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing

RefVideo-6M is a new large-scale reference-guided editing dataset that includes 5 million video editing samples and 1 million image editing samples, each paired with about 6 million visual references. The dataset is constructed to avoid artifacts by using real, artifact‑free videos as targets and filtering input conditions with multiple editing experts, thereby providing reliable supervision. It enables models to learn fine‑grained visual correspondence beyond text‑only instructions and supports the training of a reference‑guided video editing model, Ref‑MoT, which shows improved visual quality, controllability, and reference consistency.

By Bojia Zi, Xiaoyan Yang, Yu Zhou, Ruijie Sun, Lihan Zhang, Bin Liang, Kam-Fai Wong, Haibin Huang, Chi Zhang, Xuelong Li
arXiv Computer Vision
Aug 25

Sa2VA: Marrying SAM2 with MLLM for Dense Grounded Understanding of Images and Videos

arXiv:2501.04001v4 Announce Type: replace Abstract: This work presents Sa2VA, the first comprehensive, unified model for dense grounded understanding of both images and videos. Unlike existing multi-...

By Haobo Yuan, Xiangtai Li, Tao Zhang, Yueyi Sun, Zilong Huang, Shilin Xu, Shunping Ji, Yunhai Tong, Lu Qi, Jiashi Feng, Ming-Hsuan Yang
arXiv AI
Aug 19

SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation

SemComp-Bench introduces a new video generation task called Semantic Task Completion, where a model must produce a video that achieves a specified outcome while maintaining semantic alignment with a reference image. The benchmark includes the SemComp-Data dataset, spanning six domains, and a four-stage curation pipeline that transforms raw videos into standardized instances. Evaluation is performed via a vision‑language model that answers structured binary questions, yielding Outcome Achievement (OA) and Generation Reliability (GR) scores.

By Keyu Tu, Zhuowei Chen, Mengqi Huang, Yuxin Wang, Jiahao Zhu, Zhendong Mao, Yongdong Zhang
arXiv Computer Vision
Aug 27

AdaVDR: Adaptive Tool Use and Reflection for Video Deep Research

AdaVDR is an adaptive video deep research agent that selects and reflects on tool usage based on the task and the model’s capabilities. It constructs a specialized data pipeline to generate high‑quality QA pairs and uses model‑conditioned filtering to remove unnecessary tool calls. The agent is trained with supervised fine‑tuning and reinforcement learning, achieving top performance on the VDR‑EE benchmark and significant gains on VideoDR.

By Xintong Zhang, Xiaomeng Fan, Shilin Yan, Ekko He, Zicheng Liu, Zijian Zou, Guannan Zhang, Yuwei Wu, Zhi Gao, Hongwei Xue