arXiv Computer Vision

Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios

Hugging Face Trending Papers
Jun 3

VCIFBench: Evaluating Complex Instruction Following for Video Understanding

Multimodal large language models have made rapid progress in video understanding, yet existing benchmarks largely rely on simple prompts and provide limited evidence about whether models can satisfy explicit output constraints. We introduce VCIFBench, a benchmark for evaluating complex instruction following in video understanding.

arXiv AI
Jun 8

Watch, Remember, Reason: Human-View Video Understanding with MLLMs

arXiv:2606. 07433v1 Announce Type: cross Abstract: Video understanding is being rapidly transformed by multimodal large language models (MLLMs), as research moves from short clips to long, multimodal, and knowledge-intensive video scenarios.

By Jiahao Meng, Yue Tan, Qi Xu, Kuan Gao, Weisong Liu, Yanwei Li, Jason Li, Lingdong Kong, Haochen Wang, Qianyu Zhou, Jiangning Zhang, Guangliang Cheng, Yunhai Tong, Lu Qi, Minghsuan Yang
arXiv Computer Vision
4d ago

Sa2VA: Marrying SAM2 with MLLM for Dense Grounded Understanding of Images and Videos

arXiv:2501.04001v4 Announce Type: replace Abstract: This work presents Sa2VA, the first comprehensive, unified model for dense grounded understanding of both images and videos. Unlike existing multi-...

By Haobo Yuan, Xiangtai Li, Tao Zhang, Yueyi Sun, Zilong Huang, Shilin Xu, Shunping Ji, Yunhai Tong, Lu Qi, Jiashi Feng, Ming-Hsuan Yang
arXiv AI
Aug 19

SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation

SemComp-Bench introduces a new video generation task called Semantic Task Completion, where a model must produce a video that achieves a specified outcome while maintaining semantic alignment with a reference image. The benchmark includes the SemComp-Data dataset, spanning six domains, and a four-stage curation pipeline that transforms raw videos into standardized instances. Evaluation is performed via a vision‑language model that answers structured binary questions, yielding Outcome Achievement (OA) and Generation Reliability (GR) scores.

By Keyu Tu, Zhuowei Chen, Mengqi Huang, Yuxin Wang, Jiahao Zhu, Zhendong Mao, Yongdong Zhang
Hugging Face Trending Papers
Jun 1

WALL-WM: Carving World Action Modeling at the Event Joints

WALL-WM is a World Action Model that shifts video-action learning from chunk-centric optimization to event-grounded Vision-Language-Action pretraining, using semantically coherent action events as the atomic unit of learning. Existing WAMs commonly initialize from multimodal or video foundation models and then optimize fixed-length action chunks conditioned directly on the current observation and instruction.

arXiv Computer Vision
5d ago

OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs

OmniAssistBench is a new benchmark for evaluating omni-modal large language models (Omni-LLMs) as real‑time video assistants that actively guide users toward goals. The benchmark addresses the challenge of dynamic interaction paths by providing models with predefined priors from source videos, forcing them to follow the same routes as users. The dataset was constructed by reverse‑engineering existing Internet videos into multi‑turn clips, a process that required over 1,000 expert person‑hours. Results show that proprietary Gemini‑3‑Pro scores 66.4/100 while open‑source Qwen3‑Omni‑Instruct scores 51.2, revealing that current models often give incorrect or incomplete answers, struggle with visual prompts, and fail to maintain context or delay responses until target events.

By Xianyun Sun, Chaoyou Fu, Zhengye Zhang, Feiyang Duan, Qingyuan Cao, Yonghui Niu, Sihang Yuan, Ge Zhang, Caifeng Shan