arXiv Computer Vision

Thinking with Cameras: Active Visual Reasoning via Dynamic Viewpoint Control for Surveillance Video Understanding

The paper introduces CamVLM, a framework that equips large vision‑language models with the ability to actively control camera viewpoints for improved surveillance video understanding. It presents two new datasets: CCTV‑Anomaly, a large‑scale surveillance video collection with detailed captions and event annotations, and CamTrack‑53K, an object‑centric viewpoint trajectory dataset for learning camera actions. Using reinforcement learning, CamVLM learns long‑horizon observation strategies, achieving state‑of‑the‑art performance in both passive and dynamic viewpoint settings.

arXiv AI
3d ago

MVVBench: Benchmarking 4D Reasoning in Vision-Language Models

MVVBench is a new benchmark for multi‑view video reasoning that tests vision‑language models on tasks requiring integration of spatial and temporal evidence across multiple, often non‑overlapping camera streams. The benchmark contains questions that cannot be answered from any single view or single moment, forcing models to jointly reason across views and time. It evaluates six capabilities—including attribute identification, relative distance, camera pose, and compositional counting—and provides human‑authored QA, rigorous verification, and detailed error analysis. "whyItMatters":"The benchmark offers a rigorous evaluation of 4D multi‑view reasoning and a foundation for future progress toward reliable embodied perception."

By Hyungjin Chung, Byeongjun Park, Joonseok Lee, Hojun Kim, Jaeho Choi, Byung-Hoon Kim
arXiv Machine Learning
Sep 17

ActiveScale: Scaling Active Perception for Robots across Model, Data, and Hardware

ActiveScale is a framework that enhances active perception for robots by integrating model, data, and hardware innovations. It augments vision‑language‑action models with historical video observations and explicit camera‑pose supervision, and introduces a scalable human‑robot mid‑training recipe using 1000 hours of egocentric and robotic data. The Active‑perception Mobile‑manipulation Platform (AMP) enables single‑operator teleoperation for scalable demonstration collection, leading to improved success rates on active‑perception tasks.

By Shuai Zhou, Kaisheng Pang, Wenxuan Song, Wenjie Zhang, Xinhu Zheng, Haoang Li
arXiv Computer Vision
Sep 23

TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action

arXiv:2505.01583v2 Announce Type: replace Abstract: Understanding causal event relationships and achieving fine-grained temporal grounding in videos remain challenging for vision-language models (VLM...

By Jen-Hao Cheng, Yi-Hao Peng, Huapeng Zhou, Vivian Wang, Huayu Wang, Hsiang-Wei Huang, Wenhao Chai, Hou-I Liu, Kuang-Ming Chen, Cheng-Yen Yang, Yi-Ling Chen, Vibhav Vineet, Qin Cai, Jenq-Neng Hwang
arXiv Computer Vision
Sep 10

VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Models

arXiv:2609.09396v1 Announce Type: new Abstract: As Vision-Language Models (VLMs) advance toward physical deployment, the focus has remained on action-oriented Embodied AI evaluated on subject-centric...

By Zaid Pervaiz Bhat, Nimra Nayyar, Arihant Jain, Lap Fung Chan, John Suchanek, Yu Wang, Varun Praveen, Tomasz Kornuta, Vidya Nariyambut Murali
arXiv AI
Aug 28

A Very Big Video Reasoning Suite

The paper introduces the Very Big Video Reasoning (VBVR) Dataset, a large-scale collection of over one million video clips organized into 200 curated reasoning tasks. It also presents VBVR-Bench, a benchmark framework that uses rule-based, human-aligned scorers for reproducible evaluation of video reasoning models. The authors conduct a large-scale scaling study, noting early signs of emergent generalization to unseen reasoning tasks, and make all resources publicly available.

By Maijunxian Wang, Ruisi Wang, Juyi Lin, Ran Ji, Thadd\"aus Wiedemer, Qingying Gao, Dezhi Luo, Yaoyao Qian, Lianyu Huang, Zelong Hong, Jiahui Ge, Qianli Ma, Hang He, Yifan Zhou, Lingzi Guo, Lantao Mei, Jiachen Li, Hanwen Xing, Tianqi Zhao, Fengyuan Yu, Weihang Xiao, Yizheng Jiao, Jianheng Hou, Danyang Zhang, Pengcheng Xu, Boyang Zhong, Zehong Zhao, Gaoyun Fang, John Kitaoka, Yile Xu, Hua Xu, Kenton Blacutt, Tin Nguyen, Siyuan Song, Haoran Sun, Shaoyue Wen, Linyang He, Runming Wang, Yanzhi Wang, Mengyue Yang, Ziqiao Ma, Rapha\"el Milli\`ere, Freda Shi, Nuno Vasconcelos, Daniel Khashabi, Alan Yuille, Yilun Du, Ziming Liu, Bo Li, Dahua Lin, Ziwei Liu, Vikash Kumar, Yijiang Li, Lei Yang, Zhongang Cai, Hokin Deng
Hugging Face Trending Papers
Sep 3

SV-WAM: An Efficient Surround-View World-Action Model for End-to-End Autonomous Driving

SV-WAM is a surround‑view world‑action model that keeps all six camera views for autonomous driving while enabling efficient inference by discarding the video branch during deployment. It uses future‑video prediction as dense training supervision and introduces an action‑centered causal mask to prevent future‑video tokens from influencing action tokens during joint denoising. A differentiable drivable‑area compliance regularizer further improves safety by penalizing vehicle‑footprint corners that approach or cross drivable boundaries. Experiments on NAVSIMv2 and nuScenes show state‑of‑the‑art planning performance with low latency and strong zero‑shot transfer.