OpenSAL360 is an open‑source platform that enables scalable, low‑cost collection of 360° video saliency data using only a standard screen, mouse, and internet connection. It bypasses the need for VR headsets, allowing parallel data collection from crowdsourced assessors. The authors validated the protocol against seven VR eye‑tracking datasets, performed ablation studies, and released a new dataset of 500 omnidirectional videos annotated by over 2,000 assessors, the largest in the field to date.
By Alexey Bryncev, Andrey Moskalenko, Kira Shilovskaya, Ivan Kosmynin, Dmitriy Vatolin
arXiv:2608. 08947v1 Announce Type: cross Abstract: Current hazard detection systems in autonomous driving may develop mesa objectives, learned internal goals that achieve high training performance through spurious correlations rather than genuine hazard recognition.
By Lennox Anderson, Ahmed Boutar, Jonah Mulcrone, Tal Erez
EgoHRV is a method that estimates heart rate variability (HRV) and heart rate (HR) from the gaze cameras in egocentric headsets. It uses a 3D backbone and a low–high decomposition module to extract the blood volume pulse signal from gaze video, and aligns frequency‑domain representations of contact‑based and camera‑derived signals through cross‑domain pretraining. The approach achieves state‑of‑the‑art accuracy for HR and HRV estimation and, when integrated into EgoExo4D’s proficiency estimator, improves accuracy by 17.8%.
TRACE (Temporal Audit and Condition-aware Evaluation) is a new benchmark and evaluation framework for streaming video understanding that explicitly records when evidence becomes valid, how visual history is maintained, and how responses are triggered. It combines temporally audited visual tasks, evidence timing, instruction-dependent trigger annotations, a unified causal Core–Adapter protocol, and multidimensional reporting of answer quality, timeliness, response-selection behavior, workload, completion, and reliability. Using 1,240 records from 517 videos, TRACE evaluated eight publicly available models in eight configurations, revealing that similar QA accuracy can hide significant differences in completion, answer validity, generation workload, and proactive performance metrics such as response delay, false alarms, and missed target windows.
By Yibo Ma, Qianqian Zhang, Peng Liu, Tiancheng Zhao
arXiv:2610.00922v1 Announce Type: new
Abstract: Gaze estimation under natural head-eye motion underpins applications from driver monitoring to human-computer interaction. Single-frame methods predict...
By Jungmin Lee, Niamat Ullah, Yoseob Han
arXiv:2610.00960v1 Announce Type: new
Abstract: A video benchmark should reward the capability it claims to measure, yet models can exploit answer options, question text, or partial visual evidence....
By Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
arXiv:2509.01167v3 Announce Type: replace-cross
Abstract: Vision-language models (VLMs) can ingest only a limited number of video frames, making frame selection a practical necessity. But do current...
By Hyunjong Ok, Jaeho Lee
arXiv:2610.02205v1 Announce Type: new
Abstract: Programmable world models separate executable dynamics from visual generation, offering a promising foundation for next-generation game engines. Howeve...
By Zheng-Hui Huang, Guixu Lin, Yu-Ju Tsai, Jian-Kai Zhu, Fengbo Lan, Yu-Lun Liu, Yung-Yu Chuang, Kaipeng Zhang, Zhixiang Wang
ProactiveBench evaluates streaming video models on their ability to interact proactively, rather than reactively. It tests models at one‑second intervals without explicit cues, using six subtasks that vary trigger ambiguity, timing tolerance, and response patterns. The benchmark measures both response and silence rates, distinguishing early, in‑window, and missed responses, and penalizes omissions and repetitions.
By Kaixuan Du, Xin Wan, YuKun Wang, Hang Zhang, Meng Cao, Dai Guan, Ming Chen, Ni Li
arXiv:2607. 15621v1 Announce Type: cross Abstract: Large language models bring instruction following and scene reasoning to end-to-end driving, but their inference latency collides with the control rate a vehicle requires.
By Yun Li, Jiachen Gong, Simon Thompson, Ehsan Javanmardi, Qunli Zhang, Zifan Zeng, Shiming Liu, Peng Wang, Zixuan Guo, Manabu Tsukada
The paper introduces the Causal Context-Gated Forecaster (CCGF) for predicting a driver's gaze during dashboard-mounted tracker dropouts. CCGF uses a 60‑frame history of gaze and head pose combined with DINOv3 scene features, and a learned reliability gate adjusts the influence of these inputs as the dropout progresses. Experiments on 2,047 naturalistic driving events show that live scene updates reduce median error by 33% compared to history‑only forecasting, while frozen scene input yields higher error, demonstrating the value of real‑time scene information.
By Shabnam Shabani, Ghazal Farhani
LiveProBench evaluates streaming video models on their ability to interact proactively, assessing whether they respond at appropriate times without explicit cues. The benchmark tests models at one‑second intervals across six subtasks that vary trigger ambiguity and timing tolerance, measuring response accuracy, silence rates, and duplicate responses. Results show that many models issue premature responses more often than missed ones, highlighting a significant shortfall in human‑like temporal decision making.
By Kaixuan Du, Xin Wan, Hang Zhang, Meng Cao, Dai Guan, Ming Chen, YuKun Wang