LiveProBench evaluates streaming video models on their ability to interact proactively, assessing whether they respond at appropriate times without explicit cues. The benchmark tests models at one‑second intervals across six subtasks that vary trigger ambiguity and timing tolerance, measuring response accuracy, silence rates, and duplicate responses. Results show that many models issue premature responses more often than missed ones, highlighting a significant shortfall in human‑like temporal decision making.
By Kaixuan Du, Xin Wan, Hang Zhang, Meng Cao, Dai Guan, Ming Chen, YuKun Wang
TRACE (Temporal Audit and Condition-aware Evaluation) is a new benchmark and evaluation framework for streaming video understanding that explicitly records when evidence becomes valid, how visual history is maintained, and how responses are triggered. It combines temporally audited visual tasks, evidence timing, instruction-dependent trigger annotations, a unified causal Core–Adapter protocol, and multidimensional reporting of answer quality, timeliness, response-selection behavior, workload, completion, and reliability. Using 1,240 records from 517 videos, TRACE evaluated eight publicly available models in eight configurations, revealing that similar QA accuracy can hide significant differences in completion, answer validity, generation workload, and proactive performance metrics such as response delay, false alarms, and missed target windows.
By Yibo Ma, Qianqian Zhang, Peng Liu, Tiancheng Zhao
arXiv:2607. 01751v1 Announce Type: cross Abstract: Existing medical video benchmarks primarily evaluate whether a model produces the correct answer, but rarely assess whether it answers at the right time.
By Yuan Wang, Shujian Gao, Songtao Jiang, Zhengyu Hu, Zuozhu Liu
arXiv:2603. 13402v3 Announce Type: replace-cross Abstract: Current text-to-video models can make individual frames look convincing while still getting simple interactions wrong: objects move before contact, an intended action is skipped, a placed object keeps drifting, or a support relation breaks.
By Chika Maduabuchi, Jindong Wang
arXiv:2608. 06361v1 Announce Type: new Abstract: Real-world video benchmarks provide broad coverage, but their fixed clips entangle event count, rate, duration, and visual complexity, making failure modes hard to isolate.
By Sarvesh Baskar, Zikui Cai, Shayan Shabihi, Anirudh Satheesh, Muhammad R. Islam, Udari Madhushani Sehwag, Tom Goldstein, Furong Huang
arXiv:2607. 25961v1 Announce Type: cross Abstract: Ambivalence and hesitancy (A/H) are conflicting affective states that precede the delay or abandonment of health behaviour change.
By Podakanti Satyajith Chary, Barath Parthiban, Pranesh Velmurugan, Adeeba Khan, Nagarajan Ganapathy
arXiv:2606. 06991v1 Announce Type: cross Abstract: Online Video Large Language Models (Video-LLMs) have advanced toward seamless human-AI interaction through frame-by-frame processing and proactive responding.
By Zhenyu Yang, Kairui Zhang, Shengsheng Qian, Weiming Dong, Changsheng Xu
arXiv:2606. 14777v1 Announce Type: cross Abstract: Many moments in the real world do not wait for a user to ask.
By Dingyu Yao, Junhao Zhou, Chenxu Yang, Chuanyu Qin, Haowen Hou, Zheming Liang, Congcong Wang, Yuhang Cao, Shenglong Ye, Shuai Xie, Shuhuan Gu, Haoyang Huang, Qingyi Si, Nan Duan, Jiaqi Wang
TimeInteract introduces a new regime called Time-Series Interaction, enabling models to continuously perceive incoming time-series data and user intent, decide when to respond, and keep processing new observations during response generation. The system employs a dual-view streaming encoder, a response control mechanism, and a decoupled inference pipeline to avoid blocking. Evaluated on the newly created StreamTSI-34K dataset, TimeInteract outperforms existing LLMs, VLMs, and TSLMs across four interaction levels, achieving significant gains in accuracy, response triggering, and inference speed.
By Sheng Pan, Yongli Gu, Yiqing Guo, Warren Jin, Bo Du, Shirui Pan, Ming Jin
The paper introduces Temporal Context Routing (TCR), a method that aligns script timing with the shared temporal axis used for audio and video generation. By routing prompt guidance to specific temporal positions in both modalities, TCR dramatically improves shot boundary accuracy and dialogue timing on a set of test scripts, while preserving visual quality and audio‑visual sync. A user study confirms that participants prefer TCR across multiple evaluation dimensions.
By Yichen Liu, Quanwei Zhang, Haozhe Wang, Donghao Zhou, Xiaojie Li, Yang Shi, Jiaming Liu, Ruihua Huang, Yingtian Zou, Daquan Zhou
AVTrace is a diagnostic suite designed to evaluate audio‑visual temporal reasoning in omni models. It covers tasks such as onset and span grounding, synchronization, next‑step prediction, cross‑modal localization, chain parsing, and event‑conditioned comprehension, providing 34,114 training examples and balanced development and test splits. Five open omni models were tested, all scoring below the majority‑label baseline on synchronization verification and showing low performance on chain parsing and event‑conditioned tasks, while parameter‑efficient temporal post‑training improved some metrics.
By Longyin Zhang, Parth Sakhare Mahendra, Chengwei Wei, Ning Zhang, Lim Ming Chong, Sirui He, Ai Ti Aw
TempCloze is a video cloze benchmark designed to evaluate visual temporal reasoning in Video-LLMs. The task presents a video’s beginning and ending clips and asks models to select the correct missing middle from four candidates, focusing on semantic, alignment, and progression aspects while minimizing appearance cues. Evaluation of 31 models shows that temporal alignment is the main challenge, with models performing better on semantic content and event progression but struggling to place events correctly in time.
By Wenqi Pei, Henry Hengyuan Zhao, Yilai Liu, Jiahao Meng, Han Chen, Ziyu Wang, Hongyang Du