arXiv Machine Learning

ProactiveBench: Can Streaming Video Models Really Interact Like Humans?

ProactiveBench evaluates streaming video models on their ability to interact proactively, rather than reactively. It tests models at one‑second intervals without explicit cues, using six subtasks that vary trigger ambiguity, timing tolerance, and response patterns. The benchmark measures both response and silence rates, distinguishing early, in‑window, and missed responses, and penalizes omissions and repetitions.

arXiv Machine Learning
Sep 23

LiveProBench: Can Streaming Video Models Really Interact Like Humans?

LiveProBench evaluates streaming video models on their ability to interact proactively, assessing whether they respond at appropriate times without explicit cues. The benchmark tests models at one‑second intervals across six subtasks that vary trigger ambiguity and timing tolerance, measuring response accuracy, silence rates, and duplicate responses. Results show that many models issue premature responses more often than missed ones, highlighting a significant shortfall in human‑like temporal decision making.

By Kaixuan Du, Xin Wan, Hang Zhang, Meng Cao, Dai Guan, Ming Chen, YuKun Wang
arXiv Computation and Language
6d ago

TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding

TRACE (Temporal Audit and Condition-aware Evaluation) is a new benchmark and evaluation framework for streaming video understanding that explicitly records when evidence becomes valid, how visual history is maintained, and how responses are triggered. It combines temporally audited visual tasks, evidence timing, instruction-dependent trigger annotations, a unified causal Core–Adapter protocol, and multidimensional reporting of answer quality, timeliness, response-selection behavior, workload, completion, and reliability. Using 1,240 records from 517 videos, TRACE evaluated eight publicly available models in eight configurations, revealing that similar QA accuracy can hide significant differences in completion, answer validity, generation workload, and proactive performance metrics such as response delay, false alarms, and missed target windows.

By Yibo Ma, Qianqian Zhang, Peng Liu, Tiancheng Zhao
arXiv Machine Learning
Jul 1

Event-Driven Video Generation

arXiv:2603. 13402v3 Announce Type: replace-cross Abstract: Current text-to-video models can make individual frames look convincing while still getting simple interactions wrong: objects move before contact, an intended action is skipped, a placed object keeps drifting, or a support relation breaks.

By Chika Maduabuchi, Jindong Wang
arXiv Machine Learning
Sep 23

TimeInteract: Towards Real-Time Interactive Intelligence for Streaming Time Series

TimeInteract introduces a new regime called Time-Series Interaction, enabling models to continuously perceive incoming time-series data and user intent, decide when to respond, and keep processing new observations during response generation. The system employs a dual-view streaming encoder, a response control mechanism, and a decoupled inference pipeline to avoid blocking. Evaluated on the newly created StreamTSI-34K dataset, TimeInteract outperforms existing LLMs, VLMs, and TSLMs across four interaction levels, achieving significant gains in accuracy, response triggering, and inference speed.

By Sheng Pan, Yongli Gu, Yiqing Guo, Warren Jin, Bo Du, Shirui Pan, Ming Jin
arXiv Computer Vision
Sep 3

The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation

The paper introduces Temporal Context Routing (TCR), a method that aligns script timing with the shared temporal axis used for audio and video generation. By routing prompt guidance to specific temporal positions in both modalities, TCR dramatically improves shot boundary accuracy and dialogue timing on a set of test scripts, while preserving visual quality and audio‑visual sync. A user study confirms that participants prefer TCR across multiple evaluation dimensions.

By Yichen Liu, Quanwei Zhang, Haozhe Wang, Donghao Zhou, Xiaojie Li, Yang Shi, Jiaming Liu, Ruihua Huang, Yingtian Zou, Daquan Zhou
arXiv AI
Sep 18

AVTrace: Diagnosing Audio-Visual Temporal Reasoning in Omni Models

AVTrace is a diagnostic suite designed to evaluate audio‑visual temporal reasoning in omni models. It covers tasks such as onset and span grounding, synchronization, next‑step prediction, cross‑modal localization, chain parsing, and event‑conditioned comprehension, providing 34,114 training examples and balanced development and test splits. Five open omni models were tested, all scoring below the majority‑label baseline on synchronization verification and showing low performance on chain parsing and event‑conditioned tasks, while parameter‑efficient temporal post‑training improved some metrics.

By Longyin Zhang, Parth Sakhare Mahendra, Chengwei Wei, Ning Zhang, Lim Ming Chong, Sirui He, Ai Ti Aw
arXiv AI
Sep 2

TempCloze: Can Video-LLMs Identify the Missing Middle?

TempCloze is a video cloze benchmark designed to evaluate visual temporal reasoning in Video-LLMs. The task presents a video’s beginning and ending clips and asks models to select the correct missing middle from four candidates, focusing on semantic, alignment, and progression aspects while minimizing appearance cues. Evaluation of 31 models shows that temporal alignment is the main challenge, with models performing better on semantic content and event progression but struggling to place events correctly in time.

By Wenqi Pei, Henry Hengyuan Zhao, Yilai Liu, Jiahao Meng, Han Chen, Ziyu Wang, Hongyang Du