arXiv Computer Vision

NBA_Streaming: A Large-Scale Benchmark for Fine-Grained Basketball Commentary Generation in Continuous Streams

arXiv Computation and Language
Aug 21

StreamSoccer: Event-Driven Memory for Streaming Soccer Commentary

arXiv:2608. 19723v1 Announce Type: cross Abstract: Streaming video understanding requires models to causally update state as video arrives and organize growing history into semantic units that can evolve, persist, and be recalled under bounded computation and memory.

By Chenxi Shao, Bozhong Wang, Jiaxin Huang, Zhao Liu, Sunwei Zhu, Tianxin Hang, Gaoqi He, Yang Li, Changbo Wang
arXiv AI
Sep 18

Semantic Action Graph: A Shared Representation for Agent Grounding and Human Interpretation of Sports Highlights

The paper introduces the Semantic Action Graph, a lightweight domain schema that models a sports match using performer, action, recipient, moment, and state nodes linked by role, temporal, and outcome edges. This structure supports both an agentic pipeline for generating narrated highlights and a visual interface that lets viewers query and inspect the same representation. In a prototype called SportSAGE, 12 soccer fans reported satisfaction with the generated highlights and used the graph interface to search, navigate, and interpret match moments.

By Tica Lin, Deepak Chandran, Gauri Jagatap, Chen Chen, Andrea Fanelli, David Gunawan, Josh Kimball
arXiv Computer Vision
Aug 21

ID-VTG: Image-Disambiguated Video Temporal Grounding

arXiv:2608. 20127v1 Announce Type: new Abstract: Video Temporal Grounding (VTG) faces significant challenges when natural language queries must distinguish between multiple events involving visually similar entities, particularly when relying on fine-grained visual attributes that are difficult to describe accurately in words alone.

By Minghang Zheng, Jingli Wei, Hongyi Yang, Yang Liu
arXiv Computer Vision
Sep 24

MultiVENT-Raw: A Benchmark for Retrieval and Reasoning over Raw Videos

MultiVENT‑Raw is a new multilingual benchmark comprising nearly 120,000 raw videos—continuous footage from cell phones, hand‑held cameras, or CCTV—totaling over 5,300 hours. The dataset includes 130 events and 222 event‑centric queries, along with human‑annotated relevance judgments and extracted key facts for relevant videos. It supports two tasks: retrieving videos relevant to a query event and generating a coherent report summarizing event‑related videos for a target user, with baseline models showing these tasks remain challenging.

By Reno Kriz, David Etter, Alexander Martin, Cameron Carpenter, Debashish Chakraborty, Hannah Recknor, Reihaneh Iranmanesh, Matthew Maciejewski, Kenton Murray, Eugene Yang, Benjamin Van Durme, Aaron Steven White, Andrew Yates, William Walden
arXiv Computer Vision
2d ago

OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction

OneStreamer is a streaming video model that jointly learns to record evidence and respond to tasks through a shared proactive generation process. Its Proactive Hierarchical Caption Memory creates time‑grounded local‑detail captions and event summaries, while Proactive State Transition Learning reduces waiting states by supervising all output anchors. The authors also built a large OneStreamer‑1M dataset and show that a 4B model outperforms baselines on eight streaming video benchmarks, with ablations confirming the benefits of generated captions and PSTL.

By Xiangyu Zeng, Yuandong Yang, Zhiqiu Zhang, Yuhan Zhu, Xinhao Li, Qingyi Si, Dingyu Yao, Changlian Ma, Haoran Chen, Xinyu Chen, Yansong Shi, Junhao Zhou, Yifei Li, Jun Zhang, Chuanyu Qin, Chenxu Yang, Xinlei Yu, Kun Ouyang, Yuchen Shao, Qianshan Wei, Changhai Zhou, Jun Gao, Jiaqi Wang, Limin Wang