SportsGrounder: Proposal-Aided Interleaved Grounding for Dense Sports Video Reasoning
arXiv:2608. 07932v2 Announce Type: replace Abstract: Sports video analysis is crucial for athletic analytics and broadcasting enhancement.
arXiv:2606. 09327v1 Announce Type: cross Abstract: Football event data constitute a rich spatiotemporal source for quantitative analysis of player actions in team sports.
arXiv:2608. 07932v2 Announce Type: replace Abstract: Sports video analysis is crucial for athletic analytics and broadcasting enhancement.
arXiv:2607. 21267v1 Announce Type: new Abstract: Comprehensive basketball video understanding requires resolving not only what event occurs, but also who is responsible and when the key evidence appears.
Precise Event Spotting (PES) requires distinguishing visually similar yet semantically distinct adjacent frames, making it fundamentally different from image classification and coarse action recognition. Although self-distillation methods such as DINO have shown strong representation learning ability in images, we find that directly applying them to PES is ineffective: without supervised guidance, subtle but crucial motion cues are often suppressed as noise, leading to representations that are insensitive to precise event boundaries.
arXiv:2608. 19646v1 Announce Type: new Abstract: Visual understanding in sports has emerged as a hot topic in computer vision in recent years.
Action Quality Assessment (AQA) aims to objectively evaluate performance quality from action videos. Most existing methods follow a ``one-by-one'' paradigm, training a separate model for each action type.
arXiv:2607. 21290v1 Announce Type: cross Abstract: Multi-task learning (MTL) is a promising approach for prediction tasks derived from video game state data, as modern game telemetry provides multiple related supervision signals from the same structured observations.
arXiv:2606. 09289v1 Announce Type: new Abstract: Understanding tactical organisation of association football, hereafter referred to as football, requires identifying distinct match phases.
arXiv:2606. 11860v1 Announce Type: new Abstract: In this paper, we introduce Representation Prediction via Autoencoding using Iterative Refinement (RePAIR) - a novel self-supervised representation learning architecture that synthesizes Masked Autoencoders (MAE), Joint Embedding Predictive Architectures (JEPA), and Bidirectional Encoder Representations from Transformers (BERT).
arXiv:2603. 15212v2 Announce Type: replace Abstract: Evaluating football player transfers is challenging because player actions depend strongly on tactical systems, teammates, and match context.
arXiv:2505.01583v2 Announce Type: replace Abstract: Understanding causal event relationships and achieving fine-grained temporal grounding in videos remain challenging for vision-language models (VLM...
arXiv:2608.09200v3 Announce Type: replace Abstract: Live basketball commentary generation requires determining when an event is sufficiently observable and describing it before subsequent events unfo...
arXiv:2606. 07433v1 Announce Type: cross Abstract: Video understanding is being rapidly transformed by multimodal large language models (MLLMs), as research moves from short clips to long, multimodal, and knowledge-intensive video scenarios.