Prompt, Probe, Train, or Annotate? Single-camera sports video understanding in amateur settings
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
Video understanding is usually benchmarked on curated, single-actor, or professionally filmed clips, and a strong score there is routinely read as evidence a model is robust enough for deployment. Ama...
arXiv:2607. 14616v1 Announce Type: new Abstract: Vision--language models have become increasingly capable of interpreting visual scenes, but it remains unclear whether they can use information to make strategically effective decisions.
arXiv:2607. 21267v1 Announce Type: new Abstract: Comprehensive basketball video understanding requires resolving not only what event occurs, but also who is responsible and when the key evidence appears.
The paper investigates whether open‑source Vision‑Language Models (VLMs) can perform zero‑shot action quality assessment (AQA) on Olympic diving videos. Using the AQA‑7 benchmark, the authors propose a regression framework that combines VLM‑generated semantic reasoning, phase‑level sub‑scores, TF‑IDF vectorization, dimensionality reduction, and ensemble learning to predict final competition scores. While individual VLMs achieve moderate Spearman correlations (<0.32), the ensemble approach boosts performance to 0.67, demonstrating that VLM‑derived textual reasoning features are more informative than raw numerical sub‑scores for AQA. whyItMatters":"The study shows that VLMs can serve as explainable, semi‑automated tools for evaluating sports performance, potentially aiding expert judging in complex, subjective Olympic events."
arXiv:2609.37938v1 Announce Type: cross Abstract: Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggr...
Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggregate cross-video accuracy conflates failures of lo...