arXiv AI

SportD: Can VLMs Physically Strategize?

arXiv:2607. 14616v1 Announce Type: new Abstract: Vision--language models have become increasingly capable of interpreting visual scenes, but it remains unclear whether they can use information to make strategically effective decisions.

arXiv Machine Learning
Aug 31

Biases in Expected Goals Models Confound Finishing Ability

The paper investigates the reliability of Expected Goals (xG) as a measure of finishing skill in soccer, arguing that the common practice of comparing cumulative xG to actual goals is flawed. It presents three hypotheses: high variance and small sample sizes make the deviation metric inadequate, including all shot types can mask true finishing ability, and inherent biases in xG models reduce the apparent gap between expected and actual goals for top finishers. Using an AI‑fairness technique to calibrate xG across player subgroups, the authors demonstrate that standard models underestimate Messi’s goal‑adjusted xG (GAX) by 17% and that his GAX is 27% higher than that of typical elite high‑shot‑volume attackers, revealing him as an even more exceptional finisher than previously thought.

By Jesse Davis, Pieter Robberechts
arXiv AI
Jun 10

Monte Carlo Pass Search: Using Trajectory Generation for 3D Counterfactual Pass Evaluation in Football

arXiv:2606. 11120v1 Announce Type: new Abstract: We recast pass evaluation in football (soccer) as a Monte Carlo Tree Search (MCTS)-like evaluation problem whose components mostly exist in the literature under different names: a value model (possession value), a world model (multi-agent trajectories with ball interactions), and a policy over counterfactual actions (sampling pass variants with noise).

By Andrew Kang, Priya Narasimhan
arXiv AI
Sep 18

Can Vision-Language Models Judge Olympic Diving? From Reasoning to Scores in Zero-Shot Action Quality Assessment

The paper investigates whether open‑source Vision‑Language Models (VLMs) can perform zero‑shot action quality assessment (AQA) on Olympic diving videos. Using the AQA‑7 benchmark, the authors propose a regression framework that combines VLM‑generated semantic reasoning, phase‑level sub‑scores, TF‑IDF vectorization, dimensionality reduction, and ensemble learning to predict final competition scores. While individual VLMs achieve moderate Spearman correlations (<0.32), the ensemble approach boosts performance to 0.67, demonstrating that VLM‑derived textual reasoning features are more informative than raw numerical sub‑scores for AQA. whyItMatters":"The study shows that VLMs can serve as explainable, semi‑automated tools for evaluating sports performance, potentially aiding expert judging in complex, subjective Olympic events."

By Henry O. Velesaca, David Freire-Obregon, Luigi Miranda, Abel Reyes-Angulo