The paper investigates the reliability of Expected Goals (xG) as a measure of finishing skill in soccer, arguing that the common practice of comparing cumulative xG to actual goals is flawed. It presents three hypotheses: high variance and small sample sizes make the deviation metric inadequate, including all shot types can mask true finishing ability, and inherent biases in xG models reduce the apparent gap between expected and actual goals for top finishers. Using an AI‑fairness technique to calibrate xG across player subgroups, the authors demonstrate that standard models underestimate Messi’s goal‑adjusted xG (GAX) by 17% and that his GAX is 27% higher than that of typical elite high‑shot‑volume attackers, revealing him as an even more exceptional finisher than previously thought.
By Jesse Davis, Pieter Robberechts
arXiv:2608.29563v1 Announce Type: cross
Abstract: School coaches prepare for opponents with game film and intuition. The analytics tools of professional teams stay out of reach. We ask how far public...
By Yibo Gong, Cong Guo, Jiacheng Ding
arXiv:2607. 26061v1 Announce Type: new Abstract: Pre-match tactical decision-making in professional football relies heavily on subjective expert analysis and identity-based scouting systems that cannot generalize to unseen teams.
By Mouad Zemzoumi, Amine Abouaomar
arXiv:2608. 05030v1 Announce Type: new Abstract: Football score forecasting combines a strong statistical core with a difficult contextual edge.
By Shaopeng Liang
arXiv:2607. 14616v1 Announce Type: new Abstract: Vision--language models have become increasingly capable of interpreting visual scenes, but it remains unclear whether they can use information to make strategically effective decisions.
By Jasin Cekinmez, Addison J. Wu, Haotian Xia, Akshaya Bharadhwaj, Anay Putty, Anirudh Ravishankar, Jaewoong Lee, Jinglin Xiao, Kyumin Andrew Shim, Mishika Ahuja, Nisarga Patil, Leo Liu, Zhuohan Liu, Weining Shen
Football score forecasting combines a strong statistical core with a difficult contextual edge. Dynamic Poisson-family models estimate team strength, expected goals, and coherent score probabilities, but do not directly understand roles, tactical matchups, motivation, or how a first goal changes behaviour.
The study introduces a framework that reconciles minute‑resolution athlete monitoring data with injury labels that are only available at the session level. By creating fixed‑time landmarks (10, 20, 30, 40, 50, and 60 minutes) and generating a single representation per athlete‑session up to each landmark, the authors evaluate several machine‑learning models and data‑augmentation strategies on elite women’s football data. Results show that discrimination varies across landmarks, with TabPFN outperforming Logistic Regression at later landmarks but not consistently beating Random Forest, and that synthetic augmentation offers benefits only in specific conditions.
By Evangelos Chatzidimitriou, Konstantinos Tserpes
arXiv:2608. 12926v1 Announce Type: cross Abstract: Traditional player evaluation in professional handball relies on basic box-score metrics or heuristic indices, which fail to credit the multi-player build-up chain.
By Julius Broermann, Oliver M\"uller, Michael D\"oring, Jochen Baumeister
The paper introduces a method for verifying whether two gameplay replays in Counter‑Strike 2 belong to the same player by extracting a behavioral fingerprint that captures crosshair control, movement‑stop‑fire coordination, economy, combat engagement, and temporal rhythm. Using 1,330 demos and 13,300 observations, the authors train a pairwise model that achieves an ROC AUC of 0.931 and 0.722 recall at 95% precision, with low‑level mechanical habits providing the strongest identity signals. Aggregating multiple demos further improves performance, raising AUC to 0.986 when ten historical demos are considered.
By Xuchen Zhang
Skill Profiling with Attributable Reasoning (SPAR) is a wearable system that uses an eight‑IMU garment and pressure‑insole sensors to classify boxing punches as expert or novice. It provides multi‑tier explanations—joint‑level attributions for analysts, counterfactual kinetic‑chain layers for coaches, and plain‑language narratives for athletes—to enable actionable feedback. In a study of 17 participants and 4,713 punches, SPAR achieved a leave‑one‑participant‑out AUC of 0.842, and thematic analysis of coach interviews identified six key feedback themes.
By Nibraas Khan, Hanchen David Wang, Enya Bullard, Ritam Ghosh, Ruj Haan, Aarav Agrawal, Meiyi Ma, Nilanjan Sarkar
arXiv:2609.25569v1 Announce Type: new
Abstract: Soccer tactics are interactive: an attacking action changes the opponent's defensive problem, and the observed response depends on the multi-agent matc...
By Abel A. Reyes-Angulo, Henry O. Velesaca, Steven Araujo
The paper proposes a method for verifying whether two gameplay replays in Counter‑Strike 2 belong to the same player by extracting a behavioral fingerprint that captures crosshair control, movement, economy, combat, and rhythm. Using a six‑fold evaluation on two datasets (Perfect and Professional), the pairwise model achieves ROC AUCs of 0.926 and 0.956, with aiming and low‑level mechanics providing the strongest identity signals. Aggregating evidence across multiple historical demos further improves account‑history AUC, reaching 0.982 on Perfect and 0.975 on Professional.
By Xuchen Zhang