arXiv AI

From Score Matrices to Football-Aware Match-State Simulation: An Auditable LLM Harness for Exact-Score Reranking

arXiv:2608. 05030v1 Announce Type: new Abstract: Football score forecasting combines a strong statistical core with a difficult contextual edge.

arXiv Computation and Language
Sep 14

Information Specialization and Constrained Synthesis in Multi-Agent LLM Forecasting: A Prospective Live-Study of the 2026 FIFA World Cup

The study evaluates a multi‑agent large language model system for forecasting outcomes of the 2026 FIFA World Cup. Two specialist agents—one quantitative and one news‑focused—produce forecasts that are then reviewed by a critic and combined by a meta‑agent. Results show the news specialist performs best, matching betting market accuracy, while the meta‑agent adds little beyond the specialists’ predictions.

By Julian Varghese, Lucas Bickmann, Sarah Sandmann
Hugging Face Trending Papers
Aug 19

FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents

FM‑Bench is a new benchmark that tests large language model agents on long‑horizon decision‑making by having them run a football club for 20 in‑game years. The agent must manage a squad, trade players, negotiate contracts, invest in facilities, set lineups, and respond to a board that can fire it, all while a deterministic engine aggregates the outcomes into a final score without human judgment. The benchmark evaluates six behavioral capabilities and compares 15 frontier models in solo and arena tracks, revealing that managerial behavior—not computational scale—drives performance.

arXiv AI
Jul 17

SportD: Can VLMs Physically Strategize?

arXiv:2607. 14616v1 Announce Type: new Abstract: Vision--language models have become increasingly capable of interpreting visual scenes, but it remains unclear whether they can use information to make strategically effective decisions.

By Jasin Cekinmez, Addison J. Wu, Haotian Xia, Akshaya Bharadhwaj, Anay Putty, Anirudh Ravishankar, Jaewoong Lee, Jinglin Xiao, Kyumin Andrew Shim, Mishika Ahuja, Nisarga Patil, Leo Liu, Zhuohan Liu, Weining Shen
arXiv Machine Learning
Aug 31

Biases in Expected Goals Models Confound Finishing Ability

The paper investigates the reliability of Expected Goals (xG) as a measure of finishing skill in soccer, arguing that the common practice of comparing cumulative xG to actual goals is flawed. It presents three hypotheses: high variance and small sample sizes make the deviation metric inadequate, including all shot types can mask true finishing ability, and inherent biases in xG models reduce the apparent gap between expected and actual goals for top finishers. Using an AI‑fairness technique to calibrate xG across player subgroups, the authors demonstrate that standard models underestimate Messi’s goal‑adjusted xG (GAX) by 17% and that his GAX is 27% higher than that of typical elite high‑shot‑volume attackers, revealing him as an even more exceptional finisher than previously thought.

By Jesse Davis, Pieter Robberechts