arXiv:2607. 14616v1 Announce Type: new Abstract: Vision--language models have become increasingly capable of interpreting visual scenes, but it remains unclear whether they can use information to make strategically effective decisions.
By Jasin Cekinmez, Addison J. Wu, Haotian Xia, Akshaya Bharadhwaj, Anay Putty, Anirudh Ravishankar, Jaewoong Lee, Jinglin Xiao, Kyumin Andrew Shim, Mishika Ahuja, Nisarga Patil, Leo Liu, Zhuohan Liu, Weining Shen
arXiv:2607. 17765v1 Announce Type: cross Abstract: We introduce WC2026-Agents, a benchmark and dataset for evaluating large language models (LLMs) as autonomous forecasting agents on real, future events.
By Jiacheng Ding, Cong Guo, Jason Xu
arXiv:2608. 05030v1 Announce Type: new Abstract: Football score forecasting combines a strong statistical core with a difficult contextual edge.
By Shaopeng Liang
The paper presents a statistical and machine‑learning framework for estimating expected goals (xG) and attributing offensive involvement in professional box lacrosse. Using 1,006 manually annotated shot attempts from 13 Rochester Knighthawks games, the authors evaluate logistic regression, random forest, and extremely randomized trees, finding that a contextual baseline random forest achieves the lowest log loss and Brier score. They also introduce a shot‑based expected assist metric and an exploratory expected pick value (xPV) metric, noting that xPV’s magnitude is indistinguishable from noise but shows a directional pattern in permutation tests.
By Robert Jimerson Jr
Football score forecasting combines a strong statistical core with a difficult contextual edge. Dynamic Poisson-family models estimate team strength, expected goals, and coherent score probabilities, but do not directly understand roles, tactical matchups, motivation, or how a first goal changes behaviour.
arXiv:2608. 03416v1 Announce Type: new Abstract: Large language models (LLMs) are now regularly asked to forecast real-world events, but comparisons are often difficult because models receive different information, use different tools, and are evaluated under different rules.
By Jonaid Shianifar, Iias Faiud
arXiv:2607. 18084v1 Announce Type: new Abstract: Predicting a football match before kickoff requires more than knowing past results: a model must use changing information and make a clear prediction before the answer is available.
By Zhaokai Wang, Tianlin Gui, Jiayuan Rao, Shangzhe Di, Yihong Tang, Dingli Liang
arXiv:2607. 26061v1 Announce Type: new Abstract: Pre-match tactical decision-making in professional football relies heavily on subjective expert analysis and identity-based scouting systems that cannot generalize to unseen teams.
By Mouad Zemzoumi, Amine Abouaomar
arXiv:2607. 11548v1 Announce Type: cross Abstract: Spatial football metrics such as pitch control assume access to the positions of all 22 players, yet the most widely available source of positional data -- the broadcast main camera -- shows only 10-16 of them at any moment.
By Seongjin Choi
arXiv:2607. 14616v3 Announce Type: replace Abstract: Vision-language models (VLMs) can describe a scene, but can they act well within one?
By Jasin Cekinmez, Addison J. Wu, Haotian Xia, Kyumin Andrew Shim, Anay Putty, Jinglin Xiao, Zhuohan Liu, Leo Liu, Weining Shen
Predicting a football match before kickoff requires more than knowing past results: a model must use changing information and make a clear prediction before the answer is available. We present WorldCupArena, a dynamic benchmark for language models and deep-research agents.
The article argues that for frontier language models, precision—how consistently outputs cluster around the target—should be the key metric rather than capability, which measures average performance. It proposes a simple, non‑circular method to quantify precision by repeatedly scoring deterministic tasks and computing outcome consistency, and demonstrates how this metric can guide decisions about model improvements. The study shows that precision can reveal whether failures are due to systemic misalignment or random noise, and that real‑world measurement is more valuable than rule‑based benchmarks.
By George Andrikopoulos