arXiv:2607. 14616v1 Announce Type: new Abstract: Vision--language models have become increasingly capable of interpreting visual scenes, but it remains unclear whether they can use information to make strategically effective decisions.
By Jasin Cekinmez, Addison J. Wu, Haotian Xia, Akshaya Bharadhwaj, Anay Putty, Anirudh Ravishankar, Jaewoong Lee, Jinglin Xiao, Kyumin Andrew Shim, Mishika Ahuja, Nisarga Patil, Leo Liu, Zhuohan Liu, Weining Shen
arXiv:2607. 17765v1 Announce Type: cross Abstract: We introduce WC2026-Agents, a benchmark and dataset for evaluating large language models (LLMs) as autonomous forecasting agents on real, future events.
By Jiacheng Ding, Cong Guo, Jason Xu
arXiv:2608. 05030v1 Announce Type: new Abstract: Football score forecasting combines a strong statistical core with a difficult contextual edge.
By Shaopeng Liang
The paper presents a statistical and machine‑learning framework for estimating expected goals (xG) and attributing offensive involvement in professional box lacrosse. Using 1,006 manually annotated shot attempts from 13 Rochester Knighthawks games, the authors evaluate logistic regression, random forest, and extremely randomized trees, finding that a contextual baseline random forest achieves the lowest log loss and Brier score. They also introduce a shot‑based expected assist metric and an exploratory expected pick value (xPV) metric, noting that xPV’s magnitude is indistinguishable from noise but shows a directional pattern in permutation tests.
By Robert Jimerson Jr
Football score forecasting combines a strong statistical core with a difficult contextual edge. Dynamic Poisson-family models estimate team strength, expected goals, and coherent score probabilities, but do not directly understand roles, tactical matchups, motivation, or how a first goal changes behaviour.
arXiv:2608. 03416v1 Announce Type: new Abstract: Large language models (LLMs) are now regularly asked to forecast real-world events, but comparisons are often difficult because models receive different information, use different tools, and are evaluated under different rules.
By Jonaid Shianifar, Iias Faiud