arXiv AI

WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting

arXiv:2607. 18084v1 Announce Type: new Abstract: Predicting a football match before kickoff requires more than knowing past results: a model must use changing information and make a clear prediction before the answer is available.

arXiv Computation and Language
Sep 14

Information Specialization and Constrained Synthesis in Multi-Agent LLM Forecasting: A Prospective Live-Study of the 2026 FIFA World Cup

The study evaluates a multi‑agent large language model system for forecasting outcomes of the 2026 FIFA World Cup. Two specialist agents—one quantitative and one news‑focused—produce forecasts that are then reviewed by a critic and combined by a meta‑agent. Results show the news specialist performs best, matching betting market accuracy, while the meta‑agent adds little beyond the specialists’ predictions.

By Julian Varghese, Lucas Bickmann, Sarah Sandmann
arXiv Computation and Language
Aug 21

StreamSoccer: Event-Driven Memory for Streaming Soccer Commentary

arXiv:2608. 19723v1 Announce Type: cross Abstract: Streaming video understanding requires models to causally update state as video arrives and organize growing history into semantic units that can evolve, persist, and be recalled under bounded computation and memory.

By Chenxi Shao, Bozhong Wang, Jiaxin Huang, Zhao Liu, Sunwei Zhu, Tianxin Hang, Gaoqi He, Yang Li, Changbo Wang
arXiv AI
Jul 17

SportD: Can VLMs Physically Strategize?

arXiv:2607. 14616v1 Announce Type: new Abstract: Vision--language models have become increasingly capable of interpreting visual scenes, but it remains unclear whether they can use information to make strategically effective decisions.

By Jasin Cekinmez, Addison J. Wu, Haotian Xia, Akshaya Bharadhwaj, Anay Putty, Anirudh Ravishankar, Jaewoong Lee, Jinglin Xiao, Kyumin Andrew Shim, Mishika Ahuja, Nisarga Patil, Leo Liu, Zhuohan Liu, Weining Shen