arXiv:2607. 18084v1 Announce Type: new Abstract: Predicting a football match before kickoff requires more than knowing past results: a model must use changing information and make a clear prediction before the answer is available.
By Zhaokai Wang, Tianlin Gui, Jiayuan Rao, Shangzhe Di, Yihong Tang, Dingli Liang
arXiv:2608. 03416v1 Announce Type: new Abstract: Large language models (LLMs) are now regularly asked to forecast real-world events, but comparisons are often difficult because models receive different information, use different tools, and are evaluated under different rules.
By Jonaid Shianifar, Iias Faiud
Predicting a football match before kickoff requires more than knowing past results: a model must use changing information and make a clear prediction before the answer is available. We present WorldCupArena, a dynamic benchmark for language models and deep-research agents.
arXiv:2607. 17765v1 Announce Type: cross Abstract: We introduce WC2026-Agents, a benchmark and dataset for evaluating large language models (LLMs) as autonomous forecasting agents on real, future events.
By Jiacheng Ding, Cong Guo, Jason Xu
arXiv:2608. 05030v1 Announce Type: new Abstract: Football score forecasting combines a strong statistical core with a difficult contextual edge.
By Shaopeng Liang
Football score forecasting combines a strong statistical core with a difficult contextual edge. Dynamic Poisson-family models estimate team strength, expected goals, and coherent score probabilities, but do not directly understand roles, tactical matchups, motivation, or how a first goal changes behaviour.
The study evaluates a multi‑agent large language model system for forecasting outcomes of the 2026 FIFA World Cup. Two specialist agents—one quantitative and one news‑focused—produce forecasts that are then reviewed by a critic and combined by a meta‑agent. Results show the news specialist performs best, matching betting market accuracy, while the meta‑agent adds little beyond the specialists’ predictions.
By Julian Varghese, Lucas Bickmann, Sarah Sandmann
arXiv:2608. 19723v1 Announce Type: cross Abstract: Streaming video understanding requires models to causally update state as video arrives and organize growing history into semantic units that can evolve, persist, and be recalled under bounded computation and memory.
By Chenxi Shao, Bozhong Wang, Jiaxin Huang, Zhao Liu, Sunwei Zhu, Tianxin Hang, Gaoqi He, Yang Li, Changbo Wang
arXiv:2603. 15212v2 Announce Type: replace Abstract: Evaluating football player transfers is challenging because player actions depend strongly on tactical systems, teammates, and match context.
By Miru Hong, Minho Lee, Geonhee Jo, Hyeokje Cho, Hyunsung Kim, Pascal Bauer, Sang-Ki Ko
arXiv:2606. 18686v1 Announce Type: new Abstract: Forecasting benchmarks for general-purpose AI systems usually inherit the constraints of the real world: outcomes resolve slowly, tail events are rare, and counterfactual questions are difficult to score.
By Jaeho Lee, Nick Merrill, Ezra Karger
LEAP (Likelihood Elicitation and Aggregation for Probabilistic forecasting) is a new approach that reorganizes how evidence is used in LLM-based forecasting systems. Instead of a monolithic prediction that aggregates all evidence at once, LEAP examines each evidence item separately, elicits likelihood parameters, and combines them with an explicit prior to produce a posterior distribution. The method supports continuous, single-choice, and multi-choice forecasts and has been shown to improve prediction and calibration metrics across models on a benchmark covering forecasting, information-seeking, and browsing tasks.
By Yufei Chen, Yiran Zhao, Xiaogang Xu, Qipeng Xie, Jiafei Wu, Zhe Liu
arXiv:2607. 06495v1 Announce Type: cross Abstract: Live sports commentary is grounded generation under a deadline: statements concern real, named athletes, the grounding state changes every few seconds, and no reference text exists at generation time.
By Juan S. Santillana (Independent Researcher)