arXiv AI By Jiacheng Ding, Cong Guo, Jason Xu

FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches

Read the original on arXiv AI →

arXiv:2607. 17765v1 Announce Type: cross Abstract: We introduce WC2026-Agents, a benchmark and dataset for evaluating large language models (LLMs) as autonomous forecasting agents on real, future events.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 14

Information Specialization and Constrained Synthesis in Multi-Agent LLM Forecasting: A Prospective Live-Study of the 2026 FIFA World Cup

The study evaluates a multi‑agent large language model system for forecasting outcomes of the 2026 FIFA World Cup. Two specialist agents—one quantitative and one news‑focused—produce forecasts that are then reviewed by a critic and combined by a meta‑agent. Results show the news specialist performs best, matching betting market accuracy, while the meta‑agent adds little beyond the specialists’ predictions.

By Julian Varghese, Lucas Bickmann, Sarah Sandmann