arXiv AI By Jiacheng Ding, Cong Guo, Jason Xu

FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches

Read the original on arXiv AI →

arXiv:2607. 17765v1 Announce Type: cross Abstract: We introduce WC2026-Agents, a benchmark and dataset for evaluating large language models (LLMs) as autonomous forecasting agents on real, future events.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.