arXiv AI

FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches

arXiv:2607. 17765v1 Announce Type: cross Abstract: We introduce WC2026-Agents, a benchmark and dataset for evaluating large language models (LLMs) as autonomous forecasting agents on real, future events.

arXiv Computation and Language
Sep 14

Information Specialization and Constrained Synthesis in Multi-Agent LLM Forecasting: A Prospective Live-Study of the 2026 FIFA World Cup

The study evaluates a multi‑agent large language model system for forecasting outcomes of the 2026 FIFA World Cup. Two specialist agents—one quantitative and one news‑focused—produce forecasts that are then reviewed by a critic and combined by a meta‑agent. Results show the news specialist performs best, matching betting market accuracy, while the meta‑agent adds little beyond the specialists’ predictions.

By Julian Varghese, Lucas Bickmann, Sarah Sandmann
arXiv AI
Aug 20

FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents

FM‑Bench is a new benchmark that tests large language model agents on long‑horizon decision making by having them run a football club for 20 in‑game years. The agent must manage a squad, trade players, negotiate contracts, invest in facilities and youth, set lineups, and respond to a board that can fire it, all using 26 tools and roughly 340–400 decision stops, with a deterministic engine producing a final score without human or LLM judges. The benchmark includes a solo track where each of 15 frontier models competes against a frozen scripted world, and an Arena track where the same models plus a scripted anchor share one 20‑year world, allowing the first head‑to‑head evaluation at this scale. whyItMatters":"FM‑Bench provides a rigorous, large‑scale test of sustained, cumulative decision‑making in language‑model agents, revealing that managerial strategy—not computational scale or vendor—drives performance over long horizons."

By Tianyou Wang, Chongyang Gao, Kezhen Chen, Chen Dong, Yinghao He, Donghan Li, Wangcheng Xu, Hongjiu Zhang, Chi Li
Hugging Face Trending Papers
Aug 19

FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents

FM‑Bench is a new benchmark that tests large language model agents on long‑horizon decision‑making by having them run a football club for 20 in‑game years. The agent must manage a squad, trade players, negotiate contracts, invest in facilities, set lineups, and respond to a board that can fire it, all while a deterministic engine aggregates the outcomes into a final score without human judgment. The benchmark evaluates six behavioral capabilities and compares 15 frontier models in solo and arena tracks, revealing that managerial behavior—not computational scale—drives performance.

arXiv Machine Learning
Sep 24

Evaluation Choices Decide the Forecasting Leaderboard: Evidence from a Production Marketplace Panel

The paper demonstrates that the outcome of a forecasting leaderboard is largely determined by the evaluator’s design choices rather than the models themselves. By fixing the data, horizon, and period, the authors varied three key evaluation decisions—unit of analysis, error pooling, and scoring metric—and showed that each can reverse or eliminate the apparent superiority of any forecasting method. The study also evaluates the practical impact of these choices on a deployed system, revealing that the selection rule captures a significant portion of the potential performance gain, and confirms the findings on an external public dataset.

By Md Rezwanul Islam, Wael Mohammed