arXiv:2607. 17765v1 Announce Type: cross Abstract: We introduce WC2026-Agents, a benchmark and dataset for evaluating large language models (LLMs) as autonomous forecasting agents on real, future events.
By Jiacheng Ding, Cong Guo, Jason Xu
The paper demonstrates that the outcome of a forecasting leaderboard is largely determined by the evaluator’s design choices rather than the models themselves. By fixing the data, horizon, and period, the authors varied three key evaluation decisions—unit of analysis, error pooling, and scoring metric—and showed that each can reverse or eliminate the apparent superiority of any forecasting method. The study also evaluates the practical impact of these choices on a deployed system, revealing that the selection rule captures a significant portion of the potential performance gain, and confirms the findings on an external public dataset.
By Md Rezwanul Islam, Wael Mohammed
arXiv:2609. 20193v1 Announce Type: new Abstract: Retrieval plug-ins supply a deep forecaster with information its lookback window cannot carry.
By Mert Onur Cakiroglu, Elham Buxton, Mehmet Dalkilic, Hasan Kurban
arXiv:2607. 10972v1 Announce Type: new Abstract: Many evaluations of model outputs rely either on contracts checkable at evaluation time or on feedback that arrives within the operating loop.
By Aleh Manchuliantsau
arXiv:2608. 03416v1 Announce Type: new Abstract: Large language models (LLMs) are now regularly asked to forecast real-world events, but comparisons are often difficult because models receive different information, use different tools, and are evaluated under different rules.
By Jonaid Shianifar, Iias Faiud
MetroLLM-Bench is a 955‑case benchmark designed to evaluate language models as the policy layer of transit kiosks across six real metro systems, covering routing, fare calculation, disruptions, accessibility, and adversarial input. The benchmark includes 14 deterministic scoring components (Tier 1) and 8 semantic‑quality components (Tier 2), with a 75/25 split for training‑data generation and held‑out evaluation. Twenty‑six models from six vendors were tested, and a 4B Qwen 3.5 student fine‑tuned via PEFT outperformed GPT‑5.6 on Tier 1 and matched GPT‑5.4 on the combined score, while larger models offered no further improvement.
By Remco Hendriks (Continker)