The paper introduces an agentic forecasting environment built on 2,100+ resolved Polymarket questions, where a language model (Qwen3.5-35B-A3B) learns to gather evidence during rollout via web search, page reading, and financial time series, all filtered to avoid post‑cutoff leaks. Training with single‑epoch GRPO and a Brier‑score reward improves calibration by 30‑40% and reduces search attempts, while the trained policy outperforms four frontier models in evidence‑based forecasting, achieving lower soft‑Brier scores at roughly 5% of the inference cost. The authors release the environment, dataset, and per‑rollout records as a reusable harness for temporal forecasting agents.
By Yusuf Afifi, Artur Kiulian, Anton Polishko, Mykola Khandoga, Hamudi Naanaa, Alina Krasnobrizha
arXiv:2606. 02497v1 Announce Type: new Abstract: Time series forecasting has advanced rapidly, especially with the emergence of foundation models that show strong zero-shot performance on numerical extrapolation.
By Yuhua Liao, Zetian Wang, Qiangqiang Nie, Zhenhua Zhang
The paper investigates how different proper scoring rules influence the performance and behavior of large language model (LLM) forecasters. Five scoring rules were compared as training objectives for binary forecasts of real-world events, revealing that while they all theoretically incentivize truthful probability reporting, they produce models with varying calibration, probability usage, and bias, information, and noise profiles. The Brier-trained model achieved the lowest Brier score and highest AUC-ROC, whereas the log-trained model achieved the best log score and lowest calibration error, indicating that scoring rule choice can shape both forecast accuracy and error structure.
By Benjamin Turtel, Paul Wilczewski, Kris Skotheim, Ville A. Satop\"a\"a, Philip E. Tetlock
arXiv:2608.23058v1 Announce Type: new
Abstract: Large language models (LLMs) now support forecasting systems that combine language-based reasoning with temporal data, evidence retrieval, external too...
By Xiaogang Xu, Jiaqi Tang, Jianmin Chen, Yingying Yan, Zhenchao Tang, Xiangxin Zhou, Xiaobin Hu, Wei Wei, Jinfeng Wu, Qifeng Chen, Lu Zhou, Jiafei Wu, Zhe Liu, Jianwei Yin, Weimin Zheng
Text-conditioned time-series forecasting predicts a series from both its numerical history and natural-language context, allowing forecasts to account for events and constraints that the past alone cannot reveal. This requires both reliable numerical forecasting and the ability to interpret contextual information.
arXiv:2607. 24892v1 Announce Type: cross Abstract: Text-conditioned time-series forecasting predicts a series from both its numerical history and natural-language context, allowing forecasts to account for events and constraints that the past alone cannot reveal.
By Huu Hiep Nguyen, Dung Nguyen, Minh Hoang Nguyen, Dai Do, Hung Le
Forecast-Dojo is a replayable environment designed to benchmark and train large language model (LLM) forecasting agents. It integrates resolved prediction‑market questions with dated news, enabling agents to research events and revisit predictions at successive historical dates. The platform includes 1,568 Polymarket events, 18.8 million dated news articles, and supports repeated evaluation, training interactions, and outcome feedback, with evidence that research tools lower Brier scores across 12 tested models, though all models still lag behind historical market forecasts.
By Liqin Ye, Haorui Wang, Fardin Ahmed, Rongzhi Zhang, Yuan He, Ziyuan Lin, Yanbin Yin, Jing Peng, Michael Galarnyk, Sudheer Chava, Chao Zhang
LEAP (Likelihood Elicitation and Aggregation for Probabilistic forecasting) is a new approach that reorganizes how evidence is used in LLM-based forecasting systems. Instead of a monolithic prediction that aggregates all evidence at once, LEAP examines each evidence item separately, elicits likelihood parameters, and combines them with an explicit prior to produce a posterior distribution. The method supports continuous, single-choice, and multi-choice forecasts and has been shown to improve prediction and calibration metrics across models on a benchmark covering forecasting, information-seeking, and browsing tasks.
By Yufei Chen, Yiran Zhao, Xiaogang Xu, Qipeng Xie, Jiafei Wu, Zhe Liu
arXiv:2609.05905v1 Announce Type: cross
Abstract: LLM agents are increasingly used for live forecasting, where they retrieve up-to-date information and produce estimates for unresolved future events....
By Yuanpu Cao, Yongkang Du, Yurui Chang, Lu Lin, Jinghui Chen
arXiv:2609.36689v1 Announce Type: new
Abstract: Large language models have achieved significant progress in event forecasting, yet their probability outputs exhibit systematic calibration bias that v...
By Wenjin Liu, Chenxi Wang, Yue Lu, Zhe Cui, Haoran Luo
BridgeMem is a new method for temporal knowledge graph forecasting that focuses on pair‑specific transition evidence, adding a residual correction to the log scores of a frozen full‑vocabulary forecaster. It retrieves and encodes prior events between a query actor and candidate, converting them into a likelihood‑ratio correction via a support‑adaptive empirical‑Bayes reader. Across five benchmarks, BridgeMem outperforms nine baselines from 2021–2026, improving filtered MRR and Hits@{1,3,10} metrics by up to 0.0216.
By Zeyan Li, Libing Chen, Shengda Zhuo, Yin Tang, Jianfeng Xu
arXiv:2609.36914v1 Announce Type: new
Abstract: Post-training has been shown to significantly improve language models' performance on tasks with verifiable outcomes, including mathematical reasoning,...
By Jiacheng Guo, Suozhi Huang, Shuzhen Li, Yunlong Gao, Zerui Cheng, Jason Ge, Shushu Liang, Zihao Li, Hao Lu, Ming Yin, Shilong Liu, Jiashuo Liu, Xu Kuang, Mengdi Wang