arXiv Machine Learning

Detecting Lookahead Bias in LLM Forecasts

arXiv:2512. 23847v2 Announce Type: replace-cross Abstract: We develop a statistical procedure to detect lookahead bias in economic forecasts generated by large language models (LLMs).

arXiv Machine Learning
Sep 25

A Fast and Effective Solution to the Problem of Look-ahead Bias in LLMs

The paper addresses look‑ahead bias in large language models (LLMs) used for financial prediction, which arises because LLMs are trained on long time‑series data. It proposes a low‑cost solution that adjusts the logits of a base model at inference time using two smaller, specialized models—one fine‑tuned to forget certain information and another to retain it. Experiments show that this method removes both verbatim and semantic knowledge, corrects biases, and outperforms previous approaches.

By Humzah Merchant, Bradford Levy
arXiv AI
4d ago

Alignment Forecasting: Predicting Misalignment From Training Data

The paper introduces Alignment Forecasting, a method for predicting whether fine‑tuning a language model on a given dataset will increase specific alignment failures such as deception or sycophancy. It presents ALIGNMENTFORECASTBENCH, a benchmark of over 5,000 forecasting questions across many models, datasets, and failure modes, and shows that a simple forecasting scaffold using an LLM’s assessment of dataset bias can outperform baseline forecasters. The authors demonstrate that filtering out high‑risk training examples identified by the forecaster can improve alignment in multiple‑choice evaluations, though benefits in open‑ended conversations remain uncertain.

By Chen Yueh-Han, Bruce W. Lee, Ilia Sucholutsky, Tomek Korbak
arXiv AI
Aug 25

LLM-based Agents for Forecasting and Prediction: Methods, Training, Evaluation, and Applications

arXiv:2608.23058v1 Announce Type: new Abstract: Large language models (LLMs) now support forecasting systems that combine language-based reasoning with temporal data, evidence retrieval, external too...

By Xiaogang Xu, Jiaqi Tang, Jianmin Chen, Yingying Yan, Zhenchao Tang, Xiangxin Zhou, Xiaobin Hu, Wei Wei, Jinfeng Wu, Qifeng Chen, Lu Zhou, Jiafei Wu, Zhe Liu, Jianwei Yin, Weimin Zheng
arXiv Machine Learning
Aug 18

Macroeconomic Forecasting with Large Language Models

arXiv:2407. 00890v5 Announce Type: replace-cross Abstract: This paper presents a comparative analysis evaluating the accuracy of Large Language Models (LLMs) against traditional macro time series forecasting approaches.

By Andrea Carriero, Davide Pettenuzzo, Shubhranshu Shekhar
arXiv Machine Learning
Aug 4

Latent-Regime Bias Auditing for Volatility Forecasting

arXiv:2608. 01599v1 Announce Type: new Abstract: Volatility forecasts are commonly evaluated with aggregate accuracy metrics such as RMSE and MAE, but these metrics can hide conditional failures that matter for risk management.

By Arthur Chagas, Pedro Bento, Yan Aquino, Arthur Buzelin, Wagner Meira Jr., Cristiano Arbex Valle
arXiv Machine Learning
Sep 7

How Faithful Is Attribution for Sales Forecasting? A Counterfactual Study

The paper introduces a counterfactual interpretability layer for multi‑series WaveNet forecasters used in sales forecasting. It decomposes each forecast into contributions that exactly sum to the predicted value, avoiding allocation artifacts seen with SHAP‑style methods. The authors evaluate faithfulness via deletion/insertion tests, showing significant gaps that confirm the attributions reflect genuine model behavior, and they analyze where attribution is informative and where it is not.

By Glib Kechyn
arXiv AI
Sep 16

Repurposing Deep Limit Order Book Forecasting for Scenario-Conditioned Market Impact Modeling

The paper demonstrates that deep limit order book forecasting models can be repurposed to quantify scenario-conditioned market impact without retraining. By injecting counterfactual order‑book messages into a trained Transformer forecaster, the authors compare predictive distributions before and after the injection, defining a short‑horizon model‑implied market impact. The approach achieves a Spearman correlation of 0.99 and 97.2% directional agreement with historical outcomes for non‑neutral scenarios, and captures incremental sequence‑dependent variation beyond scenario identity and pre‑event forecasts.

By Eljas Linna, Kestutis Baltakys, Derrick Manoharan, Alexandros Iosifidis, Juho Kanniainen
arXiv AI
Sep 25

Forecast-Dojo: Replayable Environments for Benchmarking and Training LLM Forecasting Agents

Forecast-Dojo is a replayable environment designed to benchmark and train large language model (LLM) forecasting agents. It integrates resolved prediction‑market questions with dated news, enabling agents to research events and revisit predictions at successive historical dates. The platform includes 1,568 Polymarket events, 18.8 million dated news articles, and supports repeated evaluation, training interactions, and outcome feedback, with evidence that research tools lower Brier scores across 12 tested models, though all models still lag behind historical market forecasts.

By Liqin Ye, Haorui Wang, Fardin Ahmed, Rongzhi Zhang, Yuan He, Ziyuan Lin, Yanbin Yin, Jing Peng, Michael Galarnyk, Sudheer Chava, Chao Zhang
arXiv Computation and Language
4d ago

Can Language Models Learn to Forecast Stock Prices

arXiv:2609.36914v1 Announce Type: new Abstract: Post-training has been shown to significantly improve language models' performance on tasks with verifiable outcomes, including mathematical reasoning,...

By Jiacheng Guo, Suozhi Huang, Shuzhen Li, Yunlong Gao, Zerui Cheng, Jason Ge, Shushu Liang, Zihao Li, Hao Lu, Ming Yin, Shilong Liu, Jiashuo Liu, Xu Kuang, Mengdi Wang