The paper introduces LiveMacroEval, a live benchmark that tests large language model (LLM) agents’ ability to produce hourly nowcasts for sixteen major U.S. macroeconomic indicators before their official release. It evaluates LLM performance against institutional nowcasts, Bloomberg ECOS consensus, and an auto-ARIMA baseline using a LiveMacro Score linked to announcement-window equity returns and a LiveBetting Score from simulated Polymarket-style trading. Over six months, state-of-the-art LLMs with web search achieved overall accuracy comparable to professional benchmarks, though performance varied across indicators.
By Xinyue Zhao, Ruiyi Zhang, Liqin Ye, Rui Cao, Pengtao Xie, Sudheer Chava
arXiv:2512. 23847v2 Announce Type: replace-cross Abstract: We develop a statistical procedure to detect lookahead bias in economic forecasts generated by large language models (LLMs).
By Zhenyu Gao, Wenxi Jiang, Yutong Yan
arXiv:2606. 24950v1 Announce Type: new Abstract: Financial decision-making is contextual: forecasting prices, valuing companies, and assessing event exposure weigh price history, accounting fundamentals, macroeconomic regime, and contemporaneous text.
By Patara Trirat, Jin Myung Kwak, Jay Heo, Heejun Lee, Sung Ju Hwang
arXiv:2607. 23146v1 Announce Type: new Abstract: Inspired by recent breakthroughs in large language models for natural language processing, foundation models have emerged as a promising paradigm for zero-shot time series forecasting, enabling accurate predictions on datasets never seen during pre-training.
By Morad Laglil, Bertrand Pracca, Emilie Devijver, Eric Gaussier
arXiv:2508. 09904v3 Announce Type: replace-cross Abstract: Real-world forecasting requires models to integrate not only historical data but also relevant contextual information provided in textual form.
By Arjun Ashok, Andrew Robert Williams, Vincent Zhihao Zheng, Irina Rish, Nicolas Chapados, \'Etienne Marcotte, Valentina Zantedeschi, Alexandre Drouin
arXiv:2608.23058v1 Announce Type: new
Abstract: Large language models (LLMs) now support forecasting systems that combine language-based reasoning with temporal data, evidence retrieval, external too...
By Xiaogang Xu, Jiaqi Tang, Jianmin Chen, Yingying Yan, Zhenchao Tang, Xiangxin Zhou, Xiaobin Hu, Wei Wei, Jinfeng Wu, Qifeng Chen, Lu Zhou, Jiafei Wu, Zhe Liu, Jianwei Yin, Weimin Zheng
arXiv:2602. 02288v3 Announce Type: replace Abstract: Current time-series forecasting models are primarily based on transformer-style neural networks.
By Zheng Li, Jerry Cheng, Huanying Gu
arXiv:2606. 28670v1 Announce Type: cross Abstract: We introduce MACROCAST, a lightweight Time Series Foundation Model (TSFM) for real-time macroeconomic forecasting.
By Andrea Carriero, Davide Pettenuzzo, Shubhranshu Shekhar
arXiv:2608. 01599v1 Announce Type: new Abstract: Volatility forecasts are commonly evaluated with aggregate accuracy metrics such as RMSE and MAE, but these metrics can hide conditional failures that matter for risk management.
By Arthur Chagas, Pedro Bento, Yan Aquino, Arthur Buzelin, Wagner Meira Jr., Cristiano Arbex Valle
While deep learning models, particularly transformer-based architectures, have shown impressive performance in time series forecasting, the application of retrieval-augmented generation (RAG) in this domain remains limited. Since RAG has proven effective in enhancing the capabilities of large language models by incorporating relevant external information, retrieving similar time series sequences as references might also improve accuracy in time series forecasting tasks.
arXiv:2608. 03259v1 Announce Type: cross Abstract: As time-series foundation models have emerged, the need for benchmarks that can evaluate their forecasting ability in meaningful ways has become increasingly important.
By Jaehoon Lee, Jun Seo, Seunghan Lee, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, Minjae Kim, Sungdong Yoo, Junhyeok Kang, Sangjun Han, Soonyoung Lee, Wonbin Ahn
The paper investigates the coherence of probabilistic forecasts produced by language models, particularly in the context of life‑decision support. Using a de Finetti‑based method, the authors elicit forecasts for events derived from stock return data and compute the maximum Dutch‑book profit via linear programming, which quantifies incoherence. They find significant incoherence, especially when events have complex logical relationships or when irrelevant context is present, and suggest that alternative training strategies could improve coherence.
By Isaiah Andrews, Suproteem Sarkar