arXiv Machine Learning

Enhancing Regime Shift Detection Using Unstructured Data: A Study on the Treasury Market

arXiv:2605. 30363v2 Announce Type: replace-cross Abstract: Regime shifts in financial markets reorganise the joint dynamics of asset prices and macro variables, breaking any single-regime calibration.

arXiv Machine Learning
Aug 19

Temporal Leakage in Financial News NLP: A Multi-Architecture Audit with a Regime-Specific M&A Signal

The paper audits the impact of temporal leakage on financial-news direction prediction across 49,799 articles and 16 feature-model combinations, including TF‑IDF, MiniLM, FinBERT, and fine‑tuned RoBERTa‑large / DeBERTa‑v3‑large, as well as zero/few‑shot and LoRA probes of Llama‑3 and Qwen2.5. Random train‑test splits inflate MCC scores by 1.1× to 6.5×, with larger models and richer features showing greater gains, while end‑to‑end FinBERT fine‑tuning actually increases the gap. Only the mergers and acquisitions (M&A) category shows a positive locked‑test signal under near‑temporal chronological evaluation, with the signal localized to 2024‑2025 European‑tilted M&A semantics and not transferring to a 2009‑2020 U.S. corpus.

By Chenhao Xue, Raslen Guesmi, Siwei Feng, Yucheng Gong, Jacob Xavier Sundram, Jordan Pang, Lan Wang, Julian Kaljuvee
arXiv AI
Sep 1

Can LLMs Take the Pulse of the Economy? A Real-Time Evaluation of LLM Nowcasts on Macroeconomic Indicators

The paper introduces LiveMacroEval, a live benchmark that tests large language model (LLM) agents’ ability to produce hourly nowcasts for sixteen major U.S. macroeconomic indicators before their official release. It evaluates LLM performance against institutional nowcasts, Bloomberg ECOS consensus, and an auto-ARIMA baseline using a LiveMacro Score linked to announcement-window equity returns and a LiveBetting Score from simulated Polymarket-style trading. Over six months, state-of-the-art LLMs with web search achieved overall accuracy comparable to professional benchmarks, though performance varied across indicators.

By Xinyue Zhao, Ruiyi Zhang, Liqin Ye, Rui Cao, Pengtao Xie, Sudheer Chava
arXiv AI
Sep 23

TimeLitmus: A Diagnostic Benchmark for Cross-Modal Understanding and Explanation Faithfulness in Event-Conditioned Time-Series Prediction

arXiv:2609.24677v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to make predictions from numerical time-series histories and textual events. Yet accuracy alone cann...

By Jie Gong, Maowei Jiang, Zhiwei Liu, Yankai Chen, Guojun Xiong, Xue Liu, Min Peng, Qianqian Xie, Sophia Ananiadou
arXiv AI
3d ago

DualCast: A Dual-Path Language Model for Bimodal Financial Time-Series Forecasting

DualCast is a dual‑path language model that forecasts financial time‑series by combining a fast numerical forecaster with an optional text‑conditioned revision mechanism. The fast path trains only new financial‑token embeddings and output heads on a frozen Qwen3‑8B backbone, while the slow path uses a LoRA adapter to incorporate news and refine predictions. In zero‑shot tests across equities and energy prices at multiple time resolutions, the slow path achieves the lowest mean absolute percentage error in most settings, especially for longer horizons, and news ablations show additional gains in many markets.

By Wentao Zhao, Hongqiang Wu, Shanghang Liu, Zhaochen Zan, Yu Zhang, Biqing Huang
arXiv Computation and Language
6d ago

PALM: Point-in-Time Adaptation for Financial Language Models

arXiv:2609.30316v1 Announce Type: cross Abstract: Language models used in financial backtests suffer from look-ahead bias, as a model trained on text published after the study period has already obse...

By Seunghan Lee, Jun Seo, Jaehoon Lee, Junhyeok Kang, Sangjun Han, Sungdong Yoo, Minjae Kim, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, Soonyoung Lee, Wonbin Ahn
arXiv AI
Jul 15

Scaling Point-in-Time Language Models

arXiv:2607. 11889v1 Announce Type: cross Abstract: Large language models trained on unrestricted internet corpora inevitably embed information from the future, introducing lookahead bias that compromises the validity of backtests and causal inference in finance and the social sciences.

By Bryan Kelly, Semyon Malamud, Johannes Schwab, Teng Andrea Xu