arXiv Machine Learning

FinBench: Time-Gated Calibration and Uncertainty Benchmarking for Agentic Financial Forecasting

arXiv:2607. 16229v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as components of agentic systems that observe, plan, and act.

arXiv Computation and Language
4d ago

Can Language Models Learn to Forecast Stock Prices

arXiv:2609.36914v1 Announce Type: new Abstract: Post-training has been shown to significantly improve language models' performance on tasks with verifiable outcomes, including mathematical reasoning,...

By Jiacheng Guo, Suozhi Huang, Shuzhen Li, Yunlong Gao, Zerui Cheng, Jason Ge, Shushu Liang, Zihao Li, Hao Lu, Ming Yin, Shilong Liu, Jiashuo Liu, Xu Kuang, Mengdi Wang
arXiv AI
Aug 5

FinVerse: Financial Time-Series Benchmark

arXiv:2608. 03259v1 Announce Type: cross Abstract: As time-series foundation models have emerged, the need for benchmarks that can evaluate their forecasting ability in meaningful ways has become increasingly important.

By Jaehoon Lee, Jun Seo, Seunghan Lee, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, Minjae Kim, Sungdong Yoo, Junhyeok Kang, Sangjun Han, Soonyoung Lee, Wonbin Ahn
arXiv Machine Learning
Jun 2

FinTSB: A Comprehensive and Practical Benchmark for Financial Time Series Forecasting

arXiv:2502. 18834v3 Announce Type: replace-cross Abstract: Financial time series (FinTS) record the behavior of human-brain-augmented decision-making, capturing valuable historical information that can be leveraged for profitable investment strategies.

By Yifan Hu, Yuante Li, Peiyuan Liu, Yuxia Zhu, Naiqi Li, Tao Dai, Shu-tao Xia, Dawei Cheng, Changjun Jiang
arXiv Computation and Language
Aug 27

FinRiskAtlas: Decision-Aligned Evaluation of Large Language Models for Financial Risk Review

FinRiskAtlas is a Chinese-language benchmark designed to evaluate large language models (LLMs) for financial risk review by focusing on decision‑aligned tasks rather than generic financial knowledge. It contains 9,742 instances across 53 task families, including 42 domain‑knowledge families and 11 downstream review operations defined by explicit evaluation contracts. The extended FinRisk‑Ask framework replays 680 pre‑action states from 104 professional trajectories, withholding future evidence during inference to assess evidence‑state control and request targeting. Results across 33 model configurations show that operation‑level evaluation yields distinct rankings and that knowledge‑based shortlisting can incur significant regret, while frequent use of the Ask branch does not necessarily improve evidence acquisition, highlighting gaps in broad financial capability scores.

By Suyang Zhong, Jingzhe Zhu, Qi Xu, Liyao Sun, Yin Wang, Qingqing Sun, Shuai Chen, Tianyi Zhang
arXiv AI
Sep 25

When Should Forecasting Agents Reason? Behavioral Stress Tests for Reliability Routing

The paper investigates when forecasting agents should employ different behaviors—retrieval, reasoning, deferring to market priors, or using historical analogs—on binary forecasting tasks. It finds that the optimal mechanism depends on the data source, with structured analogs excelling for some processes and market or conservative baselines for others. The authors propose ReliabilityRoute, a rule‑based system that steers agent behavior using reliability features, achieving competitive performance across multiple LLM versions while highlighting that more reasoning is not always better.

By Yufeng Wang