arXiv AI

Reasoning Externalization for Faithful Large Language Model Narratives of Stock Return Predictions

The paper introduces a framework that uses large language models (LLMs) to generate natural‑language narratives explaining cross‑sectional stock return predictions. It combines temporal Shapley additive explanations (SHAP) from an XGBoost model with historical regime analogs to provide context. A controlled study shows that progressively externalizing numerical and relational reasoning improves evidence faithfulness and accuracy, while historical analogs boost human‑rated usefulness.

arXiv AI
Sep 10

Evaluating and Improving Evidence-Grounded Fact-Checking in LLMs via Multi-Round Evidence Ablation

The paper introduces Fact-Ablated Evaluation (FAE), a framework that iteratively removes cited evidence to test whether large language models (LLMs) adjust their fact‑checking predictions accordingly. Experiments reveal that many off‑the‑shelf LLMs rely more on internal knowledge than on the provided evidence. To address this, the authors propose REAL, a training method that uses counterfactual evidence supervision to encourage LLMs to base veracity judgments on evidence, achieving better evidence‑dependent performance across four datasets.

By Xingyu Deng, Mingzi Cao, Nikolaos Aletras, Xi Wang, Mark Stevenson
arXiv AI
Jun 4

FinTradeBench: A Financial Reasoning Benchmark for LLMs

arXiv:2603. 19225v3 Announce Type: replace-cross Abstract: Real-world financial decision-making is a challenging problem that requires reasoning over heterogeneous signals, including company fundamentals derived from regulatory filings and trading signals computed from price dynamics.

By Yogesh Agrawal, Aniruddha Dutta, Md Mahadi Hasan, Santu Karmaker, Aritra Dutta
arXiv Computation and Language
4d ago

Can Language Models Learn to Forecast Stock Prices

arXiv:2609.36914v1 Announce Type: new Abstract: Post-training has been shown to significantly improve language models' performance on tasks with verifiable outcomes, including mathematical reasoning,...

By Jiacheng Guo, Suozhi Huang, Shuzhen Li, Yunlong Gao, Zerui Cheng, Jason Ge, Shushu Liang, Zihao Li, Hao Lu, Ming Yin, Shilong Liu, Jiashuo Liu, Xu Kuang, Mengdi Wang
arXiv Machine Learning
Sep 11

AI Economist Agent: An Agentic Framework for Evidence-Based Economic and Financial Analysis with RAG, Knowledge Graphs, and Large Language Models

The paper introduces an AI economist agent that integrates large language models, retrieval‑augmented generation, knowledge graphs, and quantitative models to conduct evidence‑based economic and financial scenario analysis. The framework orchestrates LLM agents to plan analyses, retrieve relevant evidence, and structure economic mechanisms, while registered quantitative models produce numerical outcomes and predefined tests validate intermediate results for inclusion in the final report. Applied to European macro‑financial stress scenarios and bank capital analysis, the empirical study demonstrates the agent’s ability to combine flexible evidence retrieval and scenario construction while maintaining traceability to sources and explicit model calculations.

By Masahiro Kato
arXiv AI
6d ago

The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?

The study investigates whether adding inference-time reasoning to large language models (LLMs) improves trading performance. Using a controlled experiment across DeepSeek, GPT, and Gemini models, the authors varied reasoning effort while keeping other variables constant and evaluated over a full year of U.S. equities under three input conditions. Results show that additional reasoning does not reliably increase net portfolio returns and can even lead to nonmonotonic performance and unstable outcomes.

By Jiayi Chen, Guiling Wang