arXiv AI

V-FiLLM: Verified Financial LLM Reasoning Benchmark

arXiv:2608. 11047v1 Announce Type: new Abstract: While existing benchmarks have made substantial progress in evaluating LLMs across STEM domains, financial reasoning over structured data remains comparatively less explored.

arXiv AI
Sep 10

ASDA: Automated Skill Distillation and Adaptation for Financial Reasoning

ASDA (Automated Skill Distillation and Adaptation) is a framework that improves large language models on financial reasoning tasks without fine‑tuning. It works by having a teacher model analyze a student’s failures, cluster errors, and generate structured skill artifacts—reasoning procedures, code templates, and worked examples—that are injected during inference. On the FAMMA benchmark, ASDA boosts arithmetic reasoning by up to 17.33% and non‑arithmetic reasoning by 5.95%, outperforming existing training‑free methods.

By Tik Yu Yim, Wenting Tan, Sum Yee Chan, Tak-Wah Lam, Siu Ming Yiu
arXiv Computation and Language
Sep 10

Data-Centric Post-Training for Financial Reasoning: Mining, Distillation, and Verifiable Learning

The paper introduces a data‑centric pipeline for post‑training language models on financial reasoning tasks. It mines open‑source reasoning traces, distills financial instruction data, and generates knowledge‑graph‑guided question‑answer pairs, then filters examples with lightweight classifiers and applies reinforcement learning with rule‑based verifiers. Experiments on FINESSE‑Bench show that retention‑aware adaptation—self‑distilled fine‑tuning and model merging—outperforms ordinary supervised fine‑tuning, improving accuracy by up to 3.0 points and avoiding regressions.

By Zhirayr Hayrapetyan, Andrei Kalmykov, Denis Kokosinskii, Dmitry Stanishevskii, Dmitry Zmitrovich
arXiv AI
Aug 11

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks

arXiv:2604. 10015v3 Announce Type: replace Abstract: Recent studies demonstrate that tool-calling capability enables large language models (LLMs) to interact with external environments for long-horizon financial tasks.

By Yupeng Cao, Haohang Li, Weijin Liu, Wenbo Cao, Anke Xu, Lingfei Qian, Xueqing Peng, Minxue Tang, Zhiyuan Yao, Jimin Huang, K. P. Subbalakshmi, Zining Zhu, Jordan W. Suchow, Yangyang Yu
arXiv AI
Sep 2

Dependency-Aware Chain-of-Thought Compression for Financial Reasoning

The paper introduces the Hierarchical Semantic Distillation Network (HSDN), a method for compressing chain-of-thought reasoning in financial contexts while maintaining accuracy and logical coherence. HSDN uses semantic segmentation, dependency graph construction, dual encoder importance scoring, constrained segment selection, and local boundary rewriting to reduce the length of intermediate reasoning traces. Evaluated on the AFAC2025 benchmark, HSDN achieves 91.0% accuracy with a 68.4% compression rate, outperforming strong baselines in overall score and reasoning coherence.

By Wenjun Wu, Lei Fu, Kejian Tong, Tao Ning, Sichen Zhao
arXiv AI
Jul 28

Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers

arXiv:2509. 03059v2 Announce Type: replace-cross Abstract: Recent advances in Large Language Models (LLMs) have shown that their reasoning capabilities can be significantly improved through Reinforcement Learning with Verifiable Reward (RLVR), particularly in domains like mathematics and programming, where ground-truth correctness can be automatically evaluated.

By Xingyue Huang, Rishabh, Gregor Franke, Ziyi Yang, Jiamu Bai, Weijie Bai, Jinhe Bi, Zifeng Ding, Yiqun Duan, Chengyu Fan, Wendong Fan, Xin Gao, Ruohao Guo, Yuan He, Zhuangzhuang He, Xianglong Hu, Neil Johnson, Bowen Li, Fangru Lin, Siyu Lin, Tong Liu, Yunpu Ma, Hao Shen, Hao Sun, Beibei Wang, Fangyijie Wang, Hao Wang, Haoran Wang, Yang Wang, Yifeng Wang, Zhaowei Wang, Ziyang Wang, Yifan Wu, Zikai Xiao, Chengxing Xie, Fan Yang, Junxiao Yang, Qianshuo Ye, Ziyu Ye, Guangtao Zeng, Yuwen Ebony Zhang, Zeyu Zhang, Zihao Zhu, Bernard Ghanem, Philip Torr, Guohao Li
arXiv AI
Jun 4

FinTradeBench: A Financial Reasoning Benchmark for LLMs

arXiv:2603. 19225v3 Announce Type: replace-cross Abstract: Real-world financial decision-making is a challenging problem that requires reasoning over heterogeneous signals, including company fundamentals derived from regulatory filings and trading signals computed from price dynamics.

By Yogesh Agrawal, Aniruddha Dutta, Md Mahadi Hasan, Santu Karmaker, Aritra Dutta
arXiv AI
Aug 28

CIFQA: A Deterministic Tool-Grounded Multi-Agent LLM Framework for Financial Query Answering

CIFQA is a deterministic, tool‑grounded multi‑agent framework that separates language understanding from numerical execution for financial question answering. It assigns specialized agents for interpretation, routing, parameter extraction, computation planning, and response generation, while deterministic Python tools perform the calculations. On a fixed‑deposit benchmark, CIFQA achieves 95.54% accuracy on calculation‑intensive queries and 90.87% overall, outperforming larger LLM baselines and showing that architecture, not scale, drives numerical reliability.

By Kunjesh Parekh, Anil Kumar Tiwari, Divya Saxena
arXiv AI
Sep 25

LiveMathematicianBench: A Live Benchmark for Research-Level Mathematical Reasoning with Proof Sketches

LiveMathematicianBench is a dynamic multiple‑choice benchmark for research‑level mathematical reasoning, built from recent arXiv papers published after model training cutoffs. It introduces a thirteen‑category logical taxonomy of theorem types and uses a proof‑sketch‑guided distractor pipeline to create plausible but invalid answer choices, enhancing sensitivity to genuine reasoning. Evaluation shows current large language models perform poorly, with the best model scoring 43.5% overall and only 17.6% under substitution‑resistant conditions, indicating the benchmark’s difficulty and realism.

By Linyang He, Qiyao Yu, Hanze Dong, Baohao Liao, Xinxing Xu, Micah Goldblum, Jiang Bian, Nima Mesgarani