arXiv:2608. 11047v1 Announce Type: new Abstract: While existing benchmarks have made substantial progress in evaluating LLMs across STEM domains, financial reasoning over structured data remains comparatively less explored.
By Alicia Larsen, Victoire Laurent, Aulia Kharis Rakhamsari, Lara Turgut, Nino Antulov-Fantulin
The paper introduces a data‑centric pipeline for post‑training language models on financial reasoning tasks. It mines open‑source reasoning traces, distills financial instruction data, and generates knowledge‑graph‑guided question‑answer pairs, then filters examples with lightweight classifiers and applies reinforcement learning with rule‑based verifiers. Experiments on FINESSE‑Bench show that retention‑aware adaptation—self‑distilled fine‑tuning and model merging—outperforms ordinary supervised fine‑tuning, improving accuracy by up to 3.0 points and avoiding regressions.
By Zhirayr Hayrapetyan, Andrei Kalmykov, Denis Kokosinskii, Dmitry Stanishevskii, Dmitry Zmitrovich
arXiv:2605. 05409v2 Announce Type: replace Abstract: Financial document question answering (QA) demands complex multi-step numerical reasoning over heterogeneous evidence--structured tables, textual narratives, and footnotes--scattered across corporate filings.
By Yang Shu, Yingmin Liu, Zequn Xie
ASDA (Automated Skill Distillation and Adaptation) is a framework that improves large language models on financial reasoning tasks without fine‑tuning. It works by having a teacher model analyze a student’s failures, cluster errors, and generate structured skill artifacts—reasoning procedures, code templates, and worked examples—that are injected during inference. On the FAMMA benchmark, ASDA boosts arithmetic reasoning by up to 17.33% and non‑arithmetic reasoning by 5.95%, outperforming existing training‑free methods.
By Tik Yu Yim, Wenting Tan, Sum Yee Chan, Tak-Wah Lam, Siu Ming Yiu
arXiv:2602. 08324v5 Announce Type: replace Abstract: Chain-of-Thought (CoT) reasoning successfully enhances the reasoning capabilities of Large Language Models (LLMs), yet it incurs substantial computational overhead for inference.
By Yuntian Tang, Bohan Jia, Wenxuan Huang, Lianyue Zhang, Jiao Xie, Wenxi Li, Wei Li, Jie Hu, Xinghao Chen Rongrong Ji, Shaohui Lin
arXiv:2604. 10015v3 Announce Type: replace Abstract: Recent studies demonstrate that tool-calling capability enables large language models (LLMs) to interact with external environments for long-horizon financial tasks.
By Yupeng Cao, Haohang Li, Weijin Liu, Wenbo Cao, Anke Xu, Lingfei Qian, Xueqing Peng, Minxue Tang, Zhiyuan Yao, Jimin Huang, K. P. Subbalakshmi, Zining Zhu, Jordan W. Suchow, Yangyang Yu