BizFinBench.v2: Towards Reliable LLMs in Finance via Real-User Data and Offline/Online Bilingual Evaluation
arXiv:2601. 06401v2 Announce Type: replace Abstract: Large language models are becoming increasingly significant in financial applications.
arXiv:2608. 08634v1 Announce Type: new Abstract: Open-weight language models from Chinese AI labs caught up on benchmarks relative to proprietary frontier models in recent months.
arXiv:2601. 06401v2 Announce Type: replace Abstract: Large language models are becoming increasingly significant in financial applications.
arXiv:2608. 12342v1 Announce Type: cross Abstract: Ensuring the accuracy of financial documents is critical for economic analysis, regulatory compliance, and corporate decision-making.
arXiv:2609.24002v1 Announce Type: new Abstract: Large language model agents increasingly answer financial questions by searching regulatory filings. Such questions are often deceptively under-specifi...
arXiv:2608. 06108v1 Announce Type: new Abstract: Investment competence is inherently personalized: the same market evidence can justify different actions for investors with different goals, horizons, portfolios, and risk boundaries.
arXiv:2608. 04374v1 Announce Type: cross Abstract: Large language models can produce fluent financial analysis, but fluency alone does not establish whether a report is suitable for institutional delivery.
arXiv:2607. 22841v1 Announce Type: cross Abstract: We present DS@GT's submission to FinMMEval 2026 Task 1, a multilingual financial exam question answering benchmark spanning English, Spanish, Greek, Chinese, and Hindi.
arXiv:2606. 24950v1 Announce Type: new Abstract: Financial decision-making is contextual: forecasting prices, valuing companies, and assessing event exposure weigh price history, accounting fundamentals, macroeconomic regime, and contemporaneous text.
FinRAG-QA is a new benchmark dataset for financial question answering, featuring 999 practitioner-curated questions on 10 standardised indicators drawn from 209 annual and Pillar 3 reports of 24 major European and U.S. banks between 2019 and 2023. The dataset focuses on cross‑institutional retrieval over documents averaging 198k words, making it longer than any existing financial QA resource. Experiments on a multi‑stage Retrieval‑Augmented Generation pipeline show that contextual chunk enrichment and a retrieval‑optimised embedding model significantly improve NDCG@10, while a reasoning‑optimised generator boosts answer accuracy from 44.6% to 79.0% when the correct document is retrieved.
arXiv:2608. 11683v1 Announce Type: new Abstract: AI agents are increasingly deployed for professional investment research, yet no benchmark captures the complexity of the full investor workflow.
arXiv:2608. 16386v1 Announce Type: cross Abstract: Financial agents must do more than recall domain knowledge: they must be both reliable, executing precise operations over grounded evidence, and executive, sustaining long-horizon research whose conclusions remain auditable.
arXiv:2604. 10015v3 Announce Type: replace Abstract: Recent studies demonstrate that tool-calling capability enables large language models (LLMs) to interact with external environments for long-horizon financial tasks.
CIFQA is a deterministic, tool‑grounded multi‑agent framework that separates language understanding from numerical execution for financial question answering. It assigns specialized agents for interpretation, routing, parameter extraction, computation planning, and response generation, while deterministic Python tools perform the calculations. On a fixed‑deposit benchmark, CIFQA achieves 95.54% accuracy on calculation‑intensive queries and 90.87% overall, outperforming larger LLM baselines and showing that architecture, not scale, drives numerical reliability.