arXiv AI

Overview of FinMMEval 2026 Task 2: Multilingual Financial Short-Answer Question Answering

arXiv:2607. 19867v1 Announce Type: cross Abstract: FinMMEval 2026 Task 2 evaluates short-answer financial question answering over multilingual evidence.

arXiv AI
Jul 23

Overview of FinMMEval 2026 Task 1: Multilingual Financial Multiple-Choice Question Answering

arXiv:2607. 19856v1 Announce Type: cross Abstract: FinMMEval 2026 Task 1 evaluates multilingual financial multiple-choice question answering in English, Chinese, Arabic, and Hindi.

By Zhuohan Xie, Yuyang Dai, Rania Elbadry, Vanshikaa Jani, Georgi Georgiev, Dimitar Dimitrov, Fan Zhang, Xueqing Peng, Lingfei Qian, Jimin Huang, Jiahui Geng, Yankai Chen, Ye Yuan, Haolun Wu, Yuxia Wang, Ivan Koychev, Veselin Stoyanov, Mingzi Song, Yu Chen, Xue Liu, Preslav Nakov
arXiv AI
Sep 10

IGT @ FinMMEval 2026 Task 2: Question-Type Prompting with Targeted Extraction for Multilingual Financial QA

The IGT system tackles PolyFiQA Task 2 of the FinMMEval Lab, a multilingual financial QA challenge involving English SEC filings and news in five languages. It distinguishes two question families: numeric‑structured queries are answered via keyword extraction from filings, while synthesis queries use rule‑based passage selection from news. The approach yields a development ROUGE‑1 of ~0.395, a 60% boost over a generic RAG baseline, and places third among twelve teams on the official test set.

By Yuwen Chiu (Georgia Institute of Technology)
arXiv AI
Sep 4

Enhancing Financial Question Answering: A Novel Benchmark Dataset of Banks' financial statements

FinRAG-QA is a new benchmark dataset for financial question answering, featuring 999 practitioner-curated questions on 10 standardised indicators drawn from 209 annual and Pillar 3 reports of 24 major European and U.S. banks between 2019 and 2023. The dataset focuses on cross‑institutional retrieval over documents averaging 198k words, making it longer than any existing financial QA resource. Experiments on a multi‑stage Retrieval‑Augmented Generation pipeline show that contextual chunk enrichment and a retrieval‑optimised embedding model significantly improve NDCG@10, while a reasoning‑optimised generator boosts answer accuracy from 44.6% to 79.0% when the correct document is retrieved.

By Arianna Miola, Bruno Spaccavento, Lorenzo Silotto, Marco Bianchetti, Luca Cagliero
arXiv Computation and Language
Aug 27

FinRiskAtlas: Decision-Aligned Evaluation of Large Language Models for Financial Risk Review

FinRiskAtlas is a Chinese-language benchmark designed to evaluate large language models (LLMs) for financial risk review by focusing on decision‑aligned tasks rather than generic financial knowledge. It contains 9,742 instances across 53 task families, including 42 domain‑knowledge families and 11 downstream review operations defined by explicit evaluation contracts. The extended FinRisk‑Ask framework replays 680 pre‑action states from 104 professional trajectories, withholding future evidence during inference to assess evidence‑state control and request targeting. Results across 33 model configurations show that operation‑level evaluation yields distinct rankings and that knowledge‑based shortlisting can incur significant regret, while frequent use of the Ask branch does not necessarily improve evidence acquisition, highlighting gaps in broad financial capability scores.

By Suyang Zhong, Jingzhe Zhu, Qi Xu, Liyao Sun, Yin Wang, Qingqing Sun, Shuai Chen, Tianyi Zhang
arXiv Computation and Language
Aug 31

FinExam-10K: When Retrieval Helps Financial Reasoning?

FinExam-10K is a new English benchmark for financial reasoning, comprising 10,198 expert‑reannotated questions covering CFA Levels I‑III and FRM Parts I‑II. The dataset is split into a 5,110‑question release and a 5,088‑question held‑out set for a quarterly leaderboard, with separate Full‑Coverage and Context‑Complete Reasoning tracks. Across 17 models, the best overall accuracy is 85.29 %, but performance drops on harder subsets, and retrieval‑augmented methods like Function‑Graph‑RAG provide modest gains when gated appropriately.

By Yan Lin, Jingyu Sun, Zhongliang Guo, Qing Li, Zhuohan Xie, Yuxia Wang
arXiv Computation and Language
Sep 23

FinFIRST: Benchmarking Search Agents for Financial Information Retrieval, Sourcing and Traceability

arXiv:2609.25192v1 Announce Type: new Abstract: Financial search is a highly demanding task for LLM agents, requiring not only a correct final answer but also temporally valid information retrieval,...

By Wenqing Wang, Haitao Xiang, Xinyi Zhao, Mingming Yin, Ying Zhong, Zhaoxin Huan, Qiheng Zhou, Jin Zhu, Xiaolu Zhang, Shi Chang, Jun Zhou
arXiv Computation and Language
Aug 31

XHotpotQA: A Benchmark for Cross-Lingual Knowledge Composition in Multi-Hop Question Answering

XHotpotQA is a new benchmark for cross‑lingual knowledge composition in multi‑hop question answering. It presents each instance as an evidence‑dependency graph with explicit language assignments for the question, bridge evidence, answer‑bearing evidence, and distractors, and includes 15,661 training and 7,405 validation examples with sentence‑level support supervision. The dataset reveals significant performance drops when evidence spans language boundaries, providing a diagnostic tool for systems that must integrate evidence across languages.

By Iman Barati, Arash Ghafouri, Behrouz Minaei-Bidgoli
arXiv AI
Sep 17

Version- and Scope-Aware Question Answering over Normative Documents: A Deployed System and an End-to-End Evaluation at Production Scale

The paper presents a deployed system for answering questions over normative documents that is aware of document version and scope. It evaluates a hosted retrieval service against a governed system that applies explicit rules for version and scope resolution, finding the governed system achieves a higher score (97.7 vs 88.1). The study includes a public benchmark, evaluation scripts, and reports commercial deployment metrics, such as 1,126 users and 100,000 calls per day by April 2026.

By Liuyin Wang, Shuaipeng Jin, Jiwei Shi, Jensen Hsu
arXiv Machine Learning
1d ago

Sentence Specificity Scores for Collaborative Technical Documentation: A Domain-Transfer Study

The paper investigates how sentence‑specificity scores can guide the selection of revisions in collaborative technical documentation. It compares two predictors—SpeciTeller and a target‑adapted model by Ko et al.—across Wikipedia and three technical corpora, finding that the predictors rank sentences differently and that SpeciTeller can improve direction‑valid selection rates in certain datasets. The study also shows that filtering and token‑length adjustments alter but do not reconcile these ranking differences.

By Rocker D'Antonio, Thomas Benton Townsend, Dimitrios Michael Manias