arXiv:2602. 07294v4 Announce Type: replace-cross Abstract: With the increasing deployment of Large Language Models (LLMs) in the finance domain, LLMs are increasingly expected to parse complex regulatory disclosures.
By Yidong Jiang, Junrong Chen, Eftychia Makri, Jialin Chen, Peiwen Li, Ali Maatouk, Leandros Tassiulas, Eliot Brenner, Bing Xiang, Rex Ying
The paper introduces a benchmark for evaluating large language models on trustworthy analysis of earnings call transcripts. It proposes a numeric evidence evaluation method that assesses groundedness without expert annotation, and presents an automated pipeline that builds the ECTs-100 dataset from the top 100 S&P 500 constituents. The study also explores the failure mode of conscious incompetence, where models must recognize insufficient evidence and avoid hallucinations, finding that while groundedness is strong, correctness remains a challenge.
By Yingzhu Zhao, Vlad Pandelea, Han Yuan, Bo Hu, Wuqiong Luo, Li Zhang, Zheng Ma
CITECHOICE is a causal audit that examines how the presentation of documents in an agentic search engine redistributes citation credit. Using 129 everyday‑query transcripts, the study compares structured versus prose renderings of the same source while keeping all other transcript elements fixed. The results show that structured rendering increases the target’s citation count by about half a citation per answer without adding total citations or diminishing competitors’ credit, while also revealing that rank position has a larger effect on citation rates than presentation order alone.
By Sriram Selvam, Anneswa Ghosh
FinRiskAtlas is a Chinese-language benchmark designed to evaluate large language models (LLMs) for financial risk review by focusing on decision‑aligned tasks rather than generic financial knowledge. It contains 9,742 instances across 53 task families, including 42 domain‑knowledge families and 11 downstream review operations defined by explicit evaluation contracts. The extended FinRisk‑Ask framework replays 680 pre‑action states from 104 professional trajectories, withholding future evidence during inference to assess evidence‑state control and request targeting. Results across 33 model configurations show that operation‑level evaluation yields distinct rankings and that knowledge‑based shortlisting can incur significant regret, while frequent use of the Ask branch does not necessarily improve evidence acquisition, highlighting gaps in broad financial capability scores.
By Suyang Zhong, Jingzhe Zhu, Qi Xu, Liyao Sun, Yin Wang, Qingqing Sun, Shuai Chen, Tianyi Zhang
arXiv:2608. 07400v1 Announce Type: new Abstract: Financial question answering is typically evaluated by answer correctness, yet in SEC filings a plausible and even numerically correct answer can be grounded in the wrong evidence.
By Sasan Mansouri, Daniel Saad, Mark Wahrenburg, Manu Weissel, Fabian Woebbeking
arXiv:2607. 00738v1 Announce Type: cross Abstract: Large language models can generate polished scientific text that includes unsupported claims, allowing hallucinations to enter the archival record.
By Mark Russinovich, Ram Shankar Siva Kumar, Ahmed Salem
arXiv:2607. 17797v1 Announce Type: new Abstract: Financial statements (FS) such as Balance Sheet (BS), Income Statement (IS) and Cash-flow Statement (CS) summarize the annual financial performance of a company.
By Kshitij Madhav Jadhav, Sushodhan Vaishampayan, Manoj Apte, Sachin Pawar, Nitin Ramrakhiyani, Girish Keshav Palshikar
AtomCite is an agentic framework that verifies and corrects page‑level citations in multi‑page documents by parsing answers into claims, checking each claim against the cited page image, and applying a deterministic repair policy. The authors introduce DocCite, the first benchmark for this task, built on MP‑DocVQA and DUDE, containing 928 injected instances and 1,909 verified natural errors. Across Gemini, Claude, and GPT models, AtomCite achieves about 93% verification accuracy and improves citation precision from 34% to 87‑90%, while also enhancing hallucination detection in open‑source models.
By Chen Qian, Yimeng Wang, Yu Chen, Lingfei Wu, Andreas Stathopoulos
The paper introduces a claim‑gated audit framework for generative search, ensuring that a query, source, and answer tuple is only considered resolved when relationship evidence, answer adoption, materiality, and disclosure are all present. It distinguishes this audit endpoint from citation support and review priority, tying decisions to versioned evidence spans and implementing a reference checker to enforce the contract. Experiments on a synthetic dataset confirm that the system correctly handles all 81 predicate combinations and rejects 192 malformed records, while ablation studies isolate endpoint logic from missing‑evidence handling.
By Kainan Zhou, Chuhong Xu, Gangzhen Qian, Zhaoyi Li
AnalysisBank is a library that captures expert financial analysis by pairing data signals with analytical moves and the corresponding expert text spans. The system matches input signals to these library entries during inference, enabling the generation of reports that are grounded in data-derived insights rather than generic structural templates. Experiments on financial benchmarks show that AnalysisBank produces 1.7–3.7 times more novel, data‑grounded insights than structural baselines, and the approach also transfers to scientific writing.
By Yajing Yang, Yunshan Ma, Kelvin J. L. Koa, Min-Yen Kan
Financial disclosures contain numerical claims, temporal statements, entity references, policy commitments, and risk descriptions that may conflict in qualitatively different ways. Detecting a conflict is only the first step: review workflows may also need to determine its type, since numerical, temporal, referential, factual, and normative inconsistencies require different evidence and downstream checks.
arXiv:2604. 19755v2 Announce Type: replace Abstract: Anti-money laundering (AML) transaction monitoring generates large volumes of alerts that must be rapidly triaged by investigators under strict audit and governance constraints.
By Dorothy Torres, Wei Cheng, Ke Hu