arXiv AI

AgentFinVQA: A Deployable Multi-Agent Pipeline for Auditable Financial Chart QA

arXiv:2606. 19782v1 Announce Type: new Abstract: Financial chart question answering in regulated settings demands more than accuracy: practitioners must know which answers to trust before acting on them, and many institutions cannot send client data to external model providers.

arXiv Computation and Language
Aug 27

Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World Systems

The paper argues that answer accuracy alone is insufficient for evaluating large language model (LLM) data agents, especially in structured-data tasks where a correct answer can be produced by an invalid trace. It introduces Trace Integrity as a reliability criterion that ensures the computation behind an answer is explicit, executable, schema-valid, operator-faithful, replayable, answer-consistent, and auditable. The authors operationalize this concept with execution contracts and present the CAIT (Correct Answer / Invalid Trace) Rate to quantify how often answer-only evaluations mistakenly reward unsupported outputs, demonstrating that accuracy, trace validity, and silent-failure risk are distinct signals.

By Srimonti Dutta, Akshata Kishore Moharir
arXiv Computation and Language
Sep 23

FinFIRST: Benchmarking Search Agents for Financial Information Retrieval, Sourcing and Traceability

arXiv:2609.25192v1 Announce Type: new Abstract: Financial search is a highly demanding task for LLM agents, requiring not only a correct final answer but also temporally valid information retrieval,...

By Wenqing Wang, Haitao Xiang, Xinyi Zhao, Mingming Yin, Ying Zhong, Zhaoxin Huan, Qiheng Zhou, Jin Zhu, Xiaolu Zhang, Shi Chang, Jun Zhou
arXiv Computer Vision
Sep 22

Monitorable Chart Reasoning Agents via Verifiable Process Rewards

The paper introduces Chart‑RVR, a reinforcement‑learning framework that trains chart‑reasoning agents to produce monitorable, verifiable outputs. It decomposes reasoning into three auditable stages—Structure, Evidence, and Derivation—allowing stakeholders to trace how the model reads the chart, extracts data, and computes the answer. Experiments on six benchmarks show that Chart‑RVR matches or exceeds state‑of‑the‑art accuracy while delivering rationales that are far more verifiable and evidence‑grounded than existing methods.

By Sanchit Sinha, Oana Frunza, Kashif Rasul, Aidong Zhang
arXiv AI
Aug 17

CLAIR-Fin: An Adversarial Multi-Agent Framework for Claim-Level Verification and Adaptive Debate in Cross-Modal Financial QA

arXiv:2608. 13706v1 Announce Type: cross Abstract: Existing defenses against hallucination in retrieval-augmented and multi-agent pipelines remain partial: evidence is trusted despite modality disagreement, debate verifies an aggregate report rather than individual claims, and such verification occurs only after drafting, leaving inter-agent errors undetected until the final text.

By Fatema Tuj Johora Faria, Mukaffi Bin Moin, Jubayer Al Mahmud, M. F. Mridha, Md. Alam Hossain
arXiv Computation and Language
Aug 28

Towards Expert Financial QA via Self-Improving RAG

The paper introduces Self-Improving Retrieval-Augmented Generation (RAG), a framework that splits document question answering into Retrieval, Reasoning, and Judge agents coordinated by an orchestrator. When the Judge scores an answer below a dynamic threshold, the system retries with broader retrieval, more careful prompting, and relaxed acceptance criteria, achieving 86% oracle-guided accuracy on FinanceBench with a 36.4% Lazarus Rate. The approach logs every decision with confidence scores, providing audit trails needed for regulated financial applications.

By Junjie Xiong, Shawheen Ghezavat, Aum Hirpara
arXiv Machine Learning
Aug 18

Mint-Agent: Introducing Finance-Native Agentic Foundation Models

arXiv:2608. 16386v1 Announce Type: cross Abstract: Financial agents must do more than recall domain knowledge: they must be both reliable, executing precise operations over grounded evidence, and executive, sustaining long-horizon research whose conclusions remain auditable.

By Agent Team, B. Zhang, Yaze Geng, Lei Tang, Yaoyang Yi, Zonghan Wu, Yifan Hu, Kun Wang, Qingsong Wen, Yilei Shao