The paper argues that answer accuracy alone is insufficient for evaluating large language model (LLM) data agents, especially in structured-data tasks where a correct answer can be produced by an invalid trace. It introduces Trace Integrity as a reliability criterion that ensures the computation behind an answer is explicit, executable, schema-valid, operator-faithful, replayable, answer-consistent, and auditable. The authors operationalize this concept with execution contracts and present the CAIT (Correct Answer / Invalid Trace) Rate to quantify how often answer-only evaluations mistakenly reward unsupported outputs, demonstrating that accuracy, trace validity, and silent-failure risk are distinct signals.
By Srimonti Dutta, Akshata Kishore Moharir
Answer accuracy is an insufficient reliability signal for LLM data agents. In structured-data tasks, a benchmark-correct answer can be produced by an invalid trace. This paper introduces Trace Integri...
arXiv:2609.25192v1 Announce Type: new
Abstract: Financial search is a highly demanding task for LLM agents, requiring not only a correct final answer but also temporally valid information retrieval,...
By Wenqing Wang, Haitao Xiang, Xinyi Zhao, Mingming Yin, Ying Zhong, Zhaoxin Huan, Qiheng Zhou, Jin Zhu, Xiaolu Zhang, Shi Chang, Jun Zhou
The paper introduces Chart‑RVR, a reinforcement‑learning framework that trains chart‑reasoning agents to produce monitorable, verifiable outputs. It decomposes reasoning into three auditable stages—Structure, Evidence, and Derivation—allowing stakeholders to trace how the model reads the chart, extracts data, and computes the answer. Experiments on six benchmarks show that Chart‑RVR matches or exceeds state‑of‑the‑art accuracy while delivering rationales that are far more verifiable and evidence‑grounded than existing methods.
By Sanchit Sinha, Oana Frunza, Kashif Rasul, Aidong Zhang
arXiv:2608. 13706v1 Announce Type: cross Abstract: Existing defenses against hallucination in retrieval-augmented and multi-agent pipelines remain partial: evidence is trusted despite modality disagreement, debate verifies an aggregate report rather than individual claims, and such verification occurs only after drafting, leaving inter-agent errors undetected until the final text.
By Fatema Tuj Johora Faria, Mukaffi Bin Moin, Jubayer Al Mahmud, M. F. Mridha, Md. Alam Hossain
arXiv:2606. 11537v1 Announce Type: new Abstract: Financial and tabular question answering requires more than fluent reasoning: answers must be grounded in the exact facts, formulas, units, signs, and scales that support them.
By Abdelrahman Abdallah, AbdelRahim A. Elmadany, Sameh Al Natour, Hasan Cavusoglu, Adam Jatowt, Muhammad Abdul-Mageed
The paper introduces Self-Improving Retrieval-Augmented Generation (RAG), a framework that splits document question answering into Retrieval, Reasoning, and Judge agents coordinated by an orchestrator. When the Judge scores an answer below a dynamic threshold, the system retries with broader retrieval, more careful prompting, and relaxed acceptance criteria, achieving 86% oracle-guided accuracy on FinanceBench with a 36.4% Lazarus Rate. The approach logs every decision with confidence scores, providing audit trails needed for regulated financial applications.
By Junjie Xiong, Shawheen Ghezavat, Aum Hirpara
arXiv:2609.15319v1 Announce Type: cross
Abstract: Frontier models score well on shallow document/chart reading tasks. In a controlled data-room audit, moving evidence into buried conditions reduced a...
By Luis M. S\'anchez
arXiv:2608. 16386v1 Announce Type: cross Abstract: Financial agents must do more than recall domain knowledge: they must be both reliable, executing precise operations over grounded evidence, and executive, sustaining long-horizon research whose conclusions remain auditable.
By Agent Team, B. Zhang, Yaze Geng, Lei Tang, Yaoyang Yi, Zonghan Wu, Yifan Hu, Kun Wang, Qingsong Wen, Yilei Shao
arXiv:2604. 16706v2 Announce Type: replace Abstract: Automated evaluation of tool-using large language model (LLM) agents is widely assumed to be reliable, yet this assumption is rarely validated against human annotation.
By Bhaskar Gurram
arXiv:2606. 03031v1 Announce Type: new Abstract: Structured financial audit verification is difficult for language-model agents because correctness depends on structured evidence rather than text alone.
By Yan Wang, Xuguang Ai, Jaisal Patel, Xueqing Peng, Fengran Mo, Yupeng Cao, Haohang Li, Mingyu Cao, Lingfei Qian, V\'ictor Guti\'errez-Basulto
arXiv:2608. 06108v1 Announce Type: new Abstract: Investment competence is inherently personalized: the same market evidence can justify different actions for investors with different goals, horizons, portfolios, and risk boundaries.
By Yuanhong Jiang, Jingjie Zou, Zhenghong Lin, Xusheng Yu, Qiqi Huang, Shuai Jia, Shijie Dai