arXiv:2607. 18261v1 Announce Type: new Abstract: LLM agents are increasingly used as transaction compilers: a user states an intent in natural language, and the model emits a structured object that an API can execute.
By Yin Li
The paper introduces two zero‑trust frameworks for cloud data engineering and analytical processing. The first, Zero‑Trust Agentic Data Engineering, automatically generates, deploys, and verifies complete data‑engineering solutions from natural‑language tasks, requiring evidence from repositories, deployments, runtimes, and policies. The second, Zero‑Trust Agentic OLAP, combines governed data preparation with verified online analytical processing, allowing production promotion only after rigorous validation and evidence‑bound approval, and ensuring analytical outputs are released only after same‑snapshot execution, exact result equivalence, deterministic grounding, and reflection. Both frameworks rely on three core abstractions—graph engineering for evidence‑gated workflow structure, loop engineering for bounded recovery, and agent‑harness engineering for zero‑trust execution—and are evaluated under nominal execution, controlled failures, bounded recovery, and policy‑constrained conditions to measure verified completion, recovery, authorization enforcement, production promotion, and verified OLAP execution.
By Sagar Srinivas Sakhinana, Venkataramana Runkana
arXiv:2607. 05936v1 Announce Type: cross Abstract: Integration of web APIs is a cornerstone of modern software systems, yet writing correct web API invocation code remains challenging due to complex and evolving API specifications.
By Daniel Maninger, Leon Chemnitz, Jannis Brugger, Tushar Lamba, Amir Molzam Sharifloo, Mira Mezini
arXiv:2609.13548v1 Announce Type: new
Abstract: Web agents can utilize reusable tools to reduce the cost and latency of low-level browser interaction, but automatically discovered tool collections ca...
By Xinyun Cao, Adriana Szekeres, Fazle Elahi Faisal
The paper argues that answer accuracy alone is insufficient for evaluating large language model (LLM) data agents, especially in structured-data tasks where a correct answer can be produced by an invalid trace. It introduces Trace Integrity as a reliability criterion that ensures the computation behind an answer is explicit, executable, schema-valid, operator-faithful, replayable, answer-consistent, and auditable. The authors operationalize this concept with execution contracts and present the CAIT (Correct Answer / Invalid Trace) Rate to quantify how often answer-only evaluations mistakenly reward unsupported outputs, demonstrating that accuracy, trace validity, and silent-failure risk are distinct signals.
By Srimonti Dutta, Akshata Kishore Moharir
arXiv:2607. 13705v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical.
By Zichen Ding, Jiaye Ge, Shufan Jiang, Kai Chen, Mo Li, Qingqiu Li, Zehao Li, Zonglin Li, Tiaohao Liang, Shudong Liu, Zerun Ma, Zixing Shang, Wenhui Tian, Zun Wang, Liwei Wu, Zhenyu Wu, Jun Xu, Bowen Yang, Dingbo Yuan, Qi Zhang, Songyang Zhang, Peiheng Zhou, Dongsheng Zhu
arXiv:2607. 02615v1 Announce Type: cross Abstract: Generating structured artifacts with Large Language Models - e.
By Yaniv Melamed, Yoni Zukerman, Michal Shechter, Miri Weissler, Ashwin Patil, Hani Neuvirth-Telem
arXiv:2606. 07094v1 Announce Type: cross Abstract: Scientific workflows increasingly generate structured JSON data that is easy to exchange but difficult to interpret consistently across systems due to lacking semantic interoperability.
By Felix Neubauer, Mahdi Jafarkhani, Kenichi Endo, J\"urgen Pleiss, Benjamin Uekermann
arXiv:2608. 14590v1 Announce Type: new Abstract: LLM agents increasingly perform irreversible real-world actions, including database updates, API calls, file operations, and autonomous use of tools.
By Pierre Dantas, Lucas Cordeiro, Ehsan Nowroozi, Tihanyi Norbert
Answer accuracy is an insufficient reliability signal for LLM data agents. In structured-data tasks, a benchmark-correct answer can be produced by an invalid trace. This paper introduces Trace Integri...
arXiv:2605.06445v2 Announce Type: replace-cross
Abstract: Large Language Model (LLM) agents demonstrate strong performance in autonomous code generation under loose specifications. However, productio...
By Francesco Dente, Dario Satriani, Paolo Papotti
arXiv:2606. 04769v1 Announce Type: cross Abstract: The Model Context Protocol (MCP) has emerged as a critical standard empowering Large Language Models (LLMs) to utilize external tools.
By Yutao Shi, Xiaohan Zhang, Xiangjing Zhang, Xihua Shen, Hui Ouyang, Huming Qiu, Mi Zhang, Min Yang