arXiv AI

Making Failure Safe: A Constrained, Verifiable Agent Framework for Open-Web Data Collection

arXiv:2607. 00035v1 Announce Type: new Abstract: LLMs and agents can generate web scrapers from natural-language requirements, but direct generation remains unreliable because of dependency errors, broken selectors, schema mismatches, and heterogeneous page structures.

arXiv AI
Sep 25

Graph, Loop, and Harness Engineering for Zero-Trust Agentic Data Engineering and Analytical Processing

The paper introduces two zero‑trust frameworks for cloud data engineering and analytical processing. The first, Zero‑Trust Agentic Data Engineering, automatically generates, deploys, and verifies complete data‑engineering solutions from natural‑language tasks, requiring evidence from repositories, deployments, runtimes, and policies. The second, Zero‑Trust Agentic OLAP, combines governed data preparation with verified online analytical processing, allowing production promotion only after rigorous validation and evidence‑bound approval, and ensuring analytical outputs are released only after same‑snapshot execution, exact result equivalence, deterministic grounding, and reflection. Both frameworks rely on three core abstractions—graph engineering for evidence‑gated workflow structure, loop engineering for bounded recovery, and agent‑harness engineering for zero‑trust execution—and are evaluated under nominal execution, controlled failures, bounded recovery, and policy‑constrained conditions to measure verified completion, recovery, authorization enforcement, production promotion, and verified OLAP execution.

By Sagar Srinivas Sakhinana, Venkataramana Runkana
arXiv Machine Learning
Jul 8

Mitigating Errors in LLM-Generated Web API Invocations via Retrieval-Augmented Generation and Constrained Decoding

arXiv:2607. 05936v1 Announce Type: cross Abstract: Integration of web APIs is a cornerstone of modern software systems, yet writing correct web API invocation code remains challenging due to complex and evolving API specifications.

By Daniel Maninger, Leon Chemnitz, Jannis Brugger, Tushar Lamba, Amir Molzam Sharifloo, Mira Mezini
arXiv Computation and Language
Aug 27

Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World Systems

The paper argues that answer accuracy alone is insufficient for evaluating large language model (LLM) data agents, especially in structured-data tasks where a correct answer can be produced by an invalid trace. It introduces Trace Integrity as a reliability criterion that ensures the computation behind an answer is explicit, executable, schema-valid, operator-faithful, replayable, answer-consistent, and auditable. The authors operationalize this concept with execution contracts and present the CAIT (Correct Answer / Invalid Trace) Rate to quantify how often answer-only evaluations mistakenly reward unsupported outputs, demonstrating that accuracy, trace validity, and silent-failure risk are distinct signals.

By Srimonti Dutta, Akshata Kishore Moharir
arXiv AI
Jul 16

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

arXiv:2607. 13705v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical.

By Zichen Ding, Jiaye Ge, Shufan Jiang, Kai Chen, Mo Li, Qingqiu Li, Zehao Li, Zonglin Li, Tiaohao Liang, Shudong Liu, Zerun Ma, Zixing Shang, Wenhui Tian, Zun Wang, Liwei Wu, Zhenyu Wu, Jun Xu, Bowen Yang, Dingbo Yuan, Qi Zhang, Songyang Zhang, Peiheng Zhou, Dongsheng Zhu
arXiv AI
Jun 8

MetaConfigurator: AI-Assisted RDF Authoring from JSON Data

arXiv:2606. 07094v1 Announce Type: cross Abstract: Scientific workflows increasingly generate structured JSON data that is easy to exchange but difficult to interpret consistently across systems due to lacking semantic interoperability.

By Felix Neubauer, Mahdi Jafarkhani, Kenichi Endo, J\"urgen Pleiss, Benjamin Uekermann