arXiv AI

Closed-World Resolution Against Tool Hallucination in LLM Agents

The paper investigates a new failure mode of tool‑augmented large language model agents: calling non‑existent tools with arguments that do not match any declared schema. It introduces a five‑class taxonomy of tool hallucination, presents a training‑free closed‑world resolver that checks registry membership and signatures, and demonstrates that hallucinations persist across ten hosted models and various invocation surfaces, including the Model Context Protocol. The authors release a Hallucinated‑Tools Benchmark to enable comparison of resolver methods.

arXiv AI
Sep 2

Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls

The paper introduces a benchmark for evaluating large language models (LLMs) on long‑horizon state tracking by having them compute the MD5 hash through 196 dependent tool calls across 64 rounds, carrying four 32‑bit words in context. It shows that a mixture‑of‑experts LLM can maintain the full state and produce correct digests in most runs, even when all primitive tools are replaced by another LLM. The study isolates state‑tracking difficulty from instruction interpretation and identifies key factors—contextual reasoning and worker voting—that enable success.

By Dheeraj Mohandas Pai, Lu Xian
arXiv AI
Sep 10

Typed Federated Artifacts for the Agentic Web:Sharing Tool-Routing Knowledge Across Frozen,Heterogeneous LLM Agents

The paper proposes typed federated artifacts—schema‑validated objects with per‑field privacy and dispute resolution—to enable tool‑routing knowledge sharing among frozen, heterogeneous LLM agents. By replacing flat text prompts with typed fields, the authors achieve near‑centralized routing performance on StableToolBench while reducing data size to 20 MB JSON per client. The study also highlights that a simple TF‑IDF classifier can outperform LLM routing on labeled benchmarks, indicating limitations in current evaluation methods.

By Abhijit Chakraborty, Ni Trieu, Vivek Gupta
arXiv AI
Sep 15

When Tool Calls Succeed but Workflows Fail: Anomalies at the Agent-Tool Boundary

The paper investigates how AI agents that run long workflows using external tools can experience inconsistencies when retries, speculative execution, concurrency, or partial failures occur. It introduces an effect‑history model that distinguishes between actual external events and the agent’s observations, and catalogs eight common external‑effect anomalies. The authors analyze the standard Model Context Protocol tool interface, finding that its annotations are too coarse to fully express the necessary capabilities to prevent these anomalies, thereby motivating the need for reusable transactional contracts at the agent‑tool boundary.

By Artem Trofimov, Boris Novikov
arXiv Computation and Language
Aug 31

Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction

The paper reports that a model can pass fidelity checks—verifying that extracted values match the source—without actually opening a datasheet, due to a hidden constraint that disables tool use. To address this, the authors log every tool call in an agentic benchmark and develop two instruments: a rule‑based failure‑attribution classifier and a silent‑failure detector that flags runs based solely on which tools were invoked. While the detector shows low false positives on clean extractions and recovers all planted faults, its recall against correct tool usage but incorrect answers remains unmeasured, and a partial causal chamber confirms only a subset of claims, highlighting limitations in physical verification.

By Qing Ye, Meng-Hsuan Lin