arXiv AI

Typed Federated Artifacts for the Agentic Web:Sharing Tool-Routing Knowledge Across Frozen,Heterogeneous LLM Agents

The paper proposes typed federated artifacts—schema‑validated objects with per‑field privacy and dispute resolution—to enable tool‑routing knowledge sharing among frozen, heterogeneous LLM agents. By replacing flat text prompts with typed fields, the authors achieve near‑centralized routing performance on StableToolBench while reducing data size to 20 MB JSON per client. The study also highlights that a simple TF‑IDF classifier can outperform LLM routing on labeled benchmarks, indicating limitations in current evaluation methods.

arXiv AI
4d ago

TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing

arXiv:2605.18859v3 Announce Type: replace-cross Abstract: LLM routing matters most in long-horizon applications such as coding agents, deep research systems, and computer-use agents, where a single u...

By Pei Yang, Wanyi Chen, Tongyun Yang, Pengbin Feng, Jiarong Xing, Wentao Guo, Yuhang Yao, Yuhang Han, Hanchen Li, Xu Wang, Zeyu Wang, Jie Xiao, Anjie Yang, Liang Tian, Lynn Ai, Eric Yang, Tianyu Shi
arXiv AI
Aug 28

BekchiAI: Measuring, Observing, and Controlling LLM Agents in One Click

BekchiAI introduces a benchmark and platform for evaluating large language model agents. The benchmark comprises 13 tool‑using ReAct agents across seven task categories, totaling 2,057 deterministic test tasks with verifier‑checkable gold answers. The platform offers web‑based observability, token and latency telemetry, and remote run termination for live agents.

By Mesut Toruk
arXiv AI
Sep 2

APEX-EM: Non-Parametric Online Learning for Autonomous Agents via Structured Procedural-Episodic Experience Replay

APEX-EM is a non‑parametric experience memory that stores full procedural‑episodic traces in a typed Procedural Knowledge Graph and retrieves them via semantic search, structural‑signature matching, and graph traversal. It uses a Plan‑Retrieve‑Generate‑Iterate‑Ingest workflow to produce, quality‑gate, and commit experiences, indexing both successes and failures so the agent learns what to reuse and what to avoid. Evaluations on five benchmarks with a shared GPT‑4o backbone show significant performance gains, such as +7.6 pp on BigCodeBench transfer and +1.4 pp on Lifelong Agent Bench, demonstrating that the memory adds to model capability rather than replacing it.

By Pratyay Banerjee, Masud Moshtaghi, Ankit Chadha
arXiv AI
6d ago

Cartograph: Federated Tool Discovery with Operator-Attested Retrieval for AI Agents

Cartograph is a federated Model Context Protocol (MCP) proxy that reduces AI agent tool discovery from linear catalog traversal to progressive disclosure, exposing only a few proxy tools instead of all definitions. It uses operator-attested capability cards, a three-layer confusable-cluster analysis called Rift, and a two-stage retrieval process to rank servers before tools. In a 22-server, 374-tool deployment, Cartograph achieves higher recall (R@5 = 0.816 vs. 0.592) and drastically fewer tokens (475 vs. 42,450) for discovery exchanges, with minimal latency overhead.

By Justice Owusu Agyemang, Michael Agyare, Kwame Opuni-Boachie Obour Agyekum, Kwame Agyeman-Prempeh Agyekum, Francisca Adoma Acheampong, Jerry John Kponyo
arXiv AI
Sep 18

Closed-World Resolution Against Tool Hallucination in LLM Agents

The paper investigates a new failure mode of tool‑augmented large language model agents: calling non‑existent tools with arguments that do not match any declared schema. It introduces a five‑class taxonomy of tool hallucination, presents a training‑free closed‑world resolver that checks registry membership and signatures, and demonstrates that hallucinations persist across ten hosted models and various invocation surfaces, including the Model Context Protocol. The authors release a Hallucinated‑Tools Benchmark to enable comparison of resolver methods.

By Laxmipriya Ganesh Iyer
arXiv AI
3d ago

TRACE: Trajectory Selection for Parallel Scaling of Search Agents

TRACE is a lightweight learned selector that ranks completed search trajectories by aggregating cross‑rollout evidence, preserving individual query and evidence occurrences while propagating information across shared content or document identity. Trained with answer‑level supervision over frozen text embeddings, TRACE selects an existing answer without additional search or autoregressive aggregation, and a single selector generalizes across rollout policies and agent backbones. Across six WebQA policies, six long‑horizon dataset‑backbone combinations, and multiple WebQA benchmarks, TRACE outperforms majority voting and generative aggregators, achieving higher accuracy and at least tenfold higher processing throughput.

By Qisheng Zhou, Zhen Xiong, Qiaoyu Tan