Synapse: Federated Tool Routing via Typed Compendium Artifacts
arXiv:2602. 00911v2 Announce Type: replace Abstract: The unit of collaboration in federated learning determines what guarantees are even expressible.
The paper proposes typed federated artifacts—schema‑validated objects with per‑field privacy and dispute resolution—to enable tool‑routing knowledge sharing among frozen, heterogeneous LLM agents. By replacing flat text prompts with typed fields, the authors achieve near‑centralized routing performance on StableToolBench while reducing data size to 20 MB JSON per client. The study also highlights that a simple TF‑IDF classifier can outperform LLM routing on labeled benchmarks, indicating limitations in current evaluation methods.
arXiv:2602. 00911v2 Announce Type: replace Abstract: The unit of collaboration in federated learning determines what guarantees are even expressible.
arXiv:2608. 16502v1 Announce Type: new Abstract: Large-scale agents increasingly rely on retrieval to access external capabilities.
arXiv:2608. 02680v1 Announce Type: cross Abstract: Tool-using language-model agents repeatedly rediscover procedures they have already executed, producing traces that mix reusable structure with retries, exploration, accidental ordering, and repeated lookups.
arXiv:2609.15982v1 Announce Type: cross Abstract: Skills extend an LLM agent beyond its parametric knowledge, and the gain they promise rests on picking the right one. Deployed harnesses route by pre...
arXiv:2605.18859v3 Announce Type: replace-cross Abstract: LLM routing matters most in long-horizon applications such as coding agents, deep research systems, and computer-use agents, where a single u...
arXiv:2608.21375v1 Announce Type: new Abstract: Heterogeneous agentic retrieval-augmented generation (RAG) systems increasingly orchestrate external APIs, internal databases, vector stores, and graph...
BekchiAI introduces a benchmark and platform for evaluating large language model agents. The benchmark comprises 13 tool‑using ReAct agents across seven task categories, totaling 2,057 deterministic test tasks with verifier‑checkable gold answers. The platform offers web‑based observability, token and latency telemetry, and remote run termination for live agents.
APEX-EM is a non‑parametric experience memory that stores full procedural‑episodic traces in a typed Procedural Knowledge Graph and retrieves them via semantic search, structural‑signature matching, and graph traversal. It uses a Plan‑Retrieve‑Generate‑Iterate‑Ingest workflow to produce, quality‑gate, and commit experiences, indexing both successes and failures so the agent learns what to reuse and what to avoid. Evaluations on five benchmarks with a shared GPT‑4o backbone show significant performance gains, such as +7.6 pp on BigCodeBench transfer and +1.4 pp on Lifelong Agent Bench, demonstrating that the memory adds to model capability rather than replacing it.
Cartograph is a federated Model Context Protocol (MCP) proxy that reduces AI agent tool discovery from linear catalog traversal to progressive disclosure, exposing only a few proxy tools instead of all definitions. It uses operator-attested capability cards, a three-layer confusable-cluster analysis called Rift, and a two-stage retrieval process to rank servers before tools. In a 22-server, 374-tool deployment, Cartograph achieves higher recall (R@5 = 0.816 vs. 0.592) and drastically fewer tokens (475 vs. 42,450) for discovery exchanges, with minimal latency overhead.
The paper investigates a new failure mode of tool‑augmented large language model agents: calling non‑existent tools with arguments that do not match any declared schema. It introduces a five‑class taxonomy of tool hallucination, presents a training‑free closed‑world resolver that checks registry membership and signatures, and demonstrates that hallucinations persist across ten hosted models and various invocation surfaces, including the Model Context Protocol. The authors release a Hallucinated‑Tools Benchmark to enable comparison of resolver methods.
arXiv:2606. 14476v1 Announce Type: new Abstract: A growing line of work equips large language model (LLM) agents with graph neural networks (GNNs) as callable tools, assuming the agent exercises judgment over when and how much to rely on such a tool.
TRACE is a lightweight learned selector that ranks completed search trajectories by aggregating cross‑rollout evidence, preserving individual query and evidence occurrences while propagating information across shared content or document identity. Trained with answer‑level supervision over frozen text embeddings, TRACE selects an existing answer without additional search or autoregressive aggregation, and a single selector generalizes across rollout policies and agent backbones. Across six WebQA policies, six long‑horizon dataset‑backbone combinations, and multiple WebQA benchmarks, TRACE outperforms majority voting and generative aggregators, achieving higher accuracy and at least tenfold higher processing throughput.