arXiv AI

Faithful, Not Corrective: Message-Format Effects in Multi-Hop Agent Relays Are Tier-Dependent

arXiv:2607. 09678v1 Announce Type: new Abstract: When LLM agents hand off information to one another, does the message format matter?

arXiv AI
Sep 18

Faithful, Not Corrective: Model Capability Governs Message-Format Effects in Multi-Hop Agent Relays

The study investigates how different message formats affect the fidelity of information as it passes through multiple LLM agent relays. Using a controlled testbed, the authors encode twelve atomic facts in five formats (free natural language, precision‑instructed NL, JSON, triples, key‑value) across six hops and evaluate recall against programmatic ground truth. Results show that strong relays maintain near‑lossless recall for all formats, while weaker relays exhibit significant format‑dependent recall loss, and that any injected error is faithfully propagated across all formats without causing collateral damage.

By Sicheng Zeng
arXiv AI
Aug 24

Nexus: Depth-Adaptive KV-Cache Splicing and Retrieval-Decoupled Tool Routing for Agentic LLMs on Unified Memory

Nexus introduces a depth‑adaptive KV‑cache splicing and retrieval‑decoupled tool routing mechanism for agentic large language models that reduces the time‑to‑first‑token (TTFT) by decoupling tool routing from the expensive schema re‑encoding step. It uses an INT8 semantic lookaside buffer to select tools via retrieval and generates arguments from a compressed textual signature, maintaining about 89% routing accuracy even as the tool registry scales to 250 tools. Additionally, Nexus can splice compiled schema KV blocks into the live context, repairing the seam with a depth‑adaptive suffix redecode when rotary position embedding drift exceeds a threshold, ensuring output fidelity while achieving up to 1.7× TTFT speedup at moderate depth.

By Mustafa Arslan
Hugging Face Trending Papers
Jul 30

Fidelity Is Not Safety: Gently-Compressed LLMs Pass Every Data-Free Quality Guard Yet Invent Procedure Steps in Agentic Execution

Practitioners accept a compressed language model once it clears a stack of data-cheap quality guards: perplexity within a small factor of the original, downstream accuracy (for example MMLU) inside a confidence interval, and data-free output-fidelity signals that compare the compressed and original network's internal representations under random probe inputs. This stack has a blind spot.