Parallelizing Tool Execution and LLM Generation for Low-Latency Agent Serving
arXiv:2603. 18897v2 Announce Type: replace-cross Abstract: LLM-powered agents execute tasks through a sequential loop of model generation and tool execution.
arXiv:2603. 18897v2 Announce Type: replace-cross Abstract: LLM-powered agents execute tasks through a sequential loop of model generation and tool execution.
Speculative Macro Commit (SMC) is a runtime technique for tool‑using language‑model agents that separates an authoritative actor model from a faster speculative drafter model. The drafter predicts and executes future action chains on a snapshot, storing recurring multi‑action patterns in a macro library. When the actor’s next tool call aligns with a drafted action, SMC commits the pre‑executed steps, reducing latency by up to 18.59% on certain benchmarks while maintaining accuracy.
Speculative Macro Commit (SMC) is a runtime technique that speeds up tool‑using language‑model agents by having a fast speculative drafter model predict and execute future action chains on a separate environment snapshot. The drafter’s predictions are matched against a macro library of recurring multi‑action skeletons; when the authoritative actor’s next tool call aligns with the first drafted action, SMC commits the remaining pre‑executed steps to the official trajectory. Experiments with Qwen3.5 models show that SMC maintains overall accuracy while cutting latency by up to 18.6% on telecom benchmarks and 44.9% on AppWorld compared to sequential execution.
arXiv:2607. 23933v1 Announce Type: cross Abstract: As LLM agents increasingly rely on the Model Context Protocol (MCP) to invoke isolated external sandboxes, disaggregated sandbox deployment introduces a fundamental tension between resource utilization and interactive tail latency.
arXiv:2608. 00881v1 Announce Type: new Abstract: Large language model agents increasingly act through stateful tools, yet model generation and environment execution remain serialized at every step.
arXiv:2607. 08010v1 Announce Type: cross Abstract: Production LLM agents often waste latency and reliability by regenerating code for the same procedural steps on every request.
arXiv:2602.18998v2 Announce Type: replace Abstract: LLM agents are increasingly expected to operate as general-purpose systems that resolve real-world user requests, yet their dynamic scaling behavio...
arXiv:2609.35889v1 Announce Type: cross Abstract: Tool-using language-model agents select and execute third-party artifacts. Different implementations can return the requested output while producing...
arXiv:2608. 16381v1 Announce Type: new Abstract: Agentic systems often organize execution and state around a single conversation, model invocation, or agent instance, even when real work spans many calls and stages.
arXiv:2607. 03333v1 Announce Type: cross Abstract: LLM agents are becoming a common interface for research, coding, and question answering, yet their Thought-Action-Observation loop is often serial: the model reasons, emits a tool call, then idles the GPU until the result returns.
arXiv:2607. 10059v1 Announce Type: new Abstract: Agent systems based on large language models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents know when to abstain.
arXiv:2608. 00101v1 Announce Type: cross Abstract: AI coding agents like GitHub Copilot, Claude Code, and Codex interleave multi-step LLM inference with tool execution, creating a workload different from chatbots.