TraceLab: Characterizing Coding Agent Workloads for LLM Serving
arXiv:2606. 30560v1 Announce Type: cross Abstract: Coding agents are rapidly becoming a major application of agentic LLMs, but serving them efficiently remains challenging.
The paper argues that agentic systems waste time and memory by guessing how long tool calls will take, rather than using explicit progress signals from the tools themselves. It demonstrates that tools can report their remaining work or imminent completion, and that incorporating this feedback into serving systems dramatically improves cache decisions and reduces token latency. The authors show that this approach outperforms existing predictors and works robustly across different environments.
arXiv:2606. 30560v1 Announce Type: cross Abstract: Coding agents are rapidly becoming a major application of agentic LLMs, but serving them efficiently remains challenging.
arXiv:2608. 15127v1 Announce Type: cross Abstract: Agentic applications are shifting AI serving from isolated model inference to long-running workloads in which LLMs coordinate tools, environments, and persistent state.
arXiv:2608. 05784v1 Announce Type: new Abstract: Computer-use agents pay full frontier inference to re-derive routines their user has already performed, because an agent's memory today records what the user said, not what the user did.
arXiv:2609.23058v1 Announce Type: new Abstract: Current agent runtimes that plan before acting generally execute a step once it becomes ready. We present LazyAgent, a unified execution framework for...
arXiv:2608. 00101v1 Announce Type: cross Abstract: AI coding agents like GitHub Copilot, Claude Code, and Codex interleave multi-step LLM inference with tool execution, creating a workload different from chatbots.
arXiv:2603.01209v3 Announce Type: replace Abstract: In CodeAct, language-model agents write Python that calls tools and use execution feedback to choose actions. Persistent runtimes preserve Python v...
arXiv:2609.38201v1 Announce Type: new Abstract: Long-running tools can dominate coding-agent latency: compilers, test suites, and repository commands take seconds to minutes while the agent idles. Th...
arXiv:2607. 25816v1 Announce Type: new Abstract: Large language model agents often spend substantial wall-clock time waiting for tool call results.
The paper introduces a coroutine-bridge harness that lets a language model emit a Python program to manage tool calls in the CAR-bench evaluation. By decoupling model invocations from tool round-trips, the approach reduces model calls to a median of two per task while maintaining seven agent turns, achieving a median latency of 1.8 s on a Cerebras gpt‑oss‑120b. The harness achieved 60.0 % Pass³ on the official hidden evaluation, outperforming the baseline by 4.5× and matching frontier-model agents on GPT‑5.5, all while keeping the prompt largely cached and minimizing input compute.
arXiv:2608.22510v1 Announce Type: new Abstract: Agent benchmarks often evaluate only final answers even when agents run on stateful runtimes. We argue this under-specifies what is being evaluated: th...
arXiv:2609.37315v1 Announce Type: cross Abstract: Tool-using agents are entering settings where a wrong action carries real cost, and the benchmarks certifying them grade what each simulated tool cal...
The paper investigates how prefix caching, a default optimization in open‑source LLM serving stacks, affects reproducibility when combined with weight quantization. Experiments on an eighty‑episode multi‑turn agentic tool‑use workload show that enabling the cache causes the agent’s trajectory to change in 36.2 % of episodes at 16‑bit precision and 75.0 % at 4‑bit precision, while disabling the cache yields perfectly reproducible runs. The study identifies specific cache‑related settings that drive run‑to‑run divergence and demonstrates that cached serving is deterministic only when the cache state is preserved, which is not the case in typical deployments.