arXiv AI

Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It

The paper argues that agentic systems waste time and memory by guessing how long tool calls will take, rather than using explicit progress signals from the tools themselves. It demonstrates that tools can report their remaining work or imminent completion, and that incorporating this feedback into serving systems dramatically improves cache decisions and reduces token latency. The authors show that this approach outperforms existing predictors and works robustly across different environments.

arXiv AI
Aug 18

From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems

arXiv:2608. 15127v1 Announce Type: cross Abstract: Agentic applications are shifting AI serving from isolated model inference to long-running workloads in which LLMs coordinate tools, environments, and persistent state.

By Chaokun Chang, Yukun Zhou, Kaihua Fu, Dakai An, Tianyu Feng, Hanfeng Lu, Sheng Yao, Pu Guo, Yinghao Yu, Yizhou Shan, Bo Li, Binhang Yuan, Wei Wang
arXiv AI
6d ago

Policy as Code: A Coroutine-Bridge Harness for Fast-Reasoning Reliability on CAR-bench

The paper introduces a coroutine-bridge harness that lets a language model emit a Python program to manage tool calls in the CAR-bench evaluation. By decoupling model invocations from tool round-trips, the approach reduces model calls to a median of two per task while maintaining seven agent turns, achieving a median latency of 1.8 s on a Cerebras gpt‑oss‑120b. The harness achieved 60.0 % Pass³ on the official hidden evaluation, outperforming the baseline by 4.5× and matching frontier-model agents on GPT‑5.5, all while keeping the prompt largely cached and minimizing input compute.

By Ivan Matveev
arXiv Machine Learning
Sep 7

Same Request, Different Answer: Quantization Amplifies Cache-Induced Divergence in LLM Serving

The paper investigates how prefix caching, a default optimization in open‑source LLM serving stacks, affects reproducibility when combined with weight quantization. Experiments on an eighty‑episode multi‑turn agentic tool‑use workload show that enabling the cache causes the agent’s trajectory to change in 36.2 % of episodes at 16‑bit precision and 75.0 % at 4‑bit precision, while disabling the cache yields perfectly reproducible runs. The study identifies specific cache‑related settings that drive run‑to‑run divergence and demonstrates that cached serving is deterministic only when the cache state is preserved, which is not the case in typical deployments.

By Aditi Patodiya