arXiv Machine Learning By Hao Mark Chen, Jinnan Guo, Wayne Luk, Hongxiang Fan

AOSpec: Action and Observation Co-Speculation for Low-Latency Agent Serving

Read the original on arXiv Machine Learning →

arXiv:2608. 00881v1 Announce Type: new Abstract: Large language model agents increasingly act through stateful tools, yet model generation and environment execution remain serialized at every step.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 18

From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems

arXiv:2608. 15127v1 Announce Type: cross Abstract: Agentic applications are shifting AI serving from isolated model inference to long-running workloads in which LLMs coordinate tools, environments, and persistent state.

By Chaokun Chang, Yukun Zhou, Kaihua Fu, Dakai An, Tianyu Feng, Hanfeng Lu, Sheng Yao, Pu Guo, Yinghao Yu, Yizhou Shan, Bo Li, Binhang Yuan, Wei Wang
arXiv AI
Sep 4

Speculative Macro Commit for Faster Tool-Using Agents

Speculative Macro Commit (SMC) is a runtime technique for tool‑using language‑model agents that separates an authoritative actor model from a faster speculative drafter model. The drafter predicts and executes future action chains on a snapshot, storing recurring multi‑action patterns in a macro library. When the actor’s next tool call aligns with a drafted action, SMC commits the pre‑executed steps, reducing latency by up to 18.59% on certain benchmarks while maintaining accuracy.

By Zeyu Liu, Souvik Kundu, Peter A. Beerel