arXiv AI By Zhen Xu, Qizheng Zhang, Gerry Wan, Shang Zhu, Ce Zhang

LEAP: Learning Efficient Action Proposals For LLM Agents

Read the original on arXiv AI →

LEAP: Learning Efficient Action Proposals For LLM Agents proposes a method to speed up large language model agents by training a small 0.6B drafter to predict target actions accurately. The approach uses a latency framework that balances drafting, verification, and execution costs, achieving up to 60% faster end‑to‑end wall clock time without reducing task success. LEAP can be online trained, eliminating the need for prior trace collection and making it practical for real‑world deployment.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 4

Speculative Macro Commit for Faster Tool-Using Agents

Speculative Macro Commit (SMC) is a runtime technique for tool‑using language‑model agents that separates an authoritative actor model from a faster speculative drafter model. The drafter predicts and executes future action chains on a snapshot, storing recurring multi‑action patterns in a macro library. When the actor’s next tool call aligns with a drafted action, SMC commits the pre‑executed steps, reducing latency by up to 18.59% on certain benchmarks while maintaining accuracy.

By Zeyu Liu, Souvik Kundu, Peter A. Beerel
Hugging Face Trending Papers
Sep 3

Speculative Macro Commit for Faster Tool-Using Agents

Speculative Macro Commit (SMC) is a runtime technique that speeds up tool‑using language‑model agents by having a fast speculative drafter model predict and execute future action chains on a separate environment snapshot. The drafter’s predictions are matched against a macro library of recurring multi‑action skeletons; when the authoritative actor’s next tool call aligns with the first drafted action, SMC commits the remaining pre‑executed steps to the official trajectory. Experiments with Qwen3.5 models show that SMC maintains overall accuracy while cutting latency by up to 18.6% on telecom benchmarks and 44.9% on AppWorld compared to sequential execution.

arXiv Computation and Language
Sep 21

Draft-OPD: On-Policy Distillation for Speculative Draft Models

Draft-OPD introduces an on‑policy distillation method for speculative draft models, addressing the mismatch between supervised fine‑tuning and inference by letting the target model supervise the drafter on draft‑induced states. The approach uses target‑assisted rollouts for stable continuations and replays drafting from error positions exposed during verification, enabling the drafter to learn from both accepted and rejected proposals. Experiments demonstrate that Draft‑OPD achieves more than five‑fold lossless acceleration across diverse tasks, outperforming prior draft models such as EAGLE‑3 and DFlash by 23 % and 13 % respectively.

By Haodi Lei, Yafu Li, Haoran Zhang, Shunkai Zhang, Qianjia Cheng, Xiaoye Qu, Ganqu Cui, Bowen Zhou, Ning Ding, Yun Luo, Yu Cheng