arXiv AI

Cost-Aware Speculative Execution for LLM-Agent Workflows: An Integrated Five-Dimension Method

arXiv:2606. 07846v1 Announce Type: cross Abstract: LLM-agent workflows chain model calls and tool invocations, and spend most of their wall-clock time waiting on upstream operations before downstream ones can start.

arXiv AI
Sep 4

Speculative Macro Commit for Faster Tool-Using Agents

Speculative Macro Commit (SMC) is a runtime technique for tool‑using language‑model agents that separates an authoritative actor model from a faster speculative drafter model. The drafter predicts and executes future action chains on a snapshot, storing recurring multi‑action patterns in a macro library. When the actor’s next tool call aligns with a drafted action, SMC commits the pre‑executed steps, reducing latency by up to 18.59% on certain benchmarks while maintaining accuracy.

By Zeyu Liu, Souvik Kundu, Peter A. Beerel
Hugging Face Trending Papers
Sep 3

Speculative Macro Commit for Faster Tool-Using Agents

Speculative Macro Commit (SMC) is a runtime technique that speeds up tool‑using language‑model agents by having a fast speculative drafter model predict and execute future action chains on a separate environment snapshot. The drafter’s predictions are matched against a macro library of recurring multi‑action skeletons; when the authoritative actor’s next tool call aligns with the first drafted action, SMC commits the remaining pre‑executed steps to the official trajectory. Experiments with Qwen3.5 models show that SMC maintains overall accuracy while cutting latency by up to 18.6% on telecom benchmarks and 44.9% on AppWorld compared to sequential execution.

arXiv Machine Learning
Jul 1

Certified Speculative Execution for Untrusted AI Agents

arXiv:2606. 31023v1 Announce Type: cross Abstract: Hard-constrained sequential decision systems have no certified way to spend the test-time compute of modern AI: executing the multi-step drafts of a learned policy or a frozen LLM forfeits the feasibility guarantee a trusted solver provides, while invoking the solver at every step forfeits the speed the AI offers.

By Chenyu Zhou, Qiliang Jiang, Shuning Wu, Xu Zhou
arXiv AI
Aug 19

PACE: Policy-Attested Contract Execution for Safe AI Agents in Decentralized Finance

The paper introduces PACE (Policy‑Attested Contract Execution), a framework that sits between large‑language‑model (LLM) based autonomous AI agents and on‑chain DeFi operations. PACE defines typed transaction intents, a deterministic policy verifier, and signed Policy Decision Records (PDRs) that cryptographically bind an approved intent, policy, and simulation report to the exact on‑chain execution bytes, providing replay and expiration protection. In evaluations across 40 tasks and six baselines, PACE achieves zero unsafe executions and zero false positives, outperforming unguarded agents by a large margin.

By Rabimba Karanjai (Larry), Yang Lu (Larry), Richard Williamson (Larry), Hemanth Hm (Larry), Prakhar Mehrotra (Larry), Lei Xu (Larry), Weidong (Larry), Shi
arXiv AI
Sep 12

Decoupling Readiness from Release for Tail-Aware Scheduling of Agentic LLM Workflows

The paper introduces a tail‑risk‑aware scheduling strategy for agentic LLM workflows that decouples readiness from immediate release of model turns. By jointly selecting which ready turn to release and controlling the amount of unfinished work kept in the queue, the method uses a mean‑CVaR objective to adapt to evolving tail risk and online turn‑work estimates. Experiments on real software‑engineering task traces show comparable performance to eager release under light load and a significant reduction in the 95th‑percentile workflow flow time, achieving up to a 3.5× speedup under contention.

By Bochao Feng, Jianjiang Li, Haojie Wang, Lin Qiao, Yinghui Li, Yukun Yan, Jidong Zhai
arXiv Machine Learning
Sep 7

How to Speculate about Uncertainty in Agentic Coding? A Draft-Model Gate Method

The paper introduces Speculative Uncertainty (SU), a technique that infers a failure likelihood for black‑box LLM agents by evaluating their generated token sequences with a lightweight draft model, without needing internal model details. SU extracts phase‑aware features from reasoning and action spans, calibrates them against verifiable outcomes, and produces a failure‑likelihood score usable by downstream policies. Applying a pre‑execution veto gate based on SU to software‑engineering agents such as Qwen3‑Coder‑480B and Claude 3.5 Sonnet reduced execution error rates by 6‑8 percentage points and token costs by 14‑19 %, while maintaining performance on out‑of‑distribution benchmarks and across different agent models.

By Konstantin Grotov, Valentin Malykh