arXiv AI

Speculative Evaluation of Stochastic LLMs

The paper introduces Speculative Evaluation, a method to reduce variance in evaluating stochastic large language models (LLMs) under a fixed rollout budget. It employs a Hierarchical Bayesian Neyman (HBN) policy that first runs a short uniform pilot, then pools task-level success counts via a hierarchical Bayesian model to compute posterior expectations of task-level sampling variances. Using these expectations, the method applies exact positive-integer Neyman allocation to allocate rollouts, and an asynchronous variant (HBN-async) speculatively executes continuations from partial pilot feedback to mitigate synchronization overhead. Across six checkpoints and 18 benchmark groups, Speculative Evaluation achieves 12.8%-33.6% lower variance compared to uniform allocation, outperforming empirical and independent Bayesian baselines, and demonstrates practical benefits in real-generation experiments.

arXiv Computation and Language
Sep 21

Draft-OPD: On-Policy Distillation for Speculative Draft Models

Draft-OPD introduces an on‑policy distillation method for speculative draft models, addressing the mismatch between supervised fine‑tuning and inference by letting the target model supervise the drafter on draft‑induced states. The approach uses target‑assisted rollouts for stable continuations and replays drafting from error positions exposed during verification, enabling the drafter to learn from both accepted and rejected proposals. Experiments demonstrate that Draft‑OPD achieves more than five‑fold lossless acceleration across diverse tasks, outperforming prior draft models such as EAGLE‑3 and DFlash by 23 % and 13 % respectively.

By Haodi Lei, Yafu Li, Haoran Zhang, Shunkai Zhang, Qianjia Cheng, Xiaoye Qu, Ganqu Cui, Bowen Zhou, Ning Ding, Yun Luo, Yu Cheng
arXiv Machine Learning
1d ago

Tail-Influence Sampling for CVaR Policy Evaluation

arXiv:2609.38096v1 Announce Type: new Abstract: Policies with similar mean returns can differ sharply in rare failures, yet estimating lower-tail conditional value-at-risk (CVaR) accurately can requi...

By Pauline Bourigault, Xiaotong Ji, Matthieu Zimmer, Rasul Tutunov, Haitham Bou-Ammar
arXiv Computation and Language
Aug 25

Can Large Language Models "Hyper-Thread"?

arXiv:2608.22376v1 Announce Type: new Abstract: Large language models generate tokens sequentially, but can they execute multiple tasks concurrently while forming each token? Broader attention alloca...

By Fei Ding
arXiv AI
Aug 24

Beyond End-to-End Success: Diagnosing Failures in Long-Horizon Security LLM Agents

The paper introduces a diagnostic framework for long‑horizon security LLM agents that uses checkpoints to distinguish failures occurring before and after a model’s capability is exposed, and applies controlled interventions to pinpoint upstream bottlenecks. The methodology is tested on four task families—delayed reuse of discovered information, reuse of observed state, recovery from failed strategies, and decision making after uncertain outcomes—revealing that many failures happen before the agent observes the state it later needs to reuse. Experiments with Gemini 2.5 Flash and Gemini 3.7 Flash show that targeted protocol‑disambiguation guidance can significantly alter state observation rates and that the primary source of failure can shift across model generations, underscoring the need for fine‑grained failure diagnostics rather than relying solely on overall task success.

By Wei Shao, Chongzhou Fang, Zuxiong Tan, Zequan Liang, Setareh Rafatirad, Avesta Sasan, Houman Homayoun