Speculative Evaluation of Stochastic LLMs
Read the original on arXiv AI →The paper introduces Speculative Evaluation, a method to reduce variance in evaluating stochastic large language models (LLMs) under a fixed rollout budget. It employs a Hierarchical Bayesian Neyman (HBN) policy that first runs a short uniform pilot, then pools task-level success counts via a hierarchical Bayesian model to compute posterior expectations of task-level sampling variances. Using these expectations, the method applies exact positive-integer Neyman allocation to allocate rollouts, and an asynchronous variant (HBN-async) speculatively executes continuations from partial pilot feedback to mitigate synchronization overhead. Across six checkpoints and 18 benchmark groups, Speculative Evaluation achieves 12.8%-33.6% lower variance compared to uniform allocation, outperforming empirical and independent Bayesian baselines, and demonstrates practical benefits in real-generation experiments.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.