The paper introduces Speculative Evaluation, a method to reduce variance in evaluating stochastic large language models (LLMs) under a fixed rollout budget. It employs a Hierarchical Bayesian Neyman (HBN) policy that first runs a short uniform pilot, then pools task-level success counts via a hierarchical Bayesian model to compute posterior expectations of task-level sampling variances. Using these expectations, the method applies exact positive-integer Neyman allocation to allocate rollouts, and an asynchronous variant (HBN-async) speculatively executes continuations from partial pilot feedback to mitigate synchronization overhead. Across six checkpoints and 18 benchmark groups, Speculative Evaluation achieves 12.8%-33.6% lower variance compared to uniform allocation, outperforming empirical and independent Bayesian baselines, and demonstrates practical benefits in real-generation experiments.
By Qianli Shen, Xiang Li, Ruomeng Ding, Yanxi Chen, Daoyuan Chen, Yaliang Li
arXiv:2606. 15600v1 Announce Type: cross Abstract: Cardinality-estimation (CE) research ranks estimators by q-error, yet it is well known that q-error is an imperfect proxy for query-plan quality.
By Madhulatha Mandarapu, Sandeep Kunkunuru
arXiv:2609. 29140v1 Announce Type: new Abstract: Repeated evaluation can estimate a benchmark score accurately while still requiring replication to certify narrow uncertainty.
By Yezhou Cheng, Runjia Du, Zeming Liu, Qibai Chen, Hang Lyu, Yilan Wei, Yankai Zeng, Bojun Lin
arXiv:2609.13149v1 Announce Type: new
Abstract: For local large language model agents, active context is a scarce resource: memory capacity, prefill latency, cache growth, and service objectives all...
By Aditya Karnam Gururaj Rao, Arjun Jaggi
arXiv:2607. 26253v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) is bottlenecked by rollout generation, yet many sampled prompts produce saturated groups (all responses correct or all incorrect) whose zero reward variance yields no policy-gradient signal.
By Pixel Nomand, Elena Voss, Marcus Hale, Sofia Reyes
Repeated evaluation can estimate a benchmark score accurately while still requiring replication to certify narrow uncertainty. We characterize that requirement on a fixed grid of $M$ tasks with $L$ binary paths per task under the hard budget $(M+t)K$, where each path costs at most $K$ responses or episodes.
arXiv:2608. 13087v1 Announce Type: cross Abstract: Neural combinatorial optimization (NCO) solvers report the best of many sampled solutions per instance, and the sample count is, by convention, identical for every instance.
By Jinhyung Bae
arXiv:2608. 12489v1 Announce Type: new Abstract: Organizations decide whom to treat under a budget and want to know what a targeting rule would have earned before deploying it.
By Binshuang Li
arXiv:2605. 06605v2 Announce Type: replace Abstract: Evaluating and predicting the performance of large language models (LLMs) in multi-turn conversational settings is critical yet computationally expensive; key events -- e.
By Shai Feldman, Yaniv Romano
arXiv:2607. 08665v1 Announce Type: new Abstract: Routing among large language models (LLMs) trades response quality against serving cost, motivated by the reported gap between deployed routers and a per-instance oracle.
By Teng-Ruei Chen
arXiv:2607. 13048v1 Announce Type: cross Abstract: Streaming inference pipelines increasingly pair lightweight fast models with Large Language Models (LLMs) that provide rich semantic understanding at substantial cost.
By Zhaohui Wang
arXiv:2609.07901v1 Announce Type: new
Abstract: Weight quantization largely determines the economics of serving open-weight LLMs. Its costs are usually assessed with capability benchmarks, on which 4...
By Dachi Kurtskhalia