The paper introduces Speculative Evaluation, a method to reduce variance in evaluating stochastic large language models (LLMs) under a fixed rollout budget. It employs a Hierarchical Bayesian Neyman (HBN) policy that first runs a short uniform pilot, then pools task-level success counts via a hierarchical Bayesian model to compute posterior expectations of task-level sampling variances. Using these expectations, the method applies exact positive-integer Neyman allocation to allocate rollouts, and an asynchronous variant (HBN-async) speculatively executes continuations from partial pilot feedback to mitigate synchronization overhead. Across six checkpoints and 18 benchmark groups, Speculative Evaluation achieves 12.8%-33.6% lower variance compared to uniform allocation, outperforming empirical and independent Bayesian baselines, and demonstrates practical benefits in real-generation experiments.
By Qianli Shen, Xiang Li, Ruomeng Ding, Yanxi Chen, Daoyuan Chen, Yaliang Li
arXiv:2607. 14604v1 Announce Type: new Abstract: Online controlled experiments are the gold standard for hypothesis testing in online platforms.
By Olivier Jeunen
arXiv:2609.38096v1 Announce Type: new
Abstract: Policies with similar mean returns can differ sharply in rare failures, yet estimating lower-tail conditional value-at-risk (CVaR) accurately can requi...
By Pauline Bourigault, Xiaotong Ji, Matthieu Zimmer, Rasul Tutunov, Haitham Bou-Ammar
arXiv:2608. 12489v1 Announce Type: new Abstract: Organizations decide whom to treat under a budget and want to know what a targeting rule would have earned before deploying it.
By Binshuang Li
arXiv:2608. 19383v1 Announce Type: cross Abstract: Average dose-response functions are widely used to summarize causal effects of continuous treatments, but most existing methods assume that the observed sample represents the target population.
By Jay Jojo Cheng, Guanhua Chen
The paper introduces Risk-Set Transported Synthetic Control with Difference-in-Differences Adjustment (RT‑SC‑DiD), a method for staggered treatment‑adoption studies that keeps the donor pool fixed by reallocating weights from exiting donors to similar surviving donors while applying a DiD baseline correction. It analyzes distortion from horizon‑by‑horizon re‑optimization, derives bounds on error propagation, and proposes diagnostics and a donor‑only placebo for tuning the transport penalty. Empirical simulations show that intermediate transport regularization reduces average RMSE compared to independent horizon‑specific estimation and strong anchoring, supporting the method’s bias‑variance trade‑off.
"whyItMatters":"The method offers a principled way to stabilize synthetic‑control weights over time in staggered designs, potentially improving causal inference when donor support contracts as treatments roll out."
By Mojtaba Eslami