The paper introduces Speculative Evaluation, a method to reduce variance in evaluating stochastic large language models (LLMs) under a fixed rollout budget. It employs a Hierarchical Bayesian Neyman (HBN) policy that first runs a short uniform pilot, then pools task-level success counts via a hierarchical Bayesian model to compute posterior expectations of task-level sampling variances. Using these expectations, the method applies exact positive-integer Neyman allocation to allocate rollouts, and an asynchronous variant (HBN-async) speculatively executes continuations from partial pilot feedback to mitigate synchronization overhead. Across six checkpoints and 18 benchmark groups, Speculative Evaluation achieves 12.8%-33.6% lower variance compared to uniform allocation, outperforming empirical and independent Bayesian baselines, and demonstrates practical benefits in real-generation experiments.
By Qianli Shen, Xiang Li, Ruomeng Ding, Yanxi Chen, Daoyuan Chen, Yaliang Li
arXiv:2607. 14604v1 Announce Type: new Abstract: Online controlled experiments are the gold standard for hypothesis testing in online platforms.
By Olivier Jeunen
arXiv:2609.38096v1 Announce Type: new
Abstract: Policies with similar mean returns can differ sharply in rare failures, yet estimating lower-tail conditional value-at-risk (CVaR) accurately can requi...
By Pauline Bourigault, Xiaotong Ji, Matthieu Zimmer, Rasul Tutunov, Haitham Bou-Ammar
arXiv:2608. 12489v1 Announce Type: new Abstract: Organizations decide whom to treat under a budget and want to know what a targeting rule would have earned before deploying it.
By Binshuang Li
arXiv:2608. 19383v1 Announce Type: cross Abstract: Average dose-response functions are widely used to summarize causal effects of continuous treatments, but most existing methods assume that the observed sample represents the target population.
By Jay Jojo Cheng, Guanhua Chen
The paper introduces Risk-Set Transported Synthetic Control with Difference-in-Differences Adjustment (RT‑SC‑DiD), a method for staggered treatment‑adoption studies that keeps the donor pool fixed by reallocating weights from exiting donors to similar surviving donors while applying a DiD baseline correction. It analyzes distortion from horizon‑by‑horizon re‑optimization, derives bounds on error propagation, and proposes diagnostics and a donor‑only placebo for tuning the transport penalty. Empirical simulations show that intermediate transport regularization reduces average RMSE compared to independent horizon‑specific estimation and strong anchoring, supporting the method’s bias‑variance trade‑off.
"whyItMatters":"The method offers a principled way to stabilize synthetic‑control weights over time in staggered designs, potentially improving causal inference when donor support contracts as treatments roll out."
By Mojtaba Eslami
The paper explores how data from fixed A/B tests can guide the deployment of adaptive experiments using contextual bandits. By combining off‑policy evaluation with a controlled warm‑start simulation, the authors rank pre‑specified adaptive and non‑adaptive policies using doubly robust estimators. Experiments on synthetic trials and real benchmarks show that adaptive, context‑aware policies outperform fixed allocations when heterogeneity exists, but offer little advantage otherwise.
By Jo\~ao Victor Ferreira Alves, Eduardo Rocha Laurentino, Gustavo de Oliveira Kanno, Thiago Costa Rizuti da Rocha
arXiv:2607. 07695v1 Announce Type: new Abstract: We introduce institutional red-teaming, an evaluation methodology for testing deployment rules in multi-agent AI: hold the agents, objectives, and task state fixed, vary only one rule, and attribute the resulting change in collective behavior to that rule.
By Yujiao Chen
arXiv:2607. 19692v1 Announce Type: cross Abstract: Observational causal analyses increasingly pool records across sites, vendors, and collection systems, creating vulnerability to append-only attacks in which plausible records are strategically selected to alter a reported treatment effect.
By Kwangho Kim
The paper introduces a new estimand for conditional distributional treatment effects that captures how treatments influence the entire outcome distribution, including variance and tail risks, in a covariate-dependent manner. It presents a doubly robust estimator that is minimax optimal locally and uses it to construct a test for global homogeneity of conditional potential outcome distributions. The test accommodates discrepancies beyond the maximum mean discrepancy, guarantees valid type‑1 error, is consistent against fixed alternatives, and includes a computationally efficient, permutation‑free algorithm with exact closed‑form expressions for two natural discrepancies.
By Saksham Jain, Alex Luedtke
arXiv:2606. 17165v1 Announce Type: cross Abstract: Organizations and researchers show increasing interest in using large language models (LLMs) in place of human participants in A/B tests, in the hope of experimenting faster and at lower cost.
By Joel Persson, M{\aa}rten Schultzberg, Sebastian Ankargren
arXiv:2609.38914v1 Announce Type: new
Abstract: Evaluating interactive agents is expensive. Agent behavior is stochastic, so reliability must be measured over repeated trials, but failures are rare a...
By Priyanath Maji, Spandan Ghose Chowdhury