arXiv Machine Learning

Target-Weighted Neyman Allocation: Experimental Design for Heterogeneous Treatment Effects under Population Shift

arXiv:2608. 06512v1 Announce Type: new Abstract: Randomized experiments are often run in one population to guide decisions in another.

arXiv AI
Sep 25

Speculative Evaluation of Stochastic LLMs

The paper introduces Speculative Evaluation, a method to reduce variance in evaluating stochastic large language models (LLMs) under a fixed rollout budget. It employs a Hierarchical Bayesian Neyman (HBN) policy that first runs a short uniform pilot, then pools task-level success counts via a hierarchical Bayesian model to compute posterior expectations of task-level sampling variances. Using these expectations, the method applies exact positive-integer Neyman allocation to allocate rollouts, and an asynchronous variant (HBN-async) speculatively executes continuations from partial pilot feedback to mitigate synchronization overhead. Across six checkpoints and 18 benchmark groups, Speculative Evaluation achieves 12.8%-33.6% lower variance compared to uniform allocation, outperforming empirical and independent Bayesian baselines, and demonstrates practical benefits in real-generation experiments.

By Qianli Shen, Xiang Li, Ruomeng Ding, Yanxi Chen, Daoyuan Chen, Yaliang Li
arXiv Machine Learning
4d ago

Tail-Influence Sampling for CVaR Policy Evaluation

arXiv:2609.38096v1 Announce Type: new Abstract: Policies with similar mean returns can differ sharply in rare failures, yet estimating lower-tail conditional value-at-risk (CVaR) accurately can requi...

By Pauline Bourigault, Xiaotong Ji, Matthieu Zimmer, Rasul Tutunov, Haitham Bou-Ammar
arXiv AI
Sep 18

Risk-Set Transported Synthetic Control with Difference-in-Differences Adjustment under Staggered Treatment Adoption

The paper introduces Risk-Set Transported Synthetic Control with Difference-in-Differences Adjustment (RT‑SC‑DiD), a method for staggered treatment‑adoption studies that keeps the donor pool fixed by reallocating weights from exiting donors to similar surviving donors while applying a DiD baseline correction. It analyzes distortion from horizon‑by‑horizon re‑optimization, derives bounds on error propagation, and proposes diagnostics and a donor‑only placebo for tuning the transport penalty. Empirical simulations show that intermediate transport regularization reduces average RMSE compared to independent horizon‑specific estimation and strong anchoring, supporting the method’s bias‑variance trade‑off. "whyItMatters":"The method offers a principled way to stabilize synthetic‑control weights over time in staggered designs, potentially improving causal inference when donor support contracts as treatments roll out."

By Mojtaba Eslami
arXiv Machine Learning
5d ago

Offline Policy Evaluation as a decision support tool for designing Adaptive Experiments

The paper explores how data from fixed A/B tests can guide the deployment of adaptive experiments using contextual bandits. By combining off‑policy evaluation with a controlled warm‑start simulation, the authors rank pre‑specified adaptive and non‑adaptive policies using doubly robust estimators. Experiments on synthetic trials and real benchmarks show that adaptive, context‑aware policies outperform fixed allocations when heterogeneity exists, but offer little advantage otherwise.

By Jo\~ao Victor Ferreira Alves, Eduardo Rocha Laurentino, Gustavo de Oliveira Kanno, Thiago Costa Rizuti da Rocha
arXiv Machine Learning
Jul 23

Data-Poisoning Audits for Causal Effect Estimation

arXiv:2607. 19692v1 Announce Type: cross Abstract: Observational causal analyses increasingly pool records across sites, vendors, and collection systems, creating vulnerability to append-only attacks in which plausible records are strategically selected to alter a reported treatment effect.

By Kwangho Kim
arXiv Machine Learning
Sep 23

Conditional Distributional Treatment Effects: Doubly Robust Estimation and Testing

The paper introduces a new estimand for conditional distributional treatment effects that captures how treatments influence the entire outcome distribution, including variance and tail risks, in a covariate-dependent manner. It presents a doubly robust estimator that is minimax optimal locally and uses it to construct a test for global homogeneity of conditional potential outcome distributions. The test accommodates discrepancies beyond the maximum mean discrepancy, guarantees valid type‑1 error, is consistent against fixed alternatives, and includes a computationally efficient, permutation‑free algorithm with exact closed‑form expressions for two natural discrepancies.

By Saksham Jain, Alex Luedtke