arXiv Machine Learning By Hoang Dang, Luan Pham, Minh Nguyen

Target-Weighted Neyman Allocation: Experimental Design for Heterogeneous Treatment Effects under Population Shift

Read the original on arXiv Machine Learning →

arXiv:2608. 06512v1 Announce Type: new Abstract: Randomized experiments are often run in one population to guide decisions in another.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 25

Speculative Evaluation of Stochastic LLMs

The paper introduces Speculative Evaluation, a method to reduce variance in evaluating stochastic large language models (LLMs) under a fixed rollout budget. It employs a Hierarchical Bayesian Neyman (HBN) policy that first runs a short uniform pilot, then pools task-level success counts via a hierarchical Bayesian model to compute posterior expectations of task-level sampling variances. Using these expectations, the method applies exact positive-integer Neyman allocation to allocate rollouts, and an asynchronous variant (HBN-async) speculatively executes continuations from partial pilot feedback to mitigate synchronization overhead. Across six checkpoints and 18 benchmark groups, Speculative Evaluation achieves 12.8%-33.6% lower variance compared to uniform allocation, outperforming empirical and independent Bayesian baselines, and demonstrates practical benefits in real-generation experiments.

By Qianli Shen, Xiang Li, Ruomeng Ding, Yanxi Chen, Daoyuan Chen, Yaliang Li
arXiv Machine Learning
4d ago

Tail-Influence Sampling for CVaR Policy Evaluation

arXiv:2609.38096v1 Announce Type: new Abstract: Policies with similar mean returns can differ sharply in rare failures, yet estimating lower-tail conditional value-at-risk (CVaR) accurately can requi...

By Pauline Bourigault, Xiaotong Ji, Matthieu Zimmer, Rasul Tutunov, Haitham Bou-Ammar
arXiv AI
Sep 18

Risk-Set Transported Synthetic Control with Difference-in-Differences Adjustment under Staggered Treatment Adoption

The paper introduces Risk-Set Transported Synthetic Control with Difference-in-Differences Adjustment (RT‑SC‑DiD), a method for staggered treatment‑adoption studies that keeps the donor pool fixed by reallocating weights from exiting donors to similar surviving donors while applying a DiD baseline correction. It analyzes distortion from horizon‑by‑horizon re‑optimization, derives bounds on error propagation, and proposes diagnostics and a donor‑only placebo for tuning the transport penalty. Empirical simulations show that intermediate transport regularization reduces average RMSE compared to independent horizon‑specific estimation and strong anchoring, supporting the method’s bias‑variance trade‑off. "whyItMatters":"The method offers a principled way to stabilize synthetic‑control weights over time in staggered designs, potentially improving causal inference when donor support contracts as treatments roll out."

By Mojtaba Eslami