arXiv Machine Learning

Counterfactual Online Conformal Prediction Under Adaptive Logging

The paper addresses the failure of online conformal prediction when predictions influence actions that determine which outcomes are used for calibration. It introduces Propensity-Weighted Online Conformal Prediction (PW‑OCP), an inverse‑propensity‑weighted recursion that debiases calibration, and a doubly robust variant (DR‑OCP) that further reduces bias. Experiments on synthetic decision tasks, open bandit data, and financial rebalancing demonstrate that PW‑OCP and DR‑OCP improve counterfactual coverage and downstream regret while preserving prediction‑set sharpness.

arXiv Machine Learning
2d ago

Offline Policy Evaluation as a decision support tool for designing Adaptive Experiments

The paper explores how data from fixed A/B tests can guide the deployment of adaptive experiments using contextual bandits. By combining off‑policy evaluation with a controlled warm‑start simulation, the authors rank pre‑specified adaptive and non‑adaptive policies using doubly robust estimators. Experiments on synthetic trials and real benchmarks show that adaptive, context‑aware policies outperform fixed allocations when heterogeneity exists, but offer little advantage otherwise.

By Jo\~ao Victor Ferreira Alves, Eduardo Rocha Laurentino, Gustavo de Oliveira Kanno, Thiago Costa Rizuti da Rocha