arXiv Machine Learning By Woojin Chae, Ezinne Nwankwo, Haitong Qin, Angela Zhou

Optimal Sequential Annotations for Off-Policy Evaluation

Read the original on arXiv Machine Learning →

The paper proposes a method for allocating a limited budget of expert annotations to optimize the accuracy of off-policy evaluation in settings where rewards are missing or noisy. By deriving variance‑optimal annotation probabilities for sequential, forward‑monotone protocols, the authors provide a batch‑adaptive implementation that can be applied to real data. Experiments on casenotes from a homelessness services nonprofit and on human‑preference votes from LMArena demonstrate substantial reductions in RMSE—up to 65% for housing placement and 68% for progress toward a housing application—when using only 40% or more of the full annotation budget.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
5d ago

Offline Policy Evaluation as a decision support tool for designing Adaptive Experiments

The paper explores how data from fixed A/B tests can guide the deployment of adaptive experiments using contextual bandits. By combining off‑policy evaluation with a controlled warm‑start simulation, the authors rank pre‑specified adaptive and non‑adaptive policies using doubly robust estimators. Experiments on synthetic trials and real benchmarks show that adaptive, context‑aware policies outperform fixed allocations when heterogeneity exists, but offer little advantage otherwise.

By Jo\~ao Victor Ferreira Alves, Eduardo Rocha Laurentino, Gustavo de Oliveira Kanno, Thiago Costa Rizuti da Rocha
arXiv AI
Jun 4

Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots

arXiv:2606. 04503v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has greatly advanced large reasoning models (LRMs), but it requires timely training on a huge fully-annotated dataset.

By Guangcheng Zhu, Shenzhi Yang, Haobo Wang, Xing Zheng, Yingfan MA, Xuening Feng, Zhongqi Chen, Bowen Song, Weiqiang Wang, Gang Chen