Optimal Sequential Annotations for Off-Policy Evaluation
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The paper proposes a method for allocating a limited budget of expert annotations to optimize the accuracy of off-policy evaluation in settings where rewards are missing or noisy. By deriving variance‑optimal annotation probabilities for sequential, forward‑monotone protocols, the authors provide a batch‑adaptive implementation that can be applied to real data. Experiments on casenotes from a homelessness services nonprofit and on human‑preference votes from LMArena demonstrate substantial reductions in RMSE—up to 65% for housing placement and 68% for progress toward a housing application—when using only 40% or more of the full annotation budget.
arXiv:2502. 10605v4 Announce Type: replace-cross Abstract: Problem definition: Estimating causal effects of interventions is central to policy and operations, but outcome data are often missing or costly to obtain.
arXiv:2609.38860v1 Announce Type: cross Abstract: Learning from human preferences is central to large language model (LLM) alignment, but human preference annotation is costly. Active preference lear...
The paper explores how data from fixed A/B tests can guide the deployment of adaptive experiments using contextual bandits. By combining off‑policy evaluation with a controlled warm‑start simulation, the authors rank pre‑specified adaptive and non‑adaptive policies using doubly robust estimators. Experiments on synthetic trials and real benchmarks show that adaptive, context‑aware policies outperform fixed allocations when heterogeneity exists, but offer little advantage otherwise.
MInTRL (Minimal Intervention Reinforcement Learning) expands exploration in on-policy reinforcement learning by inserting sparse, local corrections into rollouts via a judge-intervention policy. These interventions replace erroneous suffixes and immediately return control to the main policy, allowing the agent to explore beyond its natural trajectory while maintaining on-policy data. The method uses a sequence-level advantage-regression objective, avoiding importance sampling, and demonstrates superior performance on math and code benchmarks compared to standard on-policy and off-policy baselines.
arXiv:2606. 20206v1 Announce Type: cross Abstract: In offline Reinforcement Learning, immediate rewards in logged batch data are often unobserved due to sparse or irregular record-keeping, or censored beyond certain reward values.