arXiv AI

Risk-Aware Adaptive Evaluation: Finding High-Impact Failures Under Limited Budgets

arXiv Machine Learning
3d ago

Offline Policy Evaluation as a decision support tool for designing Adaptive Experiments

The paper explores how data from fixed A/B tests can guide the deployment of adaptive experiments using contextual bandits. By combining off‑policy evaluation with a controlled warm‑start simulation, the authors rank pre‑specified adaptive and non‑adaptive policies using doubly robust estimators. Experiments on synthetic trials and real benchmarks show that adaptive, context‑aware policies outperform fixed allocations when heterogeneity exists, but offer little advantage otherwise.

By Jo\~ao Victor Ferreira Alves, Eduardo Rocha Laurentino, Gustavo de Oliveira Kanno, Thiago Costa Rizuti da Rocha
arXiv Machine Learning
2d ago

Tail-Influence Sampling for CVaR Policy Evaluation

arXiv:2609.38096v1 Announce Type: new Abstract: Policies with similar mean returns can differ sharply in rare failures, yet estimating lower-tail conditional value-at-risk (CVaR) accurately can requi...

By Pauline Bourigault, Xiaotong Ji, Matthieu Zimmer, Rasul Tutunov, Haitham Bou-Ammar
arXiv Computation and Language
Sep 15

When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis

arXiv:2609.15309v1 Announce Type: new Abstract: Large language model (LLM) agents allocate test-time compute adaptively as they revise solutions, use tools, explore alternatives, and decide when to s...

By Kaiyuan Liu, Qiuyang Mang, Bo Peng, Wenhao Chai, Hanchen Li, Shreyas Pimpalgaonkar, Luke Zettlemoyer, Alex Dimakis, Alvin Cheung
arXiv Machine Learning
Jul 30

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR

arXiv:2607. 26253v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) is bottlenecked by rollout generation, yet many sampled prompts produce saturated groups (all responses correct or all incorrect) whose zero reward variance yields no policy-gradient signal.

By Pixel Nomand, Elena Voss, Marcus Hale, Sofia Reyes
arXiv AI
Sep 25

A General Framework for Budgeted Threshold Incentives on Request

The paper introduces a request-driven framework for designing budgeted threshold incentives on on-demand delivery platforms. It decomposes the process into four stages—conditional prediction, population reduction, trajectory integration, and budget allocation—using seven interchangeable modules that share conditional trajectory laws. The framework includes a response-correction step that reweights abundant no-offer data to match short pilot moments, and the authors prove that the end-to-end value loss is bounded by the sum of stage errors, with empirical results showing significant speedups and reduced regret compared to traditional trials.

By Zhuolin Wu, Chengrui Zhu, Wenhua Nie, Kenny Ye Liang, Junming Lin, Haiyang Li, Zhilin Li, Wenjia Geng, Zeyu Wu, Yinan Wu, Jinghua Hao, Renqing He
arXiv AI
Sep 4

Learning What Not to Forget: Long-Horizon Agent Memory from a Few Kilobytes of Learning

The paper introduces LRE (Learned Relevance Eviction), a lightweight, CPU‑only, language‑model‑free scorer that learns which parts of an agent’s interaction history are task‑critical and preserves them verbatim. In experiments, LRE matches or surpasses baseline eviction policies on accuracy‑cost trade‑offs, recovers 93% of full‑history accuracy, reduces worst‑case prompt size by 52%, and outperforms dense and token‑pruning encoders in conversational memory while being 295–1569× smaller. The method also achieves superior budgeted answer quality on LoCoMo reading and can be trained annotation‑free, recovering 95% of supervised scorer performance.

By Nusrat Jahan Lia, Aritra Mazumder
arXiv Machine Learning
Aug 31

Budget-Constrained Causal Bandits: Bridging Uplift Modeling and Sequential Decision-Making

The paper introduces Budget-Constrained Causal Bandits (BCCB), an online framework that learns individual treatment effects, explores uncertain users, and manages budget pacing simultaneously. It derives a per-arrival decision rule from a KKT condition of a Lagrangian relaxation, providing a principled algorithmic foundation. Experiments on the Criteo Uplift dataset show BCCB outperforms offline pipelines and other online baselines, especially when historical data is scarce (below 7,500 observations).

By Abhirami Pillai