arXiv Machine Learning

Best-Arm Identification with Generative Proxy

arXiv:2607. 06879v1 Announce Type: new Abstract: Best-arm identification is a canonical model for data-driven decision-making, but in many applications each reward observation is costly.

arXiv AI
6d ago

Cost-Aware Best-LLM Identification using Dueling Feedback

The paper introduces a new variant of the multi‑armed bandit problem that incorporates dueling feedback—pairwise comparisons of model responses—and heterogeneous sampling costs to identify the best large language model (LLM) from a set with varying query costs. Assuming a Condorcet winner, the authors propose a Track‑and‑Stop style algorithm that guarantees asymptotically optimal cost as the error probability approaches zero. Extensive experiments on synthetic and real‑world data show that this cost‑aware approach consistently outperforms both classical cost‑unaware algorithms and other cost‑aware extensions.

By Sarvesh Gharat, Nikhil Karamchandani, Jayakrishnan Nair
arXiv Machine Learning
Sep 14

Satisficing Regret Minimization in Bandits: Constant Rate and Light-Tailed Distribution

The paper introduces SELECT, an algorithmic framework for satisficing regret minimization in bandit problems, achieving constant expected satisficing regret when a satisficing arm exists. A variant, SELECT‑LITE, further ensures a light‑tailed satisficing regret distribution while maintaining constant expected regret in the realizable case and sub‑linear standard regret otherwise. Experiments on synthetic data and a real‑world dynamic pricing scenario demonstrate the practical effectiveness of both algorithms.

By Qing Feng, Tianyi Ma, Ruihao Zhu
arXiv AI
3d ago

Revisiting scaling laws for reward optimization

The paper presents a new scaling law for reward optimization in AI alignment, showing that performance scales as Θ(√min{log(M), K}), where M is the number of preference comparisons used to train a proxy reward model and K is the KL‑divergence budget relative to a reference policy. The authors derive this law using an information‑theoretic model, prove its tightness, and validate it with extensive experiments involving a 70B gold reward model and smaller proxy models (0.6B–4B). The empirical results demonstrate a strong fit (R² 97–99 %) across different model sizes, noise levels, and optimization methods, suggesting that reward optimization behaves like a simple selection task over IID Gaussian variables with noisy feedback.

By Ali Aouad, Aymane El Gadarri, Vivek F. Farias
arXiv Machine Learning
Jun 15

A Complexity Measure for Active Learning in Multi-group Mean Estimation

arXiv:2606. 14690v1 Announce Type: new Abstract: We study a \emph{max-risk} objective for active learning in a multi-group mean estimation $d$-armed bandits: a learner adaptively allocates a budget of $T$ samples across $d$ groups to minimize the worst-case uncertainty index $\max_{k\in[d]}\sigma_k^2/n_k$, where $\sigma_k$ is the standard deviation of the distribution of arm $d$, and $n_k$ is the number of times arm $d$ is sampled.

By Abdellah Aznag, Rachel Cummings, Adam N. Elmachtoub
Hugging Face Trending Papers
Jul 16

MESHA: Mechanism-Enforced Sequential Halving for Strategic Linear Bandits

We design and analyze \underline{M}echanism-\underline{E}nforced \underline{S}equential \underline{HA}lving (MESHA), an algorithm for Best Arm Identification (BAI) in strategic linear bandits. In this setting, each arm may strategically misreport its feature vector to maximize the probability of being identified as the best arm, when rewards are generated from the arms' true but unobservable features.