Dynamically Allocating Evaluation Effort for Model Ranking
arXiv:2608. 03437v1 Announce Type: cross Abstract: While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability.
The paper introduces a new variant of the multi‑armed bandit problem that incorporates dueling feedback—pairwise comparisons of model responses—and heterogeneous sampling costs to identify the best large language model (LLM) from a set with varying query costs. Assuming a Condorcet winner, the authors propose a Track‑and‑Stop style algorithm that guarantees asymptotically optimal cost as the error probability approaches zero. Extensive experiments on synthetic and real‑world data show that this cost‑aware approach consistently outperforms both classical cost‑unaware algorithms and other cost‑aware extensions.
arXiv:2608. 03437v1 Announce Type: cross Abstract: While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability.
arXiv:2606. 07726v1 Announce Type: new Abstract: Large Language Models are typically benchmarked by evaluating every model on every test query.
arXiv:2606. 01799v1 Announce Type: new Abstract: We study $N$-armed stochastic dueling bandits under the Condorcet-winner assumption, where three widely adopted objectives are considered: best-arm identification (BAI), weak regret, and strong regret.
arXiv:2409. 18909v2 Announce Type: replace Abstract: Motivated by real-world applications that necessitate responsible experimentation, we introduce the problem of best arm identification (BAI) with minimal regret.
arXiv:2607. 09015v1 Announce Type: cross Abstract: We study contextual bandit problems with correlated arms and access to surrogate reward signals produced by a machine learning model, motivated by applications such as large language model (LLM) routing.
arXiv:2607. 06879v1 Announce Type: new Abstract: Best-arm identification is a canonical model for data-driven decision-making, but in many applications each reward observation is costly.
arXiv:2607. 23765v1 Announce Type: cross Abstract: Large language models (LLMs) achieve impressive performance across multiple domains, but using the most capable model for every query is prohibitive at scale.
arXiv:2604. 05859v2 Announce Type: replace Abstract: We study Contextual Multi-Armed Bandits (CMABs) for non-episodic decision-making problems where the context includes both textual and numerical information (e.
arXiv:2603.21716v2 Announce Type: replace-cross Abstract: Efficient selection among multiple generative models is increasingly important in modern generative AI, where sampling from suboptimal models...
Exhaustively evaluating every candidate LLM configuration on every benchmark item to identify a high-performing one is costly. We formulate configuration selection as a cost-aware Bayesian bandit prob...
arXiv:2606. 00846v1 Announce Type: new Abstract: Users increasingly face the challenge of selecting an appropriate LLM for a given task from a rapidly growing pool of LLMs, each with distinct but often opaque latent properties.
arXiv:2605. 01961v2 Announce Type: replace Abstract: Learning from human preference data is becoming a useful tool, from fine-tuning large language models to training reinforcement learning agents.