The paper introduces a new variant of the multi‑armed bandit problem that incorporates dueling feedback—pairwise comparisons of model responses—and heterogeneous sampling costs to identify the best large language model (LLM) from a set with varying query costs. Assuming a Condorcet winner, the authors propose a Track‑and‑Stop style algorithm that guarantees asymptotically optimal cost as the error probability approaches zero. Extensive experiments on synthetic and real‑world data show that this cost‑aware approach consistently outperforms both classical cost‑unaware algorithms and other cost‑aware extensions.
By Sarvesh Gharat, Nikhil Karamchandani, Jayakrishnan Nair
arXiv:2606. 07726v1 Announce Type: new Abstract: Large Language Models are typically benchmarked by evaluating every model on every test query.
By Zifan Lyu, Chahine Nejma, Tobias Wegel, Fanny Yang, Florian E. Dorner
arXiv:2606. 05308v1 Announce Type: new Abstract: With PRECISE, we extended Prediction-Powered Inference to produce bias-corrected estimates of ranking evaluation metrics by combining a small human-labeled set with a large LLM-judged set.
By Abhishek Divekar
arXiv:2603.21716v2 Announce Type: replace-cross
Abstract: Efficient selection among multiple generative models is increasingly important in modern generative AI, where sampling from suboptimal models...
By Bahar Dibaei Nia, Farzan Farnia
arXiv:2609.06670v1 Announce Type: new
Abstract: The In-context learning (ICL) paradigm aids large language models (LLMs) to adapt to new tasks without need for fine-tuning. However, selecting an opti...
By V Venktesh, Cem levi, Avishek Anand
arXiv:2603. 09692v2 Announce Type: replace-cross Abstract: Reinforcement Learning from Human Feedback (RLHF) has become the standard for aligning Large Language Models (LLMs), yet its efficacy is bottlenecked by the high cost of acquiring preference data, especially in low-resource and expert domains.
By Davit Melikidze, Marian Schneider, Jessica Lam, Martin Wertich, Ido Hakimi, Barna P\'asztor, Andreas Krause