The paper introduces a new variant of the multi‑armed bandit problem that incorporates dueling feedback—pairwise comparisons of model responses—and heterogeneous sampling costs to identify the best large language model (LLM) from a set with varying query costs. Assuming a Condorcet winner, the authors propose a Track‑and‑Stop style algorithm that guarantees asymptotically optimal cost as the error probability approaches zero. Extensive experiments on synthetic and real‑world data show that this cost‑aware approach consistently outperforms both classical cost‑unaware algorithms and other cost‑aware extensions.
By Sarvesh Gharat, Nikhil Karamchandani, Jayakrishnan Nair
arXiv:2606. 07726v1 Announce Type: new Abstract: Large Language Models are typically benchmarked by evaluating every model on every test query.
By Zifan Lyu, Chahine Nejma, Tobias Wegel, Fanny Yang, Florian E. Dorner
arXiv:2606. 05308v1 Announce Type: new Abstract: With PRECISE, we extended Prediction-Powered Inference to produce bias-corrected estimates of ranking evaluation metrics by combining a small human-labeled set with a large LLM-judged set.
By Abhishek Divekar
arXiv:2603.21716v2 Announce Type: replace-cross
Abstract: Efficient selection among multiple generative models is increasingly important in modern generative AI, where sampling from suboptimal models...
By Bahar Dibaei Nia, Farzan Farnia
arXiv:2609.06670v1 Announce Type: new
Abstract: The In-context learning (ICL) paradigm aids large language models (LLMs) to adapt to new tasks without need for fine-tuning. However, selecting an opti...
By V Venktesh, Cem levi, Avishek Anand
arXiv:2603. 09692v2 Announce Type: replace-cross Abstract: Reinforcement Learning from Human Feedback (RLHF) has become the standard for aligning Large Language Models (LLMs), yet its efficacy is bottlenecked by the high cost of acquiring preference data, especially in low-resource and expert domains.
By Davit Melikidze, Marian Schneider, Jessica Lam, Martin Wertich, Ido Hakimi, Barna P\'asztor, Andreas Krause
arXiv:2605.24981v2 Announce Type: replace
Abstract: Choosing a Large Language Model (LLM) for a given task requires comparing many strong candidates, yet standard evaluation relies on costly annotati...
By Yavuz Durmazkeser, Patrik Okanovic, Andreas Kirsch, Torsten Hoefler, Nezihe Merve G\"urel
COBRA‑Skills is a new framework that treats skill optimization for large language model agents as a budgeted sequential problem over a dynamically evolving candidate set. It uses contextual‑bandit prioritization to focus evaluations on promising or informative candidates and refines the skill population based on execution feedback. In experiments across six agent benchmarks and three target models, COBRA‑Skills outperforms existing methods, cuts optimization cost by 55–58 % compared to SkillOpt, and requires only 50 unique optimization examples per benchmark.
By Pingchen Lu, Xiangyi Wang, Xiang Li, Jie Mao, Zikun Qu, Junfeng Luo, Yao Shu, Bryan Kian Hsiang Low, Zhongxiang Dai
arXiv:2606. 00846v1 Announce Type: new Abstract: Users increasingly face the challenge of selecting an appropriate LLM for a given task from a rapidly growing pool of LLMs, each with distinct but often opaque latent properties.
By Son Nguyen, Xinyuan Liu, Ransalu Senanayake
arXiv:2606. 13221v2 Announce Type: replace Abstract: Evaluating new large language models typically requires costly human annotation campaigns at scale.
By Bora Kargi, David Salinas
arXiv:2609.38860v1 Announce Type: cross
Abstract: Learning from human preferences is central to large language model (LLM) alignment, but human preference annotation is costly. Active preference lear...
By Zhongman Du, Huiming Zhang, Haodong Zhu, Baochang Zhang
arXiv:2605. 01961v2 Announce Type: replace Abstract: Learning from human preference data is becoming a useful tool, from fine-tuning large language models to training reinforcement learning agents.
By Maheed H. Ahmed, Mahsa Ghasemi