arXiv Machine Learning By Zifan Lyu, Chahine Nejma, Tobias Wegel, Fanny Yang, Florian E. Dorner

Cutting LLM Evaluation Costs with SySRs: A Bandit Algorithm that Provably Exploits Model Similarity

Read the original on arXiv Machine Learning →

arXiv:2606. 07726v1 Announce Type: new Abstract: Large Language Models are typically benchmarked by evaluating every model on every test query.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
6d ago

Cost-Aware Best-LLM Identification using Dueling Feedback

The paper introduces a new variant of the multi‑armed bandit problem that incorporates dueling feedback—pairwise comparisons of model responses—and heterogeneous sampling costs to identify the best large language model (LLM) from a set with varying query costs. Assuming a Condorcet winner, the authors propose a Track‑and‑Stop style algorithm that guarantees asymptotically optimal cost as the error probability approaches zero. Extensive experiments on synthetic and real‑world data show that this cost‑aware approach consistently outperforms both classical cost‑unaware algorithms and other cost‑aware extensions.

By Sarvesh Gharat, Nikhil Karamchandani, Jayakrishnan Nair
arXiv Computation and Language
6d ago

Large Language Model Selection with Limited Annotations

arXiv:2605.24981v2 Announce Type: replace Abstract: Choosing a Large Language Model (LLM) for a given task requires comparing many strong candidates, yet standard evaluation relies on costly annotati...

By Yavuz Durmazkeser, Patrik Okanovic, Andreas Kirsch, Torsten Hoefler, Nezihe Merve G\"urel