arXiv Machine Learning By Vil\'em Zouhar, Julia Kreutzer, Alon Lavie, Tom Kocmi, Matt Post, Ond\v{r}ej Bojar, Mrinmaya Sachan

Dynamically Allocating Evaluation Effort for Model Ranking

Read the original on arXiv Machine Learning →

arXiv:2608. 03437v1 Announce Type: cross Abstract: While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
6d ago

Cost-Aware Best-LLM Identification using Dueling Feedback

The paper introduces a new variant of the multi‑armed bandit problem that incorporates dueling feedback—pairwise comparisons of model responses—and heterogeneous sampling costs to identify the best large language model (LLM) from a set with varying query costs. Assuming a Condorcet winner, the authors propose a Track‑and‑Stop style algorithm that guarantees asymptotically optimal cost as the error probability approaches zero. Extensive experiments on synthetic and real‑world data show that this cost‑aware approach consistently outperforms both classical cost‑unaware algorithms and other cost‑aware extensions.

By Sarvesh Gharat, Nikhil Karamchandani, Jayakrishnan Nair
arXiv AI
Jun 2

ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning

arXiv:2603. 09692v2 Announce Type: replace-cross Abstract: Reinforcement Learning from Human Feedback (RLHF) has become the standard for aligning Large Language Models (LLMs), yet its efficacy is bottlenecked by the high cost of acquiring preference data, especially in low-resource and expert domains.

By Davit Melikidze, Marian Schneider, Jessica Lam, Martin Wertich, Ido Hakimi, Barna P\'asztor, Andreas Krause