arXiv:2606. 07726v1 Announce Type: new Abstract: Large Language Models are typically benchmarked by evaluating every model on every test query.
By Zifan Lyu, Chahine Nejma, Tobias Wegel, Fanny Yang, Florian E. Dorner
arXiv:2606. 05308v1 Announce Type: new Abstract: With PRECISE, we extended Prediction-Powered Inference to produce bias-corrected estimates of ranking evaluation metrics by combining a small human-labeled set with a large LLM-judged set.
By Abhishek Divekar
arXiv:2603. 09692v2 Announce Type: replace-cross Abstract: Reinforcement Learning from Human Feedback (RLHF) has become the standard for aligning Large Language Models (LLMs), yet its efficacy is bottlenecked by the high cost of acquiring preference data, especially in low-resource and expert domains.
By Davit Melikidze, Marian Schneider, Jessica Lam, Martin Wertich, Ido Hakimi, Barna P\'asztor, Andreas Krause
arXiv:2606. 00846v1 Announce Type: new Abstract: Users increasingly face the challenge of selecting an appropriate LLM for a given task from a rapidly growing pool of LLMs, each with distinct but often opaque latent properties.
By Son Nguyen, Xinyuan Liu, Ransalu Senanayake
arXiv:2606. 13221v2 Announce Type: replace Abstract: Evaluating new large language models typically requires costly human annotation campaigns at scale.
By Bora Kargi, David Salinas
arXiv:2605. 01961v2 Announce Type: replace Abstract: Learning from human preference data is becoming a useful tool, from fine-tuning large language models to training reinforcement learning agents.
By Maheed H. Ahmed, Mahsa Ghasemi