arXiv:2506. 07673v2 Announce Type: replace Abstract: Large language model (LLM) evaluation is increasingly costly, prompting interest in methods that speed up evaluation by shrinking benchmark datasets.
By Guanhua Zhang, Florian E. Dorner, Moritz Hardt
arXiv:2607. 09739v1 Announce Type: new Abstract: We study LLM benchmark coreset selection: selecting a small subset of prompts over multiple benchmarks whose induced model scores and rankings approximate those obtained from the full benchmark suite.
By Jihan Yao, Gantavya Bhatt, Arnav Das, Peter Jin, Ke Bao, Qiaolin Yu, Khushi Bhardwaj, Chang Su, Jialei Wang, Yikai Zhu, Sugam Devare, Damon Mosk-Aoyama, Zhen Dong, Venkat Krishna Srinivasan, Yineng Zhang, Oleksii Kuchaiev, Jiantao Jiao, Banghua Zhu, Jeff Bilmes
arXiv:2605.24981v2 Announce Type: replace
Abstract: Choosing a Large Language Model (LLM) for a given task requires comparing many strong candidates, yet standard evaluation relies on costly annotati...
By Yavuz Durmazkeser, Patrik Okanovic, Andreas Kirsch, Torsten Hoefler, Nezihe Merve G\"urel
The paper introduces a low‑rank framework for ranking large language models (LLMs) on task‑specific benchmarks using sparse pairwise comparisons. By modeling the task‑by‑model ability matrix as low rank, the method shares information across related tasks while preserving task‑specific differences, and it provides uncertainty‑aware ranking through debiased estimators and simultaneous confidence sets. Experiments on synthetic data and the Chatbot Arena benchmark demonstrate improved sample efficiency and tighter, better‑calibrated ranking certificates, especially in the sparse comparison regime typical of real LLM evaluations.
By Jiachun Li, David Simchi-Levi, Will Wei Sun
arXiv:2607. 04581v2 Announce Type: replace-cross Abstract: Text embeddings for Portuguese have no dedicated benchmark: evaluation rests on translated corpora such as English MS MARCO or on thin multilingual coverage, with native tasks scattered and unconsolidated.
By Tardelli Ronan Coelho Stekel
arXiv:2607. 04581v1 Announce Type: cross Abstract: Text embeddings for Portuguese have no dedicated benchmark: evaluation rests on translated corpora such as English MS MARCO or on thin multilingual coverage, with native tasks scattered and unconsolidated.
By Tardelli Ronan Coelho Stekel
arXiv:2609.06670v1 Announce Type: new
Abstract: The In-context learning (ICL) paradigm aids large language models (LLMs) to adapt to new tasks without need for fine-tuning. However, selecting an opti...
By V Venktesh, Cem levi, Avishek Anand
The paper argues that typical tabular machine learning benchmarks, which aggregate results by averaging scores or ranks, can hide which models are essential for achieving the best performance on specific datasets. It proposes evaluating models against a data‑centric peak performance frontier, classifying them as irreplaceable, sufficient, redundant, or fallible based on their position relative to other models. Applying this to the TabArena benchmark shows that common aggregation metrics mainly capture consistency and failure avoidance, but fail to reflect dataset‑specific strengths, leading to a misalignment between aggregate rewards and true model utility.
By Andrej Tschalzev, Stefan L\"udtke, Heiner Stuckenschmidt, Christian Bartelt
arXiv:2606. 07492v1 Announce Type: cross Abstract: The ranking of recommendation algorithms is a challenging problem since model performance is sensitive to dataset characteristics such as sparsity, sequential structure, and scale.
By Ekaterina Grishina, Stepan Kuznetsov, Askar Tsyganov, Ilya Ivanov, Daria Korovaitceva, Margarita Rusanova, Uliana Parkina, Alexander Derevyagin, Evgeny Frolov, Sergey Samsonov, Anton Lysenko
The paper investigates the "score granularity gap" in black-box large language model (LLM) classifiers, asking how finely a confidence score can be thresholded for deployment. By comparing seven confidence construction methods across 25 model-dataset pairs, the authors find that single-shot verbalized confidence, when properly converted to a probability, ranks well but offers only a few distinct threshold values, limiting operational flexibility. The study also shows that multi-query aggregation can improve weak models but may harm strong ones, and provides concrete guidance for deployment trade-offs.
By Ao Sun, Tian Sun, Jiaxing Geng
arXiv:2510. 06048v4 Announce Type: replace Abstract: Effective data selection is essential for pretraining large language models (LLMs), enhancing efficiency and improving generalization to downstream tasks.
By Jie Hao, Rui Yu, Wei Zhang, Huixia Wang, Jie Xu, Mingrui Liu
Tabular machine learning benchmarks typically summarize performance by averaging scores, ranks, or pairwise wins across datasets. Such aggregates are useful for selecting robust default models, but th...