arXiv:2608.25940v2 Announce Type: replace-cross
Abstract: Physical AI models are evaluated on suites of benchmarks that differ across model reports, leaving the model-by-benchmark matrix sparse and t...
By Zaruhi Navasardyan, Hrant Davtyan
The paper argues that typical tabular machine learning benchmarks, which aggregate results by averaging scores or ranks, can hide which models are essential for achieving the best performance on specific datasets. It proposes evaluating models against a data‑centric peak performance frontier, classifying them as irreplaceable, sufficient, redundant, or fallible based on their position relative to other models. Applying this to the TabArena benchmark shows that common aggregation metrics mainly capture consistency and failure avoidance, but fail to reflect dataset‑specific strengths, leading to a misalignment between aggregate rewards and true model utility.
By Andrej Tschalzev, Stefan L\"udtke, Heiner Stuckenschmidt, Christian Bartelt
arXiv:2609.23201v1 Announce Type: new
Abstract: Benchmark scores increasingly influence the development, marketing, and selection of large language models (LLMs). Yet an overall score is interpretabl...
By Danial Amin
Tabular machine learning benchmarks typically summarize performance by averaging scores, ranks, or pairwise wins across datasets. Such aggregates are useful for selecting robust default models, but th...
The paper introduces Balance of Benchmarks (BoB), a framework that improves task-conditioned model comparison by weighting benchmark evidence based on semantic density, equating scores across varying difficulty levels, and pooling task-relevant residuals. BoB retains all eligible benchmark data while adjusting its influence, outperforming uniform averaging on the WildScores dataset with higher Spearman correlation, lower MAE, and better shortlist hit rates. The method also reduces ranking instability when benchmarks are repeated or paraphrased, and lowers retrospective regret in model selection.
By Jhen-Ke Lin, Hong-Yun Lin
Physical AI models are evaluated on suites of benchmarks that differ across model reports, leaving the model-by-benchmark matrix sparse and the relationship between benchmarks unmeasured. We construct...
arXiv:2607. 16239v1 Announce Type: new Abstract: AI judges offer a scalable, low-cost alternative to human evaluation, but their outputs can be biased relative to human preferences and highly item-dependent, varying across judges, tasks, and domains.
By Lei Shi, Anlan Zhang, Rita Lyu, Zhengmian Hu, Tong Yu, David Arbour, Avi Feller, Saayan Mitra, Ritwik Sinha
arXiv:2609.07785v1 Announce Type: new
Abstract: An LLM-agent leaderboard invites a familiar inference: an agent ranked above another is the better agent. Public evaluation logs may not support that c...
By Wei-Jung Huang
arXiv:2607. 15190v1 Announce Type: new Abstract: AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities, rank systems, select informative examples, and diagnose benchmark quality.
By Han Jiang, Sunbeom Kwon, Jinwen Luo, Ziang Xiao, Susu Zhang
arXiv:2602. 14307v4 Announce Type: replace Abstract: As frontier Large Language Models (LLMs) increasingly saturate new benchmarks shortly after they are published, benchmarking itself is at a juncture: if frontier models keep improving, it will become increasingly hard for humans to generate discriminative tasks, provide accurate ground-truth answers, or evaluate complex solutions.
By Samuele Marro, Jialin Yu, Emanuele La Malfa, Oishi Deb, Jiawei Li, Yibo Yang, Ebey Abraham, Sunando Sengupta, Eric Sommerlade, Michael Wooldridge, Philip Torr
arXiv:2607. 28801v1 Announce Type: cross Abstract: Benchmark datasets are central to evaluating Large Language Models (LLMs), yet they are typically conceived as monolithic tasks, obscuring substantial variation in the demands of individual samples.
By Philipp D. Siedler, Jordan Sassoon
arXiv:2606. 27997v1 Announce Type: new Abstract: Benchmarks of machine learning models often include many datasets, making evaluation expensive.
By Rostislav Gusev, Alexey Zaytsev