arXiv:2608.25940v2 Announce Type: replace-cross
Abstract: Physical AI models are evaluated on suites of benchmarks that differ across model reports, leaving the model-by-benchmark matrix sparse and t...
By Zaruhi Navasardyan, Hrant Davtyan
The paper argues that typical tabular machine learning benchmarks, which aggregate results by averaging scores or ranks, can hide which models are essential for achieving the best performance on specific datasets. It proposes evaluating models against a data‑centric peak performance frontier, classifying them as irreplaceable, sufficient, redundant, or fallible based on their position relative to other models. Applying this to the TabArena benchmark shows that common aggregation metrics mainly capture consistency and failure avoidance, but fail to reflect dataset‑specific strengths, leading to a misalignment between aggregate rewards and true model utility.
By Andrej Tschalzev, Stefan L\"udtke, Heiner Stuckenschmidt, Christian Bartelt
arXiv:2609.23201v1 Announce Type: new
Abstract: Benchmark scores increasingly influence the development, marketing, and selection of large language models (LLMs). Yet an overall score is interpretabl...
By Danial Amin
Tabular machine learning benchmarks typically summarize performance by averaging scores, ranks, or pairwise wins across datasets. Such aggregates are useful for selecting robust default models, but th...
The paper introduces Balance of Benchmarks (BoB), a framework that improves task-conditioned model comparison by weighting benchmark evidence based on semantic density, equating scores across varying difficulty levels, and pooling task-relevant residuals. BoB retains all eligible benchmark data while adjusting its influence, outperforming uniform averaging on the WildScores dataset with higher Spearman correlation, lower MAE, and better shortlist hit rates. The method also reduces ranking instability when benchmarks are repeated or paraphrased, and lowers retrospective regret in model selection.
By Jhen-Ke Lin, Hong-Yun Lin
Physical AI models are evaluated on suites of benchmarks that differ across model reports, leaving the model-by-benchmark matrix sparse and the relationship between benchmarks unmeasured. We construct...