arXiv AI By YongKyung Oh

Position: State-of-the-Art Claims Require State-of-the-Art Evidence

Read the original on arXiv AI →

arXiv:2605. 17273v3 Announce Type: replace-cross Abstract: State-of-the-Art (SOTA) claims pervade Artificial Intelligence (AI) and Machine Learning (ML) research.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 20

Lost in Aggregation: How Benchmarks Overlook Irreplaceable Model Strengths

The paper argues that typical tabular machine learning benchmarks, which aggregate results by averaging scores or ranks, can hide which models are essential for achieving the best performance on specific datasets. It proposes evaluating models against a data‑centric peak performance frontier, classifying them as irreplaceable, sufficient, redundant, or fallible based on their position relative to other models. Applying this to the TabArena benchmark shows that common aggregation metrics mainly capture consistency and failure avoidance, but fail to reflect dataset‑specific strengths, leading to a misalignment between aggregate rewards and true model utility.

By Andrej Tschalzev, Stefan L\"udtke, Heiner Stuckenschmidt, Christian Bartelt
arXiv AI
Sep 21

Balance of Benchmarks: Semantic Density Reweighting for Task-Conditioned Model Comparison

The paper introduces Balance of Benchmarks (BoB), a framework that improves task-conditioned model comparison by weighting benchmark evidence based on semantic density, equating scores across varying difficulty levels, and pooling task-relevant residuals. BoB retains all eligible benchmark data while adjusting its influence, outperforming uniform averaging on the WildScores dataset with higher Spearman correlation, lower MAE, and better shortlist hit rates. The method also reduces ranking instability when benchmarks are repeated or paraphrased, and lowers retrospective regret in model selection.

By Jhen-Ke Lin, Hong-Yun Lin