Life After Benchmark Saturation: A Case Study of CORE-Bench
arXiv:2606. 26158v1 Announce Type: new Abstract: When a benchmark's accuracy saturates, it is often retired and replaced with a more challenging version.
arXiv:2510. 15236v2 Announce Type: replace Abstract: Contemporary AGI evaluations report multidomain capability profiles, yet they typically assign symmetric weights and rely on snapshot scores.
arXiv:2606. 26158v1 Announce Type: new Abstract: When a benchmark's accuracy saturates, it is often retired and replaced with a more challenging version.
The paper introduces Balance of Benchmarks (BoB), a framework that improves task-conditioned model comparison by weighting benchmark evidence based on semantic density, equating scores across varying difficulty levels, and pooling task-relevant residuals. BoB retains all eligible benchmark data while adjusting its influence, outperforming uniform averaging on the WildScores dataset with higher Spearman correlation, lower MAE, and better shortlist hit rates. The method also reduces ranking instability when benchmarks are repeated or paraphrased, and lowers retrospective regret in model selection.
The paper investigates how making responsible‑AI evaluations more efficient—through batching, quantization, and benchmark reduction—affects the stability of conclusions drawn about model behavior. By testing three dense and mixture‑of‑experts models on the BBQ and BBQ‑V datasets under seven different conditions, the authors compare accuracy, bias, reasoning quality, subgroup performance, subset‑membership stability, runtime, and GPU energy consumption against a full‑benchmark BF16 baseline. Findings show that larger batching preserves accuracy and reduces energy in most settings, INT8 largely maintains quality but can increase energy use, INT4 introduces larger, context‑dependent changes, and reduced benchmarks save resources but are highly sensitive to which items are retained, underscoring that efficient evaluation must be validated against the benchmark’s intended conclusions.
arXiv:2607. 12790v1 Announce Type: new Abstract: Self-evolving agent systems improve by creating, revising, and retiring their own skills, but every such loop rests on a hidden assumption: a reliable evaluation metric already exists.
arXiv:2608.30568v1 Announce Type: cross Abstract: Background: Population level assessments of predictive artificial intelligence (AI) can conceal performance disparities across subgroups. Fairness ev...
arXiv:2606. 30655v1 Announce Type: cross Abstract: AI-native course assessments in senior computer science courses and related fields should grade students by \emph{AI-resilient skill}: the ability to achieve outcomes beyond a strong AI baseline.
arXiv:2607. 11969v1 Announce Type: cross Abstract: Point-adjustment (PA), long the default scoring protocol in time-series anomaly detection (TSAD), was shown by Kim et al.
arXiv:2603. 00077v3 Announce Type: replace-cross Abstract: Rubric-based LLM judges have become indispensable for evaluating and optimizing systems on non-verifiable tasks, where success cannot be reduced to exact programmatic checks.
arXiv:2606. 01314v1 Announce Type: new Abstract: Recent self-evolving agents have shown that skills can be discovered, refined, and accumulated through execution.
arXiv:2606. 09878v1 Announce Type: new Abstract: Standard benchmarks report aggregate accuracy, but practitioners need to know which specific capabilities a model lacks.
Efficient evaluation changes the protocol used to support claims about model behavior, yet it is rarely tested whether those claims remain stable after the evaluation itself is made cheaper. We stress...
Self-evolving agent systems improve by creating, revising, and retiring their own skills, but every such loop rests on a hidden assumption: a reliable evaluation metric already exists. In many real applications it does not.