Stress-Testing Efficient Responsible-AI Evaluation: When Compute Savings Change Benchmark Conclusions
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The paper investigates how making responsible‑AI evaluations more efficient—through batching, quantization, and benchmark reduction—affects the stability of conclusions drawn about model behavior. By testing three dense and mixture‑of‑experts models on the BBQ and BBQ‑V datasets under seven different conditions, the authors compare accuracy, bias, reasoning quality, subgroup performance, subset‑membership stability, runtime, and GPU energy consumption against a full‑benchmark BF16 baseline. Findings show that larger batching preserves accuracy and reduces energy in most settings, INT8 largely maintains quality but can increase energy use, INT4 introduces larger, context‑dependent changes, and reduced benchmarks save resources but are highly sensitive to which items are retained, underscoring that efficient evaluation must be validated against the benchmark’s intended conclusions.
The paper introduces Balance of Benchmarks (BoB), a framework that improves task-conditioned model comparison by weighting benchmark evidence based on semantic density, equating scores across varying difficulty levels, and pooling task-relevant residuals. BoB retains all eligible benchmark data while adjusting its influence, outperforming uniform averaging on the WildScores dataset with higher Spearman correlation, lower MAE, and better shortlist hit rates. The method also reduces ranking instability when benchmarks are repeated or paraphrased, and lowers retrospective regret in model selection.
arXiv:2606.22826v2 Announce Type: replace Abstract: Evaluating LLMs across many model variants---quantized, fine-tuned, or deployment-specific---requires running large benchmarks repeatedly, a proces...
The paper introduces VTC-Bench, a five‑domain benchmark designed to evaluate multiple outputs from large language models (LLMs) by measuring Validated Task Coverage (VTC). VTC quantifies how many distinct, useful results are produced within a set number of attempts, using real‑data tasks that allow automatic, reproducible checks of output quality and task‑relevant distinctness without relying on model‑based judges. Experiments show that models which perform best on single‑draw quality do not always achieve the highest coverage, and simple output‑variation metrics fail to capture task‑relevant diversity, highlighting the importance of evaluating finite candidate sets directly.
arXiv:2606. 03938v1 Announce Type: cross Abstract: Multi-epoch training is becoming the standard now that compute is growing faster than the supply of high-quality text.
arXiv:2609.36222v1 Announce Type: new Abstract: Large language models are increasingly expensive to serve. In large-scale serving systems, autoregressive decoding is often bottlenecked by transferrin...