arXiv Machine Learning
Sep 1

Stress-Testing Efficient Responsible-AI Evaluation: When Compute Savings Change Benchmark Conclusions

The paper investigates how making responsible‑AI evaluations more efficient—through batching, quantization, and benchmark reduction—affects the stability of conclusions drawn about model behavior. By testing three dense and mixture‑of‑experts models on the BBQ and BBQ‑V datasets under seven different conditions, the authors compare accuracy, bias, reasoning quality, subgroup performance, subset‑membership stability, runtime, and GPU energy consumption against a full‑benchmark BF16 baseline. Findings show that larger batching preserves accuracy and reduces energy in most settings, INT8 largely maintains quality but can increase energy use, INT4 introduces larger, context‑dependent changes, and reduced benchmarks save resources but are highly sensitive to which items are retained, underscoring that efficient evaluation must be validated against the benchmark’s intended conclusions.

By Ahmed El Kady, Aravind Narayanan, Rehana Noorani, Yani Ioannou, Shaina Raza
arXiv AI
Sep 21

Balance of Benchmarks: Semantic Density Reweighting for Task-Conditioned Model Comparison

The paper introduces Balance of Benchmarks (BoB), a framework that improves task-conditioned model comparison by weighting benchmark evidence based on semantic density, equating scores across varying difficulty levels, and pooling task-relevant residuals. BoB retains all eligible benchmark data while adjusting its influence, outperforming uniform averaging on the WildScores dataset with higher Spearman correlation, lower MAE, and better shortlist hit rates. The method also reduces ranking instability when benchmarks are repeated or paraphrased, and lowers retrospective regret in model selection.

By Jhen-Ke Lin, Hong-Yun Lin
arXiv AI
Aug 26

Evaluating Multiple LLM Generations with Validated Task Coverage

The paper introduces VTC-Bench, a five‑domain benchmark designed to evaluate multiple outputs from large language models (LLMs) by measuring Validated Task Coverage (VTC). VTC quantifies how many distinct, useful results are produced within a set number of attempts, using real‑data tasks that allow automatic, reproducible checks of output quality and task‑relevant distinctness without relying on model‑based judges. Experiments show that models which perform best on single‑draw quality do not always achieve the highest coverage, and simple output‑variation metrics fail to capture task‑relevant diversity, highlighting the importance of evaluating finite candidate sets directly.

By Florian Le Bronnec, Rio Yokota