arXiv Machine Learning By Ofir Arviv, Kristjan Greenewald, Yotam Perlitz, Hadar Mulian, Michal Shmueli-Scheuer, Leshem Choshen

Stop Guessing When to Stop Testing: Efficient Model Evaluation with Just Enough Data

Read the original on arXiv Machine Learning →

arXiv:2607. 08522v1 Announce Type: new Abstract: The inherent rigidity of fixed-size benchmarks makes them an inefficient tool for model evaluation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 26

Evaluating Multiple LLM Generations with Validated Task Coverage

The paper introduces VTC-Bench, a five‑domain benchmark designed to evaluate multiple outputs from large language models (LLMs) by measuring Validated Task Coverage (VTC). VTC quantifies how many distinct, useful results are produced within a set number of attempts, using real‑data tasks that allow automatic, reproducible checks of output quality and task‑relevant distinctness without relying on model‑based judges. Experiments show that models which perform best on single‑draw quality do not always achieve the highest coverage, and simple output‑variation metrics fail to capture task‑relevant diversity, highlighting the importance of evaluating finite candidate sets directly.

By Florian Le Bronnec, Rio Yokota
arXiv Machine Learning
Jun 24

You Don't Need to Run Every Eval

arXiv:2606. 24020v1 Announce Type: new Abstract: A modern model release reports scores on 40+ benchmarks and the same evaluations were run many more times before it: to track training progress, compare design choices, and select the checkpoint for the release.

By Yuchen Zeng, Dimitris Papailiopoulos