Hugging Face Blog

Machine Learning Experts - Lewis Tunstall

arXiv Statistics ML
4d ago

The Benchmarking Epistemology: Validity Theory for Evaluating Machine Learning Models

The article discusses how predictive benchmarking—evaluating machine learning models by their performance and ranking—serves as a core method in machine learning research. It argues that benchmark scores only reflect performance on specific datasets and learning problems, and that drawing broader scientific conclusions requires explicit assumptions. By adapting concepts from psychological validity theory, the authors propose validity conditions to make these assumptions clear, and demonstrate their application in two case studies (ImageNet and the Fragile Families Challenge) to illustrate how benchmark results can inform inferences about research progress and limits of predictability.

By Timo Freiesleben, Sebastian Zezulka