The Benchmarking Epistemology: Validity Theory for Evaluating Machine Learning Models
Read the original on arXiv Statistics ML →The article discusses how predictive benchmarking—evaluating machine learning models by their performance and ranking—serves as a core method in machine learning research. It argues that benchmark scores only reflect performance on specific datasets and learning problems, and that drawing broader scientific conclusions requires explicit assumptions. By adapting concepts from psychological validity theory, the authors propose validity conditions to make these assumptions clear, and demonstrate their application in two case studies (ImageNet and the Fragile Families Challenge) to illustrate how benchmark results can inform inferences about research progress and limits of predictability.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Statistics ML.