arXiv Statistics ML By Timo Freiesleben, Sebastian Zezulka

The Benchmarking Epistemology: Validity Theory for Evaluating Machine Learning Models

Read the original on arXiv Statistics ML →

The article discusses how predictive benchmarking—evaluating machine learning models by their performance and ranking—serves as a core method in machine learning research. It argues that benchmark scores only reflect performance on specific datasets and learning problems, and that drawing broader scientific conclusions requires explicit assumptions. By adapting concepts from psychological validity theory, the authors propose validity conditions to make these assumptions clear, and demonstrate their application in two case studies (ImageNet and the Fragile Families Challenge) to illustrate how benchmark results can inform inferences about research progress and limits of predictability.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Statistics ML.

arXiv AI
Aug 25

The Measurement Revolution? Credible Measurement and Inference in the Age of AI

The article discusses how artificial intelligence is reshaping measurement in economics by converting unstructured data into structured variables at low cost, enabling large‑scale measurement that was previously infeasible. It outlines three stages—discovery, construct definition, and observation—where AI impacts the measurement pipeline and stresses the importance of rigorous validation to ensure credible inference. The review offers guidance on navigating the shift from a single scalable measure to multiple plausible ones that can lead to differing empirical conclusions.

By Melissa Dell, Ashesh Rambachan
arXiv Machine Learning
Aug 11

Demystifying Prediction Powered Inference

arXiv:2601. 20819v2 Announce Type: replace-cross Abstract: Machine learning predictions are increasingly used to supplement incomplete or costly-to-measure outcomes in fields such as biomedical research, environmental science, and social science.

By Yilin Song, Dan M. Kluger, Harsh Parikh, Tian Gu