Can we trust our models? Epistemic calibration in second-order classification
arXiv:2606. 10777v1 Announce Type: new Abstract: Uncertainty estimation is critical for deploying machine learning models in high-stakes settings.
The article discusses how predictive benchmarking—evaluating machine learning models by their performance and ranking—serves as a core method in machine learning research. It argues that benchmark scores only reflect performance on specific datasets and learning problems, and that drawing broader scientific conclusions requires explicit assumptions. By adapting concepts from psychological validity theory, the authors propose validity conditions to make these assumptions clear, and demonstrate their application in two case studies (ImageNet and the Fragile Families Challenge) to illustrate how benchmark results can inform inferences about research progress and limits of predictability.
arXiv:2606. 10777v1 Announce Type: new Abstract: Uncertainty estimation is critical for deploying machine learning models in high-stakes settings.
The article discusses how artificial intelligence is reshaping measurement in economics by converting unstructured data into structured variables at low cost, enabling large‑scale measurement that was previously infeasible. It outlines three stages—discovery, construct definition, and observation—where AI impacts the measurement pipeline and stresses the importance of rigorous validation to ensure credible inference. The review offers guidance on navigating the shift from a single scalable measure to multiple plausible ones that can lead to differing empirical conclusions.
arXiv:2601. 20819v2 Announce Type: replace-cross Abstract: Machine learning predictions are increasingly used to supplement incomplete or costly-to-measure outcomes in fields such as biomedical research, environmental science, and social science.
arXiv:2607. 09668v1 Announce Type: new Abstract: Ground truth datasets play a fundamental role as reference values in the training and evaluation of machine learning models.
arXiv:2609.13914v1 Announce Type: new Abstract: Machine-learning models are commonly developed under an assumption that training and test data are sufficiently complete, balanced, labelled, and drawn...
arXiv:2403. 05532v2 Announce Type: replace Abstract: We introduce Tune without Validation (Twin), a simple and effective pipeline for tuning learning rate and weight decay of homogeneous classifiers without validation sets, eliminating the need to hold out data and avoiding the two-step process.
CausalProfiler is a synthetic benchmark generator designed to evaluate causal machine learning (Causal ML) methods more rigorously and transparently. It randomly samples causal models, data, queries, and ground truths based on explicit design choices across observation, intervention, and counterfactual reasoning levels, providing coverage guarantees and transparent assumptions. The authors demonstrate its utility by testing several state‑of‑the‑art methods under diverse conditions, both within and outside the identification regime, highlighting the insights CausalProfiler can reveal.
arXiv:2605. 17273v3 Announce Type: replace-cross Abstract: State-of-the-Art (SOTA) claims pervade Artificial Intelligence (AI) and Machine Learning (ML) research.
arXiv:2502. 20502v2 Announce Type: replace Abstract: Recent advances in Artificial Intelligence (AI) have yielded powerful computational models that, by learning from vast amounts of human-generated data, are increasingly posited as approximate models of human cognition.
arXiv:2607. 26333v1 Announce Type: cross Abstract: Chest X-ray (CXR) machine learning relies heavily on automated evaluation using reference standards that aim to approximate clinical judgment.
arXiv:2605. 23595v2 Announce Type: replace-cross Abstract: The rapid advancement of machine learning has led to an unprecedented expansion of model ecosystems, making it increasingly difficult to assess the reliability of newly released models on unseen and unlabeled data.
The paper titled "Pessimistic Meta-Induction and Its Limits: Lessons from Frequentist Statistics and Machine Learning Theory" critiques the pessimistic meta-inductive argument against scientific realism by attacking its inductive step rather than its historical premise. It introduces a new challenge, drawing on frequentist statistics, machine learning, and formal epistemology to assess induction through convergence to truth. The authors argue that ordinary enumerative induction can achieve convergence everywhere, whereas meta-induction fails to achieve even almost everywhere convergence, and in contexts where meta-induction applies, no inference method can achieve almost everywhere convergence.