The Challenger: When Do New Data Sources Justify Switching Machine Learning Models?
arXiv:2512. 18390v2 Announce Type: replace Abstract: Organizations often have an incumbent predictive model in production when new data sources become available.
arXiv:2606. 01198v1 Announce Type: new Abstract: Strategic classification studies settings in which agents respond to a deployed classifier by modifying observable features at a cost.
arXiv:2512. 18390v2 Announce Type: replace Abstract: Organizations often have an incumbent predictive model in production when new data sources become available.
arXiv:2606. 30136v1 Announce Type: new Abstract: Humans facing algorithmic decision systems have been found to ``game'' them by altering their input data (at a cost to them) in order to favorably change the algorithmic outcomes they receive (at a cost to the algorithm).
arXiv:2609.06873v1 Announce Type: cross Abstract: We study how a limited labeling budget should be allocated to minimize multiclass zero-one classification risk. We consider parametric classification...
arXiv:2510. 03950v2 Announce Type: replace Abstract: Data-centric learning seeks to improve model performance from the perspective of data quality, and has been drawing increasing attention in the machine learning community.
arXiv:2605. 23595v2 Announce Type: replace-cross Abstract: The rapid advancement of machine learning has led to an unprecedented expansion of model ecosystems, making it increasingly difficult to assess the reliability of newly released models on unseen and unlabeled data.
arXiv:2606. 28204v1 Announce Type: cross Abstract: Algorithmic developments in Strategic Classification have been mostly limited to linear classifiers in settings where the best response has a closed-form solution or can be easily approximated.
arXiv:2510. 06048v4 Announce Type: replace Abstract: Effective data selection is essential for pretraining large language models (LLMs), enhancing efficiency and improving generalization to downstream tasks.
The paper investigates the consistency of surrogate loss methods for classification and policy learning when the set of admissible classifiers is constrained, such as by interpretability or fairness requirements. It shows that hinge loss is the only surrogate that preserves consistency when constraints limit only the prediction set, but consistency can fail if constraints also restrict the functional form. The authors derive conditions guaranteeing consistency for hinge-risk-minimizing classifiers and use these results to design efficient hinge-loss-based procedures for monotone classification problems.
arXiv:2606. 02671v1 Announce Type: cross Abstract: Machine learning predictors have become essential tools for guiding automated decision making.
arXiv:2607. 10694v1 Announce Type: cross Abstract: We study the problem of optimal continual fine-tuning for a pre-trained Foundation Model deployed at a resource-limited device.
arXiv:2607. 18358v1 Announce Type: cross Abstract: Document classification is a solved problem in the laboratory and an unsolved one in the enterprise.
arXiv:2606. 19587v1 Announce Type: cross Abstract: We propose a scalable method for training prediction (machine learning) models in the predict-then-optimize paradigm, where model outputs serve as coefficients for a subsequent linear optimization task.