arXiv:2209. 01754v5 Announce Type: replace-cross Abstract: The empirical risk minimization approach to data-driven decision making requires access to training data drawn under the same conditions as those that will be faced when the decision rule is deployed.
By Roshni Sahoo, Lihua Lei, Stefan Wager
arXiv:2606. 01002v1 Announce Type: cross Abstract: Engression is a recently proposed and effective framework for conditional distribution learning.
By Jiaqi Huang, Gongjun Xu, Ji Zhu
arXiv:2506. 14194v2 Announce Type: replace Abstract: We present a theory for the construction of out-of-distribution (OOD) detection features for neural networks.
By Sudeepta Mondal (Mary), Xinyi (Mary), Xie, Alex Wong, Ganesh Sundaramoorthi
arXiv:2503. 18314v5 Announce Type: replace-cross Abstract: We present LoTUS, a novel Machine Unlearning (MU) method that eliminates the influence of training samples from pre-trained models, avoiding retraining from scratch.
By Christoforos N. Spartalis, Theodoros Semertzidis, Petros Daras, Efstratios Gavves
arXiv:2606. 06772v1 Announce Type: cross Abstract: Understanding the generalization performance of over-parameterized neural networks has become a central topic in deep learning theory.
By Junyu Zhou, Puyu Wang, Yunwen Lei, Marius Kloft, Yiming Ying
arXiv:2607. 26562v1 Announce Type: cross Abstract: We study optimization under performative prediction, where deploying a model affects the future data distribution.
By Hiroki Hamaguchi, Yuya Hikima, Hiroshi Sawada, Akiko Takeda
arXiv:2607. 13550v1 Announce Type: cross Abstract: Boosting is one of the most successful learning techniques for standard classification and regression tasks.
By R\'emy Chapelle (CESP, CB, EVDG), Nicolas Vayatis (CB), Bruno Falissard (CESP), Mohammed Sedki (CESP)
arXiv:2607. 14306v1 Announce Type: new Abstract: In this paper, we study the connection between an LLM's output distribution and the data used to train it.
By Zachary Izzo
arXiv:2606. 06764v1 Announce Type: cross Abstract: Recent progress has been made in understanding the statistical generalization performance of gradient descent methods for overparameterized neural networks within the neural tangent kernel (NTK) regime.
By Junyu Zhou, Puyu Wang, Yunwen Lei, Yiming Ying, Ding-Xuan Zhou
arXiv:2503. 08038v2 Announce Type: replace-cross Abstract: In this paper, we delve deeper into the Kullback-Leibler (KL) Divergence loss and mathematically prove that it is equivalent to the Decoupled Kullback-Leibler (DKL) Divergence loss that consists of (1) a weighted Mean Square Error (wMSE) loss and (2) a Cross-Entropy loss incorporating soft labels.
By Jiequan Cui, Beier Zhu, Qingshan Xu, Zhuotao Tian, Xiaojuan Qi, Bei Yu, Hanwang Zhang, Richang Hong
arXiv:2606. 29925v1 Announce Type: new Abstract: As deep learning models are increasingly deployed in high-stakes applications, providing well-calibrated uncertainty estimates has become as critical as achieving high predictive accuracy.
By Han Zhou, Teodora Popordanoska, Matthew Blaschko
arXiv:2606. 16050v1 Announce Type: cross Abstract: Robust deep learning under heavy-tailed and impulsive noise remains challenging because conventional losses such as mean squared error (MSE) exhibit unbounded sensitivity to outliers.
By Mainak Kundu, Ria Kanjilal, Ismail Uysal