arXiv:2606. 05814v1 Announce Type: new Abstract: The support vector machine (SVM) is a widely used classifier, but choosing an appropriate loss function remains difficult.
By Yuliang Yang, Chen Chen, Yuxiang Liu, Huiru Wang
arXiv:2403. 05532v2 Announce Type: replace Abstract: We introduce Tune without Validation (Twin), a simple and effective pipeline for tuning learning rate and weight decay of homogeneous classifiers without validation sets, eliminating the need to hold out data and avoiding the two-step process.
By Lorenzo Brigato, Stavroula Mougiakakou
arXiv:2606. 22068v2 Announce Type: replace-cross Abstract: Most real-world datasets used for training supervised learning models are contaminated with noisy data and outliers leading to large prediction errors.
By Mathew Mithra Noel, Arindam Banerjee, Yug D. Oswal, Geraldine Bessie Amali D, Venkataraman Muthiah-Nakarajan
arXiv:2609.13526v1 Announce Type: cross
Abstract: We develop a spectral three-term modification of the classic Hestenes--Stiefel conjugate gradient algorithm, preserving its anti-jamming characterist...
By Saman Babaie-Kafaki, Maryam Khoshsimaye-Bargard, Ahmad Mousavi
arXiv:2604. 27742v2 Announce Type: replace Abstract: A fundamental dichotomy in the theory of classification sets smoothness against statistical efficiency: smooth surrogate losses such as the logistic loss enable fast $O(1/T)$ optimization but yield slow square-root $H$-consistency bounds, while piecewise-linear losses like the Hinge loss achieve optimal linear $H$-consistency rates but are non-differentiable.
By Mehryar Mohri, Yutao Zhong
arXiv:2609.36310v1 Announce Type: new
Abstract: Everywhere learning provides a principled framework for training AI models under constraints that must hold throughout the data distribution. In the du...
By Ignacio Boero, Jonathan Nixon, Alejandro Ribeiro
The paper investigates minimal‑norm interpolation and λ2‑regularized logistic‑loss minimization for binary classification using univariate two‑layer ReLU networks. It provides exact geometric characterizations of optimal classifiers, showing that unpenalized hidden‑layer biases yield continuous piecewise‑affine functions that tightly follow label switches, while penalized biases produce a unique, sparsest classifier with a single kink per same‑label segment. Adding a free affine skip connection does not change these function‑space solutions but guarantees that every KKT point becomes globally optimal, eliminating suboptimal KKT points that can arise without the skip connection.
By Karolina Drabik, Ben Lewis, Antoni Puch, Etienne Boursier, Piotr Hofman, Matthias Englert, Ranko Lazi\'c
arXiv:2604. 13130v2 Announce Type: replace Abstract: We study learning to learn through the lens of hyperparameter tuning.
By Saumya Goyal, Rohith Rongali, Ritabrata Ray, Barnab\'as P\'oczos
arXiv:2608. 00949v1 Announce Type: new Abstract: The pinball-loss support vector machine is robust, but its asymmetry parameter is usually fixed in advance.
By Xiaofei Wu, Kai Qi, Rongmei Liang
arXiv:2605. 20347v2 Announce Type: replace Abstract: Labeling a training set is often expensive and susceptible to errors, making the design of robust loss functions for label noise an important problem.
By Alexandre Lemire Paquin, Brahim Chaib-Draa, Philippe Gigu\`ere
arXiv:2609.38263v1 Announce Type: new
Abstract: Feature selection in neural networks remains a challenging problem, particularly in the presence of noisy or contaminated data. LassoNet is a recent ap...
By Daniela De Canditiis, Italia De Feis, Paola Stolfi
arXiv:2607. 22212v1 Announce Type: cross Abstract: Visual anomaly detection requires adaptive representations and reliable decision boundaries, particularly when anomalous training samples are scarce and class distributions are highly imbalanced.
By Alireza Dastmalchi Saei, Shervin Rahimzadeh Arashloo