arXiv Machine Learning

Symmetrization of Loss Functions for Robust Training of Neural Networks in the Presence of Noisy Labels

arXiv:2605. 20347v2 Announce Type: replace Abstract: Labeling a training set is often expensive and susceptible to errors, making the design of robust loss functions for label noise an important problem.

arXiv Machine Learning
Sep 24

Minimal-Norm Univariate Two-Layer ReLU Classification: Exact Solutions and Global Optimality with Skip Connections

The paper investigates minimal‑norm interpolation and λ2‑regularized logistic‑loss minimization for binary classification using univariate two‑layer ReLU networks. It provides exact geometric characterizations of optimal classifiers, showing that unpenalized hidden‑layer biases yield continuous piecewise‑affine functions that tightly follow label switches, while penalized biases produce a unique, sparsest classifier with a single kink per same‑label segment. Adding a free affine skip connection does not change these function‑space solutions but guarantees that every KKT point becomes globally optimal, eliminating suboptimal KKT points that can arise without the skip connection.

By Karolina Drabik, Ben Lewis, Antoni Puch, Etienne Boursier, Piotr Hofman, Matthias Englert, Ranko Lazi\'c
arXiv Computer Vision
Sep 14

Learning Sign Language Recognition under Label Noise: A Study of Noise-Robust Losses for Isolated and Continuous Settings

The paper investigates the impact of label noise on sign language recognition, comparing robust loss functions—symmetric cross entropy (SCE) and generalized cross entropy (GCE)—to standard cross entropy (CE) in both isolated (ISLR) and continuous (CSLR) settings. Experiments on ASL Citizen with injected symmetric noise show that SCE and GCE outperform CE across multiple backbones, though GCE’s optimal hyperparameter does not transfer well. In CSLR experiments on PHOENIX-2014, robust losses offer limited gains, with performance largely influenced by auxiliary weight settings rather than the loss choice.

By Akihisa Shitara, Yoichi Ochiai
arXiv AI
Jun 18

Generalized Kullback-Leibler Divergence Loss

arXiv:2503. 08038v2 Announce Type: replace-cross Abstract: In this paper, we delve deeper into the Kullback-Leibler (KL) Divergence loss and mathematically prove that it is equivalent to the Decoupled Kullback-Leibler (DKL) Divergence loss that consists of (1) a weighted Mean Square Error (wMSE) loss and (2) a Cross-Entropy loss incorporating soft labels.

By Jiequan Cui, Beier Zhu, Qingshan Xu, Zhuotao Tian, Xiaojuan Qi, Bei Yu, Hanwang Zhang, Richang Hong
arXiv Machine Learning
Aug 13

Linear-Core Surrogates: Smooth Loss Functions with Linear Rates for Classification and Structured Prediction

arXiv:2604. 27742v2 Announce Type: replace Abstract: A fundamental dichotomy in the theory of classification sets smoothness against statistical efficiency: smooth surrogate losses such as the logistic loss enable fast $O(1/T)$ optimization but yield slow square-root $H$-consistency bounds, while piecewise-linear losses like the Hinge loss achieve optimal linear $H$-consistency rates but are non-differentiable.

By Mehryar Mohri, Yutao Zhong
arXiv Machine Learning
Jun 2

Robust Learning of a Group DRO Neuron

arXiv:2601. 18115v2 Announce Type: replace Abstract: We study the problem of learning a single neuron under standard squared loss in the presence of arbitrary label noise and group-level distributional shifts, for a broad family of covariate distributions.

By Guyang Cao, Shuyao Li, Sushrut Karmalkar, Jelena Diakonikolas