arXiv Machine Learning By Alexandre Lemire Paquin, Brahim Chaib-Draa, Philippe Gigu\`ere

Symmetrization of Loss Functions for Robust Training of Neural Networks in the Presence of Noisy Labels

Read the original on arXiv Machine Learning →

arXiv:2605. 20347v2 Announce Type: replace Abstract: Labeling a training set is often expensive and susceptible to errors, making the design of robust loss functions for label noise an important problem.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 24

Minimal-Norm Univariate Two-Layer ReLU Classification: Exact Solutions and Global Optimality with Skip Connections

The paper investigates minimal‑norm interpolation and λ2‑regularized logistic‑loss minimization for binary classification using univariate two‑layer ReLU networks. It provides exact geometric characterizations of optimal classifiers, showing that unpenalized hidden‑layer biases yield continuous piecewise‑affine functions that tightly follow label switches, while penalized biases produce a unique, sparsest classifier with a single kink per same‑label segment. Adding a free affine skip connection does not change these function‑space solutions but guarantees that every KKT point becomes globally optimal, eliminating suboptimal KKT points that can arise without the skip connection.

By Karolina Drabik, Ben Lewis, Antoni Puch, Etienne Boursier, Piotr Hofman, Matthias Englert, Ranko Lazi\'c
arXiv Computer Vision
Sep 14

Learning Sign Language Recognition under Label Noise: A Study of Noise-Robust Losses for Isolated and Continuous Settings

The paper investigates the impact of label noise on sign language recognition, comparing robust loss functions—symmetric cross entropy (SCE) and generalized cross entropy (GCE)—to standard cross entropy (CE) in both isolated (ISLR) and continuous (CSLR) settings. Experiments on ASL Citizen with injected symmetric noise show that SCE and GCE outperform CE across multiple backbones, though GCE’s optimal hyperparameter does not transfer well. In CSLR experiments on PHOENIX-2014, robust losses offer limited gains, with performance largely influenced by auxiliary weight settings rather than the loss choice.

By Akihisa Shitara, Yoichi Ochiai