arXiv Machine Learning

How many labelers do you have? A closer look at gold-standard labels

The paper examines the common practice of aggregating multiple labels per instance into a single ‘true’ label for supervised learning. By creating a theoretical model, the authors show that using the full, non‑aggregated label information can make it easier to train well‑calibrated models, though the benefits depend on the specific problem. They predict when non‑aggregated labels will improve learning and validate these predictions on real datasets.

arXiv Machine Learning
Aug 24

Benchmarking noisy label detection methods

arXiv:2510.16211v2 Announce Type: replace Abstract: Label noise is a common problem in real-world datasets, affecting both model training and validation. Clean data are essential for achieving strong...

By Henrique Pickler, Jorge K. S. Kamassury, Danilo Silva
arXiv Machine Learning
Jul 30

The Advantage of Fine-Grained Training

arXiv:2509. 05130v2 Announce Type: replace Abstract: In classification problems, models are trained to predict a class label based on the input data features.

By Davide Pirovano, Federico Milanesio, Michele Caselle, Piero Fariselli, Matteo Osella
arXiv AI
Aug 25

FedCC: Towards Addressing Label Distribution Skews in Distillation-Based Federated Learning

FedCC is a new algorithm for distillation-based federated learning that tackles label distribution skew by allowing clients to mark ambiguous samples as 'unknown' instead of forcing a potentially wrong classification. By adding this extra class and calibrating pseudo-labels on a public dataset, FedCC balances confidence across majority and minority classes. Experiments show that FedCC outperforms existing methods, achieving 67.3% accuracy even when each client has data from only one of ten classes, whereas baselines drop to near-random performance.

By Wenxuan Ye, Onur Ayan, Xueli An, Georg Carle
arXiv Machine Learning
Sep 10

Approaching the Harm of Gradient Attacks While Only Flipping Labels

The paper investigates the impact of label‑flipping attacks on distributed machine learning, where an adversary can only flip a limited number of training labels. It formalizes the attack as a per‑round constrained optimization problem, derives a greedy label‑selection rule for logistic regression, and shows that this rule is provably optimal under mean aggregation. Experiments demonstrate that optimized label flipping can significantly degrade model accuracy, outperforming random flips, and that the attack transfers to other robust aggregators such as coordinate‑wise median and trimmed mean.

By Abdessamad El-Kabid, El-Mahdi El-Mhamdi