arXiv AI By Ha-Linh Nguyen, Hong-Anh Nguyen, Minh-Duc La, Phong Lam, Thu-Trang Nguyen, Son Nguyen, Hieu Dinh Vo

Noise-Aware Framework for Correcting Corrupted Labels

Read the original on arXiv AI →

arXiv:2606. 11695v1 Announce Type: cross Abstract: High-quality labeled data is essential for training reliable ML/DL models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 1

Forget or Fine-tune? A Comparative Study of Machine Unlearning Strategies for Noisy Label Correction

The paper compares five machine unlearning (MU) methods—NegGrad, Fine‑Tuning (FT), Random Labeling (RL), SalUn, and MUNBa—on noisy‑label correction across CIFAR‑10, CIFAR‑100, and Food‑101N. Results show that the best MU strategy depends on the noise type: FT works well for most closed‑set noise, RL and SalUn are robust and nearly match retraining accuracy under instance‑dependent noise, while MUNBa excels only under extreme symmetric noise. In open‑set noise, retraining on the cleaned data actually hurts performance, indicating that approximating retraining is not suitable in that regime, yet all MU methods still achieve near‑retraining accuracy on Food‑101N with much lower runtime.

By Jo\~ao L. P. Santana, Filipe R. Cordeiro
arXiv Machine Learning
Aug 24

Benchmarking noisy label detection methods

arXiv:2510.16211v2 Announce Type: replace Abstract: Label noise is a common problem in real-world datasets, affecting both model training and validation. Clean data are essential for achieving strong...

By Henrique Pickler, Jorge K. S. Kamassury, Danilo Silva
arXiv AI
Sep 18

Pre-train to Gain: Robust Learning Without Clean Labels

The paper proposes a method that first pre‑trains a feature extractor on the target dataset using in‑domain self‑supervised learning (SSL) without labels, then performs standard supervised training on the same noisy dataset. This two‑stage approach eliminates the need for a clean label subset and consistently improves classification accuracy and label‑error detection across synthetic and real‑world noise, especially as noise rates increase. Experiments show that the method matches or surpasses ImageNet and DinoV2 pre‑training, particularly under high noise conditions.

By David Szczecina, Nicholas Pellegrino, Paul Fieguth