arXiv AI By David Szczecina, Nicholas Pellegrino, Paul Fieguth

Pre-train to Gain: Robust Learning Without Clean Labels

Read the original on arXiv AI →

The paper proposes a method that first pre‑trains a feature extractor on the target dataset using in‑domain self‑supervised learning (SSL) without labels, then performs standard supervised training on the same noisy dataset. This two‑stage approach eliminates the need for a clean label subset and consistently improves classification accuracy and label‑error detection across synthetic and real‑world noise, especially as noise rates increase. Experiments show that the method matches or surpasses ImageNet and DinoV2 pre‑training, particularly under high noise conditions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.