Spectral Overfitting in Noisy Linear Probing of Pretrained Representations
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2607. 12438v1 Announce Type: new Abstract: Deep networks trained with label noise often learn clean structure before memorizing corrupted labels.
The paper proposes a method that first pre‑trains a feature extractor on the target dataset using in‑domain self‑supervised learning (SSL) without labels, then performs standard supervised training on the same noisy dataset. This two‑stage approach eliminates the need for a clean label subset and consistently improves classification accuracy and label‑error detection across synthetic and real‑world noise, especially as noise rates increase. Experiments show that the method matches or surpasses ImageNet and DinoV2 pre‑training, particularly under high noise conditions.
The paper investigates how knowledge distillation from Vision Transformers to smaller CNNs can cause dimensional collapse in the student’s representation space. Using SVD and Shannon entropy, the authors show that cosine‑based distillation leads to a drastic reduction in effective rank, while adding an InfoNCE objective can double the rank but harms downstream accuracy due to signal dilution. They further demonstrate that a label‑aware contrastive objective (Supervised Contrastive distillation) can maintain or improve accuracy without unnecessary rank expansion, indicating that effective rank alone is not a reliable indicator of representation quality.
arXiv:2608.29167v1 Announce Type: new Abstract: Robustness of segmentation models is commonly assessed through input-domain perturbations, while dependence on frequency content within learned feature...
arXiv:2608. 07921v1 Announce Type: cross Abstract: We apply Marchenko-Pastur (MP) random matrix theory to pre-trained attention weights in order to separate each projection matrix into a random-like bulk and a set of spectral outliers.
arXiv:2607.01630v2 Announce Type: replace Abstract: Dynamic expansion methods for class-incremental learning (CIL) protect task-specific knowledge by growing dedicated tokens or subnetworks, yet our...