The paper proposes a method that first pre‑trains a feature extractor on the target dataset using in‑domain self‑supervised learning (SSL) without labels, then performs standard supervised training on the same noisy dataset. This two‑stage approach eliminates the need for a clean label subset and consistently improves classification accuracy and label‑error detection across synthetic and real‑world noise, especially as noise rates increase. Experiments show that the method matches or surpasses ImageNet and DinoV2 pre‑training, particularly under high noise conditions.
By David Szczecina, Nicholas Pellegrino, Paul Fieguth
arXiv:2606. 07086v1 Announce Type: cross Abstract: Deep neural networks (DNNs) excel in computer vision tasks given large annotated datasets.
By Chen-Hsuan Fang, Wei-Hsinag Chen, Pin-Hsuan Yu, Jung-Hua Wang, Tsung-Wei Pan
arXiv:2606. 11699v1 Announce Type: new Abstract: The performance of machine learning and deep learning models largely depends on the quality of the training data.
By Ha-Linh Nguyen, Hong-Anh Nguyen, Minh-Duc La, Thu-Trang Nguyen, Son Nguyen, Hieu Dinh Vo
arXiv:2608.20710v1 Announce Type: new
Abstract: Real-world semi-supervised learning (SSL) often encounters significant challenges with long-tailed label distributions and noisy pseudo-labels, which h...
By Hongyang He, Xinyuan Song, Yan Zhong, Daizong Liu, Yanbin Li, Yang-fan He, Wenqiao Zhang
arXiv:2606. 11695v1 Announce Type: cross Abstract: High-quality labeled data is essential for training reliable ML/DL models.
By Ha-Linh Nguyen, Hong-Anh Nguyen, Minh-Duc La, Phong Lam, Thu-Trang Nguyen, Son Nguyen, Hieu Dinh Vo
PaSta introduces a Partial label-based Self‑training framework for noisy node classification on graphs. The method trains multiple annotators to generate high‑quality partial labels, then uses a partial‑label classification model with two loss functions to learn both labels and representations. A closed‑loop self‑training strategy further refines annotators, yielding an average 1.1% improvement over state‑of‑the‑art methods across five datasets.
By Yujing Liu, Yixin Liu, Yu Zheng, Yue Tan, Alan Wee-Chung Liew, Shirui Pan
Noisy node classification problem is a fundamental yet challenging task for real-world graph-related web services, where node labels are often corrupted or unreliable due to weak supervision or automa...
arXiv:2606. 14965v1 Announce Type: new Abstract: Synthetic instance-dependent label noise (IDN) benchmarks are widely used to evaluate noisy-label learning methods, yet existing approaches typically generate noise through imperfect annotators or classifier raters, leaving the source of ambiguity implicit.
By Shadman Islam, Agustinus Kristiadi, Mostafa Milani
arXiv:2606. 08718v1 Announce Type: cross Abstract: While Deep Active Learning (DAL) effectively reduces human annotation costs, its efficacy is constrained by human annotation errors.
By Md Abdullah Al Forhad, Weishi Shi
arXiv:2604. 06614v2 Announce Type: replace-cross Abstract: Prompt learning has gained significant attention as a parameter-efficient approach for adapting large pre-trained vision-language models to downstream tasks.
By Yaqi Zhao, Haoliang Sun, Yating Wang, Yongshun Gong, Yilong Yin
arXiv:2405. 03386v2 Announce Type: replace Abstract: Training with noisy class labels impairs neural networks' generalization performance.
By Marek Herde, Lukas L\"uhrs, Denis Huseljic, Bernhard Sick
arXiv:2609.14451v1 Announce Type: cross
Abstract: Modern semi-supervised learning (SSL) couples pseudo-label generation and classifier training, using the classifier's own confidence to select the ps...
By Itai David, Daphna Weinshall