arXiv:2512.12870v2 Announce Type: replace-cross
Abstract: Active Learning (AL) is commonly used in applications where labeling data is expensive or time-consuming. In practice, however, labels are of...
By Pouya Ahadi, Blair Winograd, Camille Zaug, Karunesh Arora, Lijun Wang, Kamran Paynabar
arXiv:2606. 14965v1 Announce Type: new Abstract: Synthetic instance-dependent label noise (IDN) benchmarks are widely used to evaluate noisy-label learning methods, yet existing approaches typically generate noise through imperfect annotators or classifier raters, leaving the source of ambiguity implicit.
By Shadman Islam, Agustinus Kristiadi, Mostafa Milani
The paper investigates how two signals—input‑conditional uncertainty and prediction‑label loss—detect different types of data corruption in federated learning. Experiments on ResNet‑20 with CIFAR‑10 and SVHN show that prediction‑label loss excels at spotting persistent random label flips, while expected‑entropy uncertainty better identifies additive image noise. The authors argue that effective federated data‑quality assessment must match the chosen signal to the specific corruption type rather than rely solely on uncertainty measures.
By Bradley Scott, Zeqi Luo, Edmond S. L. Ho
arXiv:2608. 03432v1 Announce Type: new Abstract: Refurbishment-based noisy-label learning mixes an observed label with a model-derived pseudo target, typically using one sample-wise cleanliness score to control both branches.
By Wenxiao Fan, Kan Li
arXiv:2609.16380v1 Announce Type: new
Abstract: Class-balanced learning and label noise create a coupled failure mode: frequency correction prevents majority classes from dominating the decision rule...
By Mushir Akhtar, Akarsh J., M. Tanveer, Mohd. Arshad
arXiv:2606. 11695v1 Announce Type: cross Abstract: High-quality labeled data is essential for training reliable ML/DL models.
By Ha-Linh Nguyen, Hong-Anh Nguyen, Minh-Duc La, Phong Lam, Thu-Trang Nguyen, Son Nguyen, Hieu Dinh Vo
arXiv:2608. 04147v1 Announce Type: cross Abstract: Label noise is common in medical imaging datasets due to factors such as inter-rater variability, annotation errors, and ambiguous cases.
By Abhishek Moturu, Babak Taati, Anna Goldenberg
arXiv:2608. 03511v1 Announce Type: cross Abstract: Active learning (AL) promises to reduce the cost of medical imaging projects by lowering the number of clinical labels required.
By Julia Machnio, Mads Nielsen, Mostafa Mehdipour Ghazi
arXiv:2606. 07128v1 Announce Type: new Abstract: Raw numerical datasets remain less systematically examined in integrity screening than images, plagiarism, or summary-statistic inconsistencies.
By Zhuphua Cao
arXiv:2510.16211v2 Announce Type: replace
Abstract: Label noise is a common problem in real-world datasets, affecting both model training and validation. Clean data are essential for achieving strong...
By Henrique Pickler, Jorge K. S. Kamassury, Danilo Silva
arXiv:2606. 11699v1 Announce Type: new Abstract: The performance of machine learning and deep learning models largely depends on the quality of the training data.
By Ha-Linh Nguyen, Hong-Anh Nguyen, Minh-Duc La, Thu-Trang Nguyen, Son Nguyen, Hieu Dinh Vo
Classifiers can make identical predictions yet require labels to compare their selective performance: confidence ranks weight the same errors differently. We quantify this requirement for the area und...