arXiv:2606. 08718v1 Announce Type: cross Abstract: While Deep Active Learning (DAL) effectively reduces human annotation costs, its efficacy is constrained by human annotation errors.
By Md Abdullah Al Forhad, Weishi Shi
arXiv:2510.16211v2 Announce Type: replace
Abstract: Label noise is a common problem in real-world datasets, affecting both model training and validation. Clean data are essential for achieving strong...
By Henrique Pickler, Jorge K. S. Kamassury, Danilo Silva
arXiv:2606. 07630v1 Announce Type: cross Abstract: Real-world datasets across image and text domains are often characterized by skewed class distributions and noisy annotations, which jointly degrade model performance, particularly on minority classes.
By Jiancheng Zhang, Meiqing Li, Qi Zhang, Yinglun Zhu
arXiv:2608. 03511v1 Announce Type: cross Abstract: Active learning (AL) promises to reduce the cost of medical imaging projects by lowering the number of clinical labels required.
By Julia Machnio, Mads Nielsen, Mostafa Mehdipour Ghazi
arXiv:2606. 14965v1 Announce Type: new Abstract: Synthetic instance-dependent label noise (IDN) benchmarks are widely used to evaluate noisy-label learning methods, yet existing approaches typically generate noise through imperfect annotators or classifier raters, leaving the source of ambiguity implicit.
By Shadman Islam, Agustinus Kristiadi, Mostafa Milani
arXiv:2606. 11699v1 Announce Type: new Abstract: The performance of machine learning and deep learning models largely depends on the quality of the training data.
By Ha-Linh Nguyen, Hong-Anh Nguyen, Minh-Duc La, Thu-Trang Nguyen, Son Nguyen, Hieu Dinh Vo
arXiv:2608. 13601v1 Announce Type: new Abstract: Active learning can reduce labeling cost by selecting informative examples, but the most uncertain examples may also be the hardest to label correctly.
By John Myron Uy
arXiv:2606. 11130v1 Announce Type: new Abstract: We study the task of agnostically learning general (as opposed to homogeneous) ReLUs under the Gaussian distribution with respect to the squared loss.
By Ilias Diakonikolas, Daniel M. Kane, Mingchen Ma
arXiv:2609.15255v1 Announce Type: new
Abstract: Ecological monitoring increasingly relies on machine learning models, whose performance depends on the quality and quantity of labelled data. However,...
By Ben McEwen, Rupa Kurinchi-Vendhan, Shiqi Zhang, Lukas Rauch, Marek Herde, Sara Beery
arXiv:2608. 03432v1 Announce Type: new Abstract: Refurbishment-based noisy-label learning mixes an observed label with a model-derived pseudo target, typically using one sample-wise cleanliness score to control both branches.
By Wenxiao Fan, Kan Li
arXiv:2606. 17805v1 Announce Type: new Abstract: Data acquisition is a major bottleneck for learning in real-time streams: analysts must decide on the fly which labels to purchase while respecting a rolling budget.
By Xiwen Huang, Pierre Pinson
arXiv:2609.26631v1 Announce Type: new
Abstract: Accurate ground-based cloud classification is important for atmospheric monitoring, solar-energy forecasting, aviation weather assessment, and climate...
By Esther Bou Dagher, Viktoriya Bu-Dager, Boguslaw Zegarlinski