arXiv Statistics ML

Estimation of the Label-Noise Transition Matrix with Performance Guarantees via Selective Classification

arXiv Machine Learning
Sep 10

Selective Posterior Margin Regularization for Forward-Corrected Classification

The paper introduces Selective Posterior Margin Regularization (SPMR), a technique that enhances Forward correction for learning with class‑conditional label noise. SPMR preserves the Forward objective while converting disagreements between the corrected likelihood’s reverse posterior and the observed annotation into a graded update on the clean classifier. Experiments on five known‑transition benchmarks show that SPMR improves performance by 2.5–7.0 percentage points over full‑length Forward and remains 0.7–2.5 percentage points better when combined with Mixup and early stopping, with gains attributed to posterior‑space coefficients, transition‑adjusted targets, and pairwise actions.

By Zexing Zhang, Jichao Li, Tianyang Lei, XiongYi Lu, Yang Kewei
arXiv Machine Learning
Sep 3

Bayes-Optimal BER and AUC: Estimation and Evaluation of Estimators

The paper introduces soft‑label‑based estimators for the Bayes‑optimal balanced error rate (BER) and area under the ROC curve (AUC), extending from a clean setting with known class priors to a realistic scenario with unknown priors and corrupted soft labels. It also adapts the FeeBee evaluation framework to assess these estimators without needing the true optimum, providing practical evaluation scores for any estimator of optimal BER or AUC. Experiments on synthetic and real datasets confirm the effectiveness of both the estimators and the evaluation method.

By Ryota Ushio, Takashi Ishida, Masashi Sugiyama
arXiv Machine Learning
Jul 8

Factorizable joint shift revisited

arXiv:2601. 15036v4 Announce Type: replace Abstract: Factorizable joint shift (FJS) represents a type of distribution shift (or dataset shift) that comprises both covariate and label shift.

By Dirk Tasche
arXiv Machine Learning
Aug 27

PaSta: Noisy Node Classification with Partial Label Learning

PaSta introduces a Partial label-based Self‑training framework for noisy node classification on graphs. The method trains multiple annotators to generate high‑quality partial labels, then uses a partial‑label classification model with two loss functions to learn both labels and representations. A closed‑loop self‑training strategy further refines annotators, yielding an average 1.1% improvement over state‑of‑the‑art methods across five datasets.

By Yujing Liu, Yixin Liu, Yu Zheng, Yue Tan, Alan Wee-Chung Liew, Shirui Pan