arXiv:2608.23960v1 Announce Type: cross
Abstract: Missing labels are usually regarded as a source of information loss in classification. We study a semi-supervised setting in which the probability of...
By You-Gan Wang, Jinran Wu, Geoffrey J. McLachlan
arXiv:2609.00774v1 Announce Type: cross
Abstract: We consider semi-supervised classification from a partially classified sample arising from a two-component Weibull mixture. The feature is observed f...
By Jinran Wu, You-Gan Wang, Geoffrey J. McLachlan
arXiv:2609.06873v1 Announce Type: cross
Abstract: We study how a limited labeling budget should be allocated to minimize multiclass zero-one classification risk. We consider parametric classification...
By F. Setoudehtanzangi, Geoffrey J. McLachlan
arXiv:2607. 24943v1 Announce Type: cross Abstract: In many classification problems, reliable instance-level labels are unavailable.
By Rapha\"el Bonnet-Guerrini, Johann Ioannou-Nikolaides, Troels Petersen, Vincenzo Piuri
arXiv:2607. 18088v1 Announce Type: new Abstract: Standard evaluation of many recognition systems contains distribution shift by construction, since benchmarks place disjoint conditions in the training and test splits.
By Weijia Han, Lisha Qu
arXiv:2511. 22823v2 Announce Type: replace-cross Abstract: Weakly supervised learning has emerged as a practical alternative to fully supervised learning when complete and accurate labels are costly or infeasible to acquire.
By Miao Zhang, Junpeng Li, Changchun Hua, Yana Yang
arXiv:2512. 03322v4 Announce Type: replace-cross Abstract: Partially labelled samples arise when features are observed for all data, but class labels are available for only a subset.
By Geoffrey J. McLachlan, Jinran Wu
arXiv:2608. 11162v1 Announce Type: new Abstract: The Naive Bayes (NB) classifier remains a standard choice for categorical data, yet its widely used smoothing rules, such as Laplace, Lidstone, Krichevsky-Trofimov, and the $m$-estimate, all prescribe a fixed smoothing strength that ignores feature cardinality, sample size, and class imbalance, inducing a non-vanishing bias on modern high-cardinality tabular data.
By Nguyen Thai Anh, Truong Viet Vu, Tran Thien Thanh, Vo Nguyen Quoc Bao, Ngo Hoang Tu
arXiv:2601. 11670v3 Announce Type: replace-cross Abstract: Pseudo-label selection in semi-supervised learning is commonly driven by maximum-confidence thresholds, yet confidence alone can be unreliable under model overconfidence and class imbalance.
By Jinshi Liu, Lei He, Pan Liu
arXiv:2606. 29471v1 Announce Type: new Abstract: Strictly proper scoring rules identify the true conditional class distribution at population level, but their curvature can alter optimization and finite-sample behavior.
By Soumyadip Sarkar
arXiv:2607. 11947v1 Announce Type: cross Abstract: Typical semi-supervised learning (SSL) methods rely on distributional assumptions, and their performance degrades when these are violated.
By Yushi Hirose, Hiroo Irobe, Takafumi Kanamori
The paper introduces soft‑label‑based estimators for the Bayes‑optimal balanced error rate (BER) and area under the ROC curve (AUC), extending from a clean setting with known class priors to a realistic scenario with unknown priors and corrupted soft labels. It also adapts the FeeBee evaluation framework to assess these estimators without needing the true optimum, providing practical evaluation scores for any estimator of optimal BER or AUC. Experiments on synthetic and real datasets confirm the effectiveness of both the estimators and the evaluation method.
By Ryota Ushio, Takashi Ishida, Masashi Sugiyama