The paper introduces soft‑label‑based estimators for the Bayes‑optimal balanced error rate (BER) and area under the ROC curve (AUC), extending from a clean setting with known class priors to a realistic scenario with unknown priors and corrupted soft labels. It also adapts the FeeBee evaluation framework to assess these estimators without needing the true optimum, providing practical evaluation scores for any estimator of optimal BER or AUC. Experiments on synthetic and real datasets confirm the effectiveness of both the estimators and the evaluation method.
By Ryota Ushio, Takashi Ishida, Masashi Sugiyama
The paper demonstrates that common binary classification metrics—Matthews' correlation coefficient, Cohen's κ, the F-score, and the Jaccard similarity—are not robust to extreme class imbalance, as the Bayes classifier’s true positive rate tends to zero when the minority class proportion vanishes. To address this, the authors propose robustified versions of these metrics that include a tuning parameter, ensuring that the Bayes-optimal classifier’s threshold remains bounded and its true positive rate stays above zero even in highly imbalanced scenarios. The study provides theoretical bounds, simulation results, and practical guidance on applying these robust metrics to real data, such as a credit‑default dataset, and discusses their relationship to ROC and precision‑recall curves.
By Hajo Holzmann, Bernhard Klar
arXiv:2608.30028v1 Announce Type: new
Abstract: This paper introduces a family of multiclass linear Perceptron classifiers with a multiplicative margin mechanism (MMPerc), as an alternative to standa...
By Dmitri Rachkovskij, Evgeny Osipov, Olexander Volkov, Daswin De Silva, Denis Kleyko
arXiv:2607. 11956v1 Announce Type: cross Abstract: Data Shapley is the standard principled answer to which training points are worth what, and its k-nearest-neighbor (KNN) specialization is the version deployed in practice: the exact estimator shipped by toolkits such as pyDVL and OpenDataVal.
By Zongye Lyu
arXiv:2409. 13007v3 Announce Type: replace-cross Abstract: Class imbalance poses a significant challenge in classification tasks, often causing standard learning algorithms to become biased toward the majority class.
By Asif Newaz, Asif Ur Rahman Adib, Taskeed Jabid
The paper introduces the Adaptive Margin Ordinal Loss (AMOL), a new loss function designed to reduce the tendency of neural networks to predict center classes in ordinal classification tasks—a problem called center‑class hedging. AMOL applies a multiplicative weight to per‑class loss terms that is large only when a candidate class is near the center while the true label is far from it, thereby discouraging hedging. The authors also propose the Center‑Hedging Rate (CHR) metric to quantify this failure mode and demonstrate that AMOL achieves state‑of‑the‑art Quadratic Weighted Kappa scores on four benchmarks, with an asymmetric variant eliminating hedging on the Abalone dataset.
By Manisha Kandel