arXiv Machine Learning By Le Peng, Yash Travadi, Chuan He, Ying Cui, Ju Sun

Exact Reformulation and Optimization for Direct Metric Optimization in Binary Imbalanced Classification

Read the original on arXiv Machine Learning →

The paper presents an exact constrained reformulation for direct metric optimization (DMO) in binary imbalanced classification, focusing on precision, recall, and F1-score under three settings: fixing precision to optimize recall, fixing recall to optimize precision, and optimizing F1-score. Unlike prior approaches that use smooth approximations, the authors introduce exact penalty methods to solve these problems efficiently. Experiments on benchmark datasets show that this exact reformulation and optimization (ERO) framework outperforms state‑of‑the‑art methods for all three DMO tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Statistics ML
Aug 25

Robust performance metrics for imbalanced classification problems

The paper demonstrates that common binary classification metrics—Matthews' correlation coefficient, Cohen's κ, the F-score, and the Jaccard similarity—are not robust to extreme class imbalance, as the Bayes classifier’s true positive rate tends to zero when the minority class proportion vanishes. To address this, the authors propose robustified versions of these metrics that include a tuning parameter, ensuring that the Bayes-optimal classifier’s threshold remains bounded and its true positive rate stays above zero even in highly imbalanced scenarios. The study provides theoretical bounds, simulation results, and practical guidance on applying these robust metrics to real data, such as a credit‑default dataset, and discusses their relationship to ROC and precision‑recall curves.

By Hajo Holzmann, Bernhard Klar
arXiv AI
Jun 2

How Hard Can It Be? Hardness-Aware Multi-Objective Unlearning

arXiv:2606. 02119v1 Announce Type: cross Abstract: Machine unlearning aims to remove the influence of specific forget training data due to privacy, copyright or bias concerns while maintaining the model performance on the remaining retain data.

By Jiangwei Chen, Xinyuan Niu, Rachael Hwee Ling Sim, Zhengyuan Liu, Nancy F. Chen, Bryan Kian Hsiang Low
arXiv Machine Learning
Sep 3

Bayes-Optimal BER and AUC: Estimation and Evaluation of Estimators

The paper introduces soft‑label‑based estimators for the Bayes‑optimal balanced error rate (BER) and area under the ROC curve (AUC), extending from a clean setting with known class priors to a realistic scenario with unknown priors and corrupted soft labels. It also adapts the FeeBee evaluation framework to assess these estimators without needing the true optimum, providing practical evaluation scores for any estimator of optimal BER or AUC. Experiments on synthetic and real datasets confirm the effectiveness of both the estimators and the evaluation method.

By Ryota Ushio, Takashi Ishida, Masashi Sugiyama