arXiv:2607. 09816v1 Announce Type: new Abstract: Class imbalance poses a fundamental challenge in risk-sensitive applications such as fraud detection and medical diagnosis, where minority-class samples are scarce yet critical for accurate classification.
By Yanxuan Yu, Dong liu, Renata Borovica-Gajic, Ying Nian Wu
arXiv:2509. 07605v2 Announce Type: replace-cross Abstract: Class imbalance poses a significant challenge to supervised classification, particularly in critical domains like medical diagnostics and anomaly detection where minority class instances are rare.
By Ali Nawaz, Amir Ahmad, Shehroz S. Khan
The paper demonstrates that common binary classification metrics—Matthews' correlation coefficient, Cohen's κ, the F-score, and the Jaccard similarity—are not robust to extreme class imbalance, as the Bayes classifier’s true positive rate tends to zero when the minority class proportion vanishes. To address this, the authors propose robustified versions of these metrics that include a tuning parameter, ensuring that the Bayes-optimal classifier’s threshold remains bounded and its true positive rate stays above zero even in highly imbalanced scenarios. The study provides theoretical bounds, simulation results, and practical guidance on applying these robust metrics to real data, such as a credit‑default dataset, and discusses their relationship to ROC and precision‑recall curves.
By Hajo Holzmann, Bernhard Klar
arXiv:2409. 13007v3 Announce Type: replace-cross Abstract: Class imbalance poses a significant challenge in classification tasks, often causing standard learning algorithms to become biased toward the majority class.
By Asif Newaz, Asif Ur Rahman Adib, Taskeed Jabid
arXiv:2605.14467v2 Announce Type: replace
Abstract: We propose a new method of learning from positive and unlabeled (PU) examples in highly imbalanced datasets. Many real-world problems, such as dise...
By Elias Zavitsanos, Georgios Paliouras
arXiv:2609.26652v1 Announce Type: cross
Abstract: Commonly, classifiers and monitoring procedures are trained from labeled data by optimizing an objective such as the misclassification rate. This may...
By Ansgar Steland
The paper investigates the consistency of surrogate loss methods for classification and policy learning when the set of admissible classifiers is constrained, such as by interpretability or fairness requirements. It shows that hinge loss is the only surrogate that preserves consistency when constraints limit only the prediction set, but consistency can fail if constraints also restrict the functional form. The authors derive conditions guaranteeing consistency for hinge-risk-minimizing classifiers and use these results to design efficient hinge-loss-based procedures for monotone classification problems.
By Toru Kitagawa, Shosei Sakaguchi, Aleksey Tetenov
Machine learning models are widely used in financial fraud and credit-risk detection, yet their adversarial robustness remains difficult to evaluate because financial tabular data involve domain-speci...
arXiv:2608.24551v1 Announce Type: cross
Abstract: Machine learning models are widely used in financial fraud and credit-risk detection, yet their adversarial robustness remains difficult to evaluate...
By Xitong Zeng, Zhaoge Bi, Yitian Yang, Huaming Chen, Quan Z. Sheng
arXiv:2607. 27143v1 Announce Type: new Abstract: High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantification under severe class imbalance and asymmetric error costs.
By Manpreet Singh, Akshatha Srikantha, Shyamal Lakhanpal
arXiv:2507. 14706v2 Announce Type: replace-cross Abstract: Detecting fraudulent credit card transactions remains a significant challenge, due to the extreme class imbalance in real-world data and the often subtle patterns that separate fraud from legitimate activity.
By Claudio Giusti, Luca Guarnera, Mirko Casu, Sebastiano Battiato
arXiv:2506. 01486v2 Announce Type: replace Abstract: Data imbalance persists as a pervasive challenge in regression tasks, introducing bias in model performance and undermining predictive reliability.
By Jelke Wibbeke, Sebastian Rohjans, Andreas Rauh