arXiv Computation and Language

How to Estimate Whether You Have Found Several Needles in a Haystack: Measuring Calibration in Multi-Label Text Classification

arXiv AI
Sep 10

The Accuracy Paradox: Empirical Diagnostic of Default Decision Thresholds in Multi-Label Enzyme Commission Prediction [With Code]

The study evaluates the use of default decision thresholds (t=0.50) in multi‑label enzyme commission (EC) number prediction across 14,096 compounds and six EC classes. It finds a high mean accuracy of 77.16% but low macro F1 (0.3976) and macro recall (0.3872), indicating severe class‑imbalance issues: majority classes are over‑predicted while minority classes, especially EC6, have zero recall despite reasonable ROC‑AUC. The authors recommend target‑specific threshold tuning and conformal calibration as post‑processing safeguards to expose and correct these hidden errors.

By Bilal Ahmad, Rajed Mehmood