arXiv AI

LWCal: Loss-Weighted Calibration for Tabular Classifiers with Noisy Calibration Labels

arXiv Machine Learning
Sep 14

Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration

The paper investigates how large language models (LLMs) used as judges in absolute scoring tasks exhibit systematic biases that compromise reliability. It shows that a judge’s task accuracy strongly predicts both its judging accuracy and its directional bias, yet more capable examinee models consistently receive more lenient judgments. To mitigate these biases, the authors propose a calibrated weighted majority voting (WMV) ensemble that estimates judges’ error rates from inter-judge agreement patterns, achieving near-oracle performance without labeled data and improving both accuracy and fairness.

By Gemma Zhang, Prachi Badarayani, Asmi Kumar, Sadid Hasan, Sulaiman Vesal
arXiv Statistics ML
Aug 24

EDGE: a closed-form directed test for the calibration of probabilistic binary classifiers

The paper introduces EDGE, a closed‑form statistical test for assessing the calibration of probabilistic binary classifiers, specifically logistic regression. EDGE uses the same binned predicted‑versus‑observed table as a reliability diagram, projects standardized bin residuals onto a small basis of smooth calibration‑distortion shapes, and yields a null distribution that is a weighted sum of chi‑square variables. The method requires only a single pass over the data and a small eigendecomposition, avoiding refitting, resampling, or tuning, and remains robust in sparse or misspecified settings where other binned tests fail.

By Ebrahim Khaled Ebrahim, Ahmed El-Kotory