arXiv Machine Learning By Jason Z Wang

ERRORQUAKE: Heavy-Tailed Error Severity Distributions in Open-Weight Large Language Models

Read the original on arXiv Machine Learning →

arXiv:2606. 05170v1 Announce Type: new Abstract: At matched accuracy, open-weight LLMs differ substantially in the shape of their error severity distribution -- a difference invisible to the scalar error rate.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
2d ago

Evaluating LLM-as-a-Judge Beyond Score Alignment: A Psychometric Analysis of Residual Judging Difficulty

The paper investigates how well large language models (LLMs) can serve as judges in summarization evaluation by applying psychometric techniques. Using Many‑Facet Rasch Models, the authors decompose human and LLM ratings into latent summary quality, rater severity, dimension severity, and rating‑scale thresholds, and introduce a residual hardness metric to capture judging difficulty. Their analysis of 17 open‑weight LLM judges on the SummEval dataset reveals that moderate alignment in latent quality does not translate to alignment in residual hardness; humans and LLMs differ in which summary–dimension units are hard, with LLMs tending to find consistency hard and humans tending to find coherence hard, and some hard cases can be predicted from observable source–summary properties.

By Longwei Cong, Sonja Hahn, Sebastian Gombert, Leon Camus, Fabian Zehner, Hendrik Drachsler, Ulf Kroehne
arXiv Computation and Language
Sep 11

Rethinking Verbalized Confidence for LLM-as-a-Judge: A Compatibility Shift on Post-2025 Proprietary Models

The paper argues that verbalized confidence—once viewed as overconfident and coarse—has become the preferred soft‑scoring method for LLM‑as‑a‑Judge on top‑tier proprietary models released after 2025. Experiments on SummEval, AggreFact, and HelpSteer2 across up to 18 LLMs show that log‑probabilities are no longer the best signal, and that adding an overconfidence advisory and self‑debate further improves calibration and robustness. The authors note that these enhancements incur little accuracy loss on post‑2025 models but do affect pre‑2025 ones, highlighting a compatibility shift in how confidence should be measured.

By Yu-Chung Hsiao
arXiv Machine Learning
Sep 17

Safety-Flag: A Unified Benchmark for the Reliability and Calibration of LLM Content Moderators

Safety-Flag is a unified benchmark that consolidates seven popular safety datasets into a single balanced flag/do‑not‑flag protocol, providing item‑level decisions and confidence scores for multiple large language models and dedicated guards. The benchmark evaluates moderator reliability across three dimensions—error direction, probability calibration, and confidence‑based error ranking—revealing that aggregate accuracy masks significant differences, such as one model flagging 85% of benign content while another misses 54% of harmful content. The study shows that general‑purpose models are overconfident, but temperature tuning can substantially improve calibration, and confidence‑based abstention can reduce selective risk, though performance varies with how well confidence ranks errors.

By Yibo Hu