arXiv AI By Keito Inoshita

Bias-Corrected Ceilings of Emotion Predictability from Human Label Variation Based on Instance-Level Fano Bounds

Read the original on arXiv AI →

arXiv:2608. 15619v1 Announce Type: new Abstract: Emotion recognition from text keeps improving on benchmarks, yet whether an accuracy ceiling has been reached is seldom asked with discipline.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 11

When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text

Large language models (LLMs) are increasingly used to assess social bias in text, but the passages they evaluate often contain surface noise such as typos and broken punctuation. This study applied five realistic noise conditions at varying intensities to 3,822 stereotype‑related responses and compared bias judgments on noisy versus original text. The findings show that noise disproportionately turns neutral judgments into biased ones—up to 120 times more likely—while rarely converting biased judgments into neutral ones, and that the most fragile LLM judge exhibits the greatest distortion at mild noise levels. As LLMs become more robust, the bias distortion tends toward parity rather than reversal, meaning bias measured on noisy text is systematically overestimated, especially in fairness‑critical categories.

By DongHyun Ryu, Jaehyeok Lee, YeongJun Hwang, JinYeong Bak
Hugging Face Trending Papers
Sep 10

When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text

Large language models used as judges for social bias are affected by noisy text, such as typos and broken punctuation. In experiments with 3,822 stereotype-related responses, noise more often turns neutral judgments into biased ones than the reverse, with up to a 120‑fold difference. The effect is strongest at mild realistic noise levels and leads to systematic overestimation of bias, especially in fairness‑critical categories.