arXiv AI By Zixiang Xu, Sixian Li, Huaxing Liu, Xiang Wang, Shuai Li, Zirui Song, Xiuying Chen

Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias

Read the original on arXiv AI →

arXiv:2607. 11871v1 Announce Type: cross Abstract: Existing studies of LLM-as-judge scoring bias work predominantly at the input-output level: they perturb inputs, measure score deltas, and propose prompt-level mitigations.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.