arXiv AI By Zongyou Yang, Yinghan Hou, Xiaokun Yang

When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability

Read the original on arXiv AI →

arXiv:2607. 08535v1 Announce Type: cross Abstract: An LLM-as-judge score can move even when the candidate responses stay fixed, simply because the evaluator has changed.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.