arXiv AI By Hiroyasu Usami, Keisuke Hara, Ayato Tsuboi, Naohiko Matsuda

LLM Judges Have Dark Current: A Psychometric Datasheet for LLM-as-a-Judge Evaluation

Read the original on arXiv AI →

arXiv:2606. 15610v1 Announce Type: cross Abstract: LLM-as-a-judge systems are now routinely used for open-ended model evaluation, where human preference annotation is costly, slow, and difficult to reproduce.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.