arXiv AI By Sher Badshah, Ali Emami, Hassan Sajjad

SCOPE: Selective Conformal Optimized Pairwise LLM Judging

Read the original on arXiv AI →

arXiv:2602. 13110v4 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used as scalable judges in pairwise evaluation, but they remain prone to miscalibration and biases.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.