arXiv AI By Zilong Zhang, Yi-Ting Hung, Weiyi He, Junxi Zhang, Lei Ding, Chi-Kuang Yeh

AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing

Read the original on arXiv AI →

arXiv:2606. 19714v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as judges for open-ended generation, as large-scale human evaluation is often expensive and difficult to scale, yet their preferences remain imperfect proxies for human judgment.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.