arXiv AI By Gaojie Jin, Yong Tao, Lijia Yu, Tianjin Huang

Margin-Adaptive Confidence Ranking for Reliable LLM Judgement

Read the original on arXiv AI →

arXiv:2605. 15416v2 Announce Type: replace-cross Abstract: Jung et al.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

Hugging Face Trending Papers
Aug 3

Aggregate-then-Calibrate for Human-centered Assessment with Theoretical Guarantees

Human-centered assessment tasks, which are essential for systematic decision-making, rely heavily on human judgment and typically lack verifiable ground truth. Existing approaches face a dilemma: methods using only human judgments suffer from heterogeneous expertise and inconsistent rating scales, while methods using only model-generated scores must learn from imperfect proxies or incomplete features.