arXiv AI By Mengting Chen, Yanshu Sun, Wanting Liang, Beidi Luan, Rui Sun, Dezhi Chen, Jing Li, Zuo Bai

CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation

Read the original on arXiv AI →

arXiv:2607. 29252v1 Announce Type: cross Abstract: Reliable evaluation of open-ended LLM outputs requires fine-grained rubrics, yet expert curation is costly and difficult to scale.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.