arXiv AI By Yuma Asato, Kiyoaki Shirai, Natthawut Kertkeidkachorn

Mitigating Scoring Bias in LLM-as-a-Judge via Random Number Generation

Read the original on arXiv AI →

arXiv:2608. 05726v1 Announce Type: cross Abstract: Large Language Models (LLMs) are often used as evaluators of text quality, known as LLM-as-a-Judge, which can outperform conventional automatic evaluation metrics that rely on reference texts.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.