arXiv Machine Learning

Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning

arXiv:2606. 19057v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used as judges for scalable evaluation, yet such LLM--as--a--Judge systems exhibit systematic biases that are decoupled from semantic quality, most notably verbosity bias.

Hugging Face Trending Papers
Jul 2

Many Voices, One Reward: Multi-Role Rubric Generation for LLM Judging and Reward Modeling

Reliable reward and preference signals are critical for evaluating and optimizing large language models on open-ended tasks. Rubric-based judges offer a transparent way to decompose such judgments into explicit evaluation criteria, but existing annotation-free rubric generators typically rely on a single generic evaluator.

arXiv AI
Aug 11

Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text Generation

arXiv:2510. 21891v2 Announce Type: replace-cross Abstract: To deploy large language models (LLMs) in high-stakes application domains that require substantively accurate responses to open-ended prompts, we need reliable, computationally inexpensive methods that assess the trustworthiness of long-form responses generated by LLMs.

By Dhrupad Bhardwaj, Julia Kempe, Tim G. J. Rudner