arXiv Machine Learning By Dennis Frauen, Athiya Deviyani, Mihaela van der Schaar, Stefan Feuerriegel

Nonparametric LLM Evaluation from Preference Data

Read the original on arXiv Machine Learning →

arXiv:2601. 21816v2 Announce Type: replace Abstract: Evaluating the performance of large language models (LLMs) from human preference data is crucial for obtaining LLM leaderboards.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 3

Distribution-Calibrated Inference Time Compute for Thinking LLM-as-a-Judge

arXiv:2512. 03019v2 Announce Type: replace-cross Abstract: Thinking Large Language Models (LLMs) used as judges for pairwise preferences remain noisy at the single-sample level, and common aggregation rules (majority vote, soft self-consistency, or instruction-based self-aggregation) are inconsistent when ties are allowed.

By Hamid Dadkhahi, Firas Trabelsi, Parker Riley, Juraj Juraska, Mehdi Mirzazadeh