arXiv Machine Learning

(Towards) Scalable Reliable Automated Evaluation with Large Language Models

arXiv:2607. 28282v1 Announce Type: cross Abstract: Evaluating the quality and relevance of textual outputs from Large Language Models (LLMs) remains challenging and resource-intensive.

arXiv AI
Jun 10

RankLLM: Weighted Ranking of LLMs by Quantifying Question Difficulty

arXiv:2602. 12424v2 Announce Type: replace-cross Abstract: Benchmarks establish a standardized evaluation framework to systematically assess the performance of large language models (LLMs), facilitating objective comparisons and driving advancements in the field.

By Ziqian Zhang, Xingjian Hu, Yue Huang, Kai Zhang, Ruoxi Chen, Yixin Liu, Qingsong Wen, Kaidi Xu, Xiangliang Zhang, Neil Zhenqiang Gong, Lichao Sun
arXiv Computation and Language
Aug 27

IDEAlign: Comparing Ideas of Large Language Models to Domain Expert

IDEAlign introduces a new protocol for evaluating the similarity of large language model (LLM) annotations to expert judgments. It uses pick‑the‑odd‑one‑out tasks to capture expert similarity and benchmarks various similarity methods—including text embeddings, topic models, and LLM-as-a-judge—against these human ratings. Applied to educational datasets, the study finds that most metrics miss nuanced expert dimensions, with LLM-as-a-judge performing best yet still insufficient for full expert alignment.

By Hyunji Nam, Lucia Langlois, James Malamut, Mei Tan, Dorottya Demszky
arXiv AI
Sep 3

Evaluating the Evaluator: Summarization Metrics and LLM-Judges beyond English

The paper introduces BASSE, a multilingual meta‑evaluation dataset containing 2,040 human‑rated abstractive summaries produced manually or by five LLMs with four prompts. Annotators scored each summary on coherence, consistency, fluency, relevance, and 5W1H using a 5‑point Likert scale. Benchmarking shows proprietary LLM‑judge models best align with human judgments, followed by criteria‑specific automatic metrics, while open‑source judge LLMs perform poorly.

By Jeremy Barnes, Naiara Perez, Alba Bonet-Jover, Bego\~na Altuna
arXiv Computation and Language
Sep 4

Evaluating Criterion-Conditioned Behaviour of Large Language Models in Content Moderation

The paper introduces DECO, a diagnostic framework that factorises content into independent moderation criteria, allowing controlled evaluation of large language models (LLMs) at the criterion level. Using pairwise evaluation across four datasets and four LLMs, the authors find that high aggregate benchmark scores can mask significant failures when decisions hinge on specific content aspects required by individual criteria. The study underscores that aggregated labels do not guarantee reliable criterion-conditioned performance, highlighting the need for evaluation methods that explicitly assess this behavior.

By Danting Zhang, Bei Peng, Robert Loftin
arXiv Computation and Language
Sep 22

LLJ Cards: Best practices for the Use of LLMs as Judges

arXiv:2609.24516v1 Announce Type: new Abstract: In recent years, large language models (LLMs) have emerged as a popular alternative for evaluation. Often referred to as LLMs as judges (LLJs), these s...

By Khaoula Chehbouni, Melina Medjdoub, Florian Carichon, Golnoosh Farnadi, Jackie Chi Kit Cheung
Hugging Face Trending Papers
Sep 3

Evaluating Criterion-Conditioned Behaviour of Large Language Models in Content Moderation

The paper introduces DECO, a diagnostic tool that factorises content into independent criteria for evaluating large language models (LLMs) on content moderation tasks. Using DECO and pairwise evaluation across four datasets and four LLMs, the authors find that high benchmark scores can mask significant failures at the criterion level, especially when decisions hinge on specific content aspects rather than overall harmfulness. The study underscores that aggregated label performance does not guarantee reliable criterion-conditioned evaluation, calling for new methods that explicitly assess this behavior.