← Back to all news
Hugging Face Blog November 19, 2024

Judge Arena: Benchmarking LLMs as Evaluators

Read the original on Hugging Face Blog →

The Flow has not summarised this story yet — read it at Hugging Face Blog.

  • llms
  • benchmarks

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

Sebastian Raschka
Oct 5, 2025

Understanding the 4 Main Approaches to LLM Evaluation (From Scratch)

Multiple-Choice Benchmarks, Verifiers, Leaderboards, and LLM Judges with Code Examples

By Sebastian Raschka, PhD
llmsbenchmarks
More like this →
arXiv AI
Jul 10

When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability

arXiv:2607. 08535v1 Announce Type: cross Abstract: An LLM-as-judge score can move even when the candidate responses stay fixed, simply because the evaluator has changed.

By Zongyou Yang, Yinghan Hou, Xiaokun Yang
llmssafety
More like this →
arXiv AI
Jun 3

Quantifying and Mitigating Self-Preference Bias of LLM Judges

arXiv:2604. 22891v4 Announce Type: replace-cross Abstract: LLM-as-a-Judge has become a dominant approach in automated evaluation systems, playing critical roles in model alignment, leaderboard construction, quality control, and so on.

By Jinming Yang, Zheng Hu, Chuxian Qiu, Zhenyu Deng, Xinshan Jiao, Tao Zhou
llmsbenchmarkssafety
More like this →
Hugging Face Blog
Dec 4, 2024

Rethinking LLM Evaluation with 3C3H: AraGen Benchmark and Leaderboard

llmsbenchmarks
More like this →
arXiv AI
Aug 10

"LLM Agent Performance" Is Not a Single Evaluation Target

arXiv:2602. 03238v3 Announce Type: replace Abstract: LLM agent benchmark scores are shaped not only by the model but also by the agent harness, environment, evaluator, and inference budget.

By Pengyu Zhu, Li Sun, Philip S. Yu, Sen Su
llmsagentsbenchmarks
More like this →
arXiv AI
Jul 17

Eta Given Delta: Defining LLM Tool Efficiency With Marginal Tool Utility

arXiv:2607. 14108v1 Announce Type: cross Abstract: This paper introduces tool efficiency, a new quantitative metric to evaluate the rate of useful tool calls in an LLM agent trajectory.

By Nyx Iskandar
llmsagentsbenchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.0.0 · bb4ee0e