Judge Arena: Benchmarking LLMs as Evaluators
Read the original on Hugging Face Blog →The Flow has not summarised this story yet — read it at Hugging Face Blog.
The Flow has not summarised this story yet — read it at Hugging Face Blog.
Multiple-Choice Benchmarks, Verifiers, Leaderboards, and LLM Judges with Code Examples
arXiv:2607. 08535v1 Announce Type: cross Abstract: An LLM-as-judge score can move even when the candidate responses stay fixed, simply because the evaluator has changed.
arXiv:2604. 22891v4 Announce Type: replace-cross Abstract: LLM-as-a-Judge has become a dominant approach in automated evaluation systems, playing critical roles in model alignment, leaderboard construction, quality control, and so on.
arXiv:2602. 03238v3 Announce Type: replace Abstract: LLM agent benchmark scores are shaped not only by the model but also by the agent harness, environment, evaluator, and inference budget.
arXiv:2607. 14108v1 Announce Type: cross Abstract: This paper introduces tool efficiency, a new quantitative metric to evaluate the rate of useful tool calls in an LLM agent trajectory.