Evaluating RAG with LLM as a Judge
Related stories
Judge Arena: Benchmarking LLMs as Evaluators
When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability
arXiv:2607. 08535v1 Announce Type: cross Abstract: An LLM-as-judge score can move even when the candidate responses stay fixed, simply because the evaluator has changed.
Diagnosing LLM Arbitration Behavior over Pre-evidence Epistemic States in RAG-based Fact-Checking
arXiv:2606. 01120v1 Announce Type: new Abstract: In RAG-based fact-checking, LLMs are increasingly used as verifiers to check given claims against retrieved evidence.
LLM Judges Have Dark Current: A Psychometric Datasheet for LLM-as-a-Judge Evaluation
arXiv:2606. 15610v1 Announce Type: cross Abstract: LLM-as-a-judge systems are now routinely used for open-ended model evaluation, where human preference annotation is costly, slow, and difficult to reproduce.
Who Drifted: the System or the Judge? Anytime-Valid Attribution in LLM Evaluation Pipelines
arXiv:2606. 15474v1 Announce Type: new Abstract: Continuous evaluation of LLM products relies on a strong LLM judge treated as ground truth: a cheap monitor scores every interaction and a team is paged when the score drifts down.
An LLM as arbiter in RAG retrieval: picking the right candidate with reasons
Enterprise Document Intelligence [Vol. 1 #7C] - One LLM call ranks the candidates with reasons.
JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment
arXiv:2605. 25240v2 Announce Type: replace-cross Abstract: Two methodologies dominate current practices of benchmarking: rubric-based scoring evaluates items against predefined criteria, whereas comparative judgment elicits pairwise preferences between outputs.
Understanding the 4 Main Approaches to LLM Evaluation (From Scratch)
Multiple-Choice Benchmarks, Verifiers, Leaderboards, and LLM Judges with Code Examples
Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators
arXiv:2606. 07874v1 Announce Type: new Abstract: LLMs-as-judges are the only way to evaluate safety at scale.
TW-LegalBench: Measuring Taiwanese Legal Understanding
arXiv:2606. 18699v1 Announce Type: cross Abstract: Large language models (LLMs) have shown impressive capabilities across diverse tasks, yet their performance on jurisdiction-specific legal reasoning remains underexplored.
