arXiv Computation and Language
Sep 3

The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment

The paper investigates whether consensus among large language model (LLM) judges truly reflects human alignment. By treating each judge’s scores as vectors, the authors measure spread, effective rank, and angles to human scores across 42 judges on Indic benchmarks, revealing that inter‑judge agreement often mirrors shared blind spots rather than human judgments. They find that while judges agree as much as humans, they only reach 58‑66% of human agreement and frequently focus on axes humans do not weight, indicating that ensemble agreement alone is insufficient evidence of alignment.

By Sourabrata Mukherjee, Hamna Hamna, Kalika Bali, Sunayana Sitaram