arXiv Computation and Language

When Residualization Helps an Audit: Format Effects, Slice Gains, and Their Limits

arXiv Machine Learning
2d ago

Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation

The paper investigates how language‑model judges can make version‑dependent errors when evaluating upgraded agents. Using 35 public coding‑agent submissions, two customer‑service agents, and over a thousand expert‑labeled trajectories, the authors show that fixed judges often reject task‑conditioned error invariance and can incorrectly approve failed patches, especially as agent capability increases. Paired audits of current outputs reduce interval width only marginally, and the study concludes that independent human patch review is still necessary.

By Jiapeng Li
arXiv AI
Sep 10

ARC-Bench: Closed-Loop Replanning Masks Broken Action Ranking in Frozen JEPA World Models

ARC‑Bench is a new benchmark that tests whether frozen JEPA‑style latent world models can correctly rank candidate actions by latent distance. The study finds that the assumption of latent rankability fails dramatically in both navigation and manipulation tasks, with the top‑scored actions often being suboptimal. Closed‑loop replanning masks this defect, but reducing replanning frequency reveals the underlying ranking failures.

By Zhengshu Zhang, Zhiyuan Li
arXiv Computation and Language
Sep 10

Judge Circuits Explain Format-Induced Inconsistency in LLM-as-a-Judge

arXiv:2605.16023v3 Announce Type: replace Abstract: LLM-as-a-judge has become the dominant paradigm for grading model outputs at scale, yet the same model assigns systematically different scores when...

By Nils Feldhus, Tanja Baeumel, Elena Golimblevskaia, Qianli Wang, Van Bach Nguyen, Aaron Louis Eidt, Selin Kahvecioglu, Christopher Ebert, Wojciech Samek, Jing Yang, Vera Schmitt, Sebastian M\"oller, Simon Ostermann