When Residualization Helps an Audit: Format Effects, Slice Gains, and Their Limits
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The paper investigates how language‑model judges can make version‑dependent errors when evaluating upgraded agents. Using 35 public coding‑agent submissions, two customer‑service agents, and over a thousand expert‑labeled trajectories, the authors show that fixed judges often reject task‑conditioned error invariance and can incorrectly approve failed patches, especially as agent capability increases. Paired audits of current outputs reduce interval width only marginally, and the study concludes that independent human patch review is still necessary.
ARC‑Bench is a new benchmark that tests whether frozen JEPA‑style latent world models can correctly rank candidate actions by latent distance. The study finds that the assumption of latent rankability fails dramatically in both navigation and manipulation tasks, with the top‑scored actions often being suboptimal. Closed‑loop replanning masks this defect, but reducing replanning frequency reveals the underlying ranking failures.
arXiv:2606. 14530v1 Announce Type: new Abstract: Large language models encode rich information in their hidden states.
arXiv:2608. 20290v1 Announce Type: new Abstract: Whether a language model has improved itself is increasingly judged not by mean accuracy but by which individual problems it gains and loses.
arXiv:2607. 03386v1 Announce Type: new Abstract: Agentic AI systems are increasingly used to edit, refine, and repair decision policies, but evaluating these edits is difficult when per-state expert action labels are unavailable.
arXiv:2606. 09046v1 Announce Type: new Abstract: Useful audits reveal not only how often a model fails, but also where its failures concentrate.