Hierarchical orchestration, in which a Manager agent reviews worker output and can send it back for revision, is the default coordination pattern in production multi-agent LLM frameworks. Classical or...
The paper investigates how language‑model judges can make version‑dependent errors when evaluating upgraded agents. Using 35 public coding‑agent submissions, two customer‑service agents, and over a thousand expert‑labeled trajectories, the authors show that fixed judges often reject task‑conditioned error invariance and can incorrectly approve failed patches, especially as agent capability increases. Paired audits of current outputs reduce interval width only marginally, and the study concludes that independent human patch review is still necessary.
By Jiapeng Li
arXiv:2607. 08065v1 Announce Type: new Abstract: LLM-as-judge (Zheng et al.
By Kaihua Ding
Agent evaluations tell us that a model picked the wrong tool, but rarely why. We introduce canary tools: diagnostic probe tools planted in an agent's Model Context Protocol (MCP) tool set, each engineered to probe one specific tool-selection weakness.
arXiv:2606. 19111v1 Announce Type: cross Abstract: Team science holds that leadership is contingent: it helps only under specific conditions, and capable, autonomous teams may need none at all.
By Haewoon Kwak
arXiv:2608. 04719v1 Announce Type: new Abstract: Agent evaluations tell us that a model picked the wrong tool, but rarely why.
By Atul Anand, Sourav Chattaraj
arXiv:2607. 15388v1 Announce Type: new Abstract: Many math- and science-oriented agent systems use hierarchical designs with specialized reviewer roles, assuming that a dedicated review stage should help turn wrong candidates into correct ones.
By Chih-Hsuan Yang, Jingyan Jiang, Vikram Vasudevan, Cheng-Hau Yang, Huihuo Zheng, Le Chen, Eliu A. Huerta, Venkatram Vishwanath, Ian T. Foster, Rajeev Thakur
The paper investigates whether widely used evaluation frameworks for large language models (LLMs) implement a defense called commit‑first judging, which requires a judge to solve a task itself before accepting a candidate answer. Across 24 configurations in eight popular frameworks, none use the full commit‑first method; nine use a weaker variant that is ineffective. In controlled experiments, the weaker variant allowed systems to game the judge, while the full commit‑first approach eliminated this vulnerability but sometimes worsened evaluation when the judge’s own answer was incorrect.
By Idil Gozel
LLM judges, models that score another system's output, can be gamed by the systems they score. Recent work identifies one defence that works: the judge solves the task itself first and commits to that...
Praxa is an evidence‑bound harness for governed AI agent execution that explicitly represents states such as proposal, authority, dispatch, verified external effect, and promotion through deterministic admission, brokered execution, external read‑back, reconciliation, and reviewed promotion. The authors report four evidence lanes: a repository‑local audit passing all unit and Workerd tests; a pilot on 12 curated tasks where both baseline and reliability‑layer arms passed 17 of 36 trials; a coordination‑proxy comparison where both baseline and a source‑authored candidate completed all 180 trials with equal accuracy but the candidate used fewer tokens and steps; and deployed source/configuration evidence showing bounded reflection, recall accounting, memory compilation, and tool‑health paths. None of the evidence demonstrates superiority in security, safety, or user benefit.
whyItMatters:"Praxa provides a testable architecture that makes authority‑to‑effect transitions explicit, offering a framework for verifying AI agent behavior, though current evidence does not prove improved security or performance."
By Stefan G. Creadore
The paper argues that using a large language model (LLM) as the sole judge in self‑improving agent pipelines is problematic, as the judge can be biased or manipulated, leading to false confidence in system performance. The authors propose a new framework, PROCTOR, which replaces the oracle judge with a deterministic, teacher‑student loop that enforces guardrails such as sandboxing, role separation, and acceptance checks to prevent cheating and ensure reliable evaluation. Experiments across contract analysis, compliance review, and code quality demonstrate that PROCTOR mitigates eleven identified failure modes that previously allowed agents to achieve perfect scores while hiding significant capability gaps.
By Vansh Wahi
The study investigates how varying the capability of a reviewer model in a large‑language‑model (LLM) execute‑review‑revise pipeline affects rejection decisions on 100 olympiad mathematics problems. A mid‑tier reviewer improves final accuracy by 12 percentage points (from 52 % to 64 %) without damaging answers, while a self‑reviewer detects errors best (85 % recall) but rejects too often and harms correct solutions. Below a certain capability threshold the reviewer becomes inert, changing none of the answers and doubling token cost.
By Faizan Tanveer