arXiv AI

When Is an Agent Evaluation Over? Outcome Finality and Cross-Unit Separation

arXiv:2608. 14940v1 Announce Type: new Abstract: Current agent evaluations score models on the state visible at the end of a stopped run which they count as one trial.

arXiv AI
Sep 7

FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflicting Financial Finality

FinalityBench is an executable benchmark that tests how agents decide on shipping, re‑capturing, refunding, or waiting when a merchant’s payment processor, ledger, ERP, and bank feed receive delayed, duplicated, dropped, or reordered messages, causing contradictory beliefs about an order. The benchmark uses a hidden canonical event log and faulted delivery streams to generate system views, scoring each episode by the merchant’s terminal economic position relative to a privileged reference. It contains 321 tasks, including 45 twin pairs where all four views are identical yet the correct disposition differs, and evaluates nine programmatic policies, revealing that a ship‑on‑first‑sign policy performs best by accuracy but worst by paired loss, while a runtime‑gated irreversible‑action policy achieves 85.4% accuracy without losing money.

By Abhishek Sharma
arXiv AI
2d ago

Global Coherence: When Every Agent Is Right and the Team Is Still Wrong - A Local-to-Global Semantic Foundation for Multi-Agent Collaboration

The paper introduces the global coherence problem, where AI agents make locally valid decisions that collectively lead to an invalid outcome due to shared state failures. It presents the Observation‑Aliasing Impossibility Theorem, establishing that a policy can guarantee a valid action only when all indistinguishable worlds share an admissible action, and shows that even with additional reasoning, roles, messages, or samples, the missing distinction cannot be recovered. The authors propose a local‑to‑global runtime semantics framework and conduct nine studies demonstrating that missing global state cannot be substituted by local intelligence.

By Xin Heng
arXiv AI
Sep 2

trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories

The paper "trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories" examines the limitations of outcome-only evaluation for large language model agents. Using a deterministic tool‑using support‑desk environment with a scripted oracle policy and a fault injector, the authors compare five different judging approaches—programmatic rules, outcome‑only, step‑rubric at two model sizes, and a self‑consistency ensemble—on metrics such as detection, step localisation, fault typing, calibration, and cost across 400 trajectories. The study finds that outcome‑only judges miss many silent faults and generate false positives, while step‑rubric judges achieve higher recall with no false alarms but at greater cost, and that none of the judges read the final reply, allowing fabricated promises to evade detection. "whyItMatters":"The findings highlight that current production‑default outcome‑only evaluations can overlook critical failures in agent behavior, underscoring the need for more nuanced, step‑level judging methods to ensure reliable LLM agent performance."

By Hadi Mohammadi
arXiv AI
Sep 25

Evaluation of Multi-Turn Consistency in LLM Agents: Survival Analysis and Failure-Rationale Taxonomy

The study evaluates how large language model agents maintain consistency over extended interactions by simulating a 20‑step delayed‑gratification task. Researchers ran 84,540 trajectories across eight model families, using survival analysis to track when agents first claim a reward and discrete‑time hazard regression to assess how factors like social visibility, persona stressors, and deliberation policy affect failure risk. They also developed a seven‑category taxonomy from 13,780 deliberation traces, revealing that early failures are impulse‑driven, later ones are fatigue‑ or cost‑benefit‑framed, and public settings elicit norm‑oriented justifications; longer deliberation correlates with higher intra‑rationale contradictions, challenging assumptions about reasoning depth and consistency.

By Igor Bogdanov, Olga Manakina, Chung-Horng Lung
arXiv AI
6d ago

Beyond Approved Actions: Runtime Validation of Persistent Outcomes in Agent Workflows

EffectMatch is a runtime system that monitors and validates persistent changes made by large language model agents during software interactions. It collects all side‑effects within a controlled execution boundary and compares them against the application’s approved state, deciding whether to commit the changes and allow subsequent steps. In tests on 206 public business tasks, EffectMatch preserved correct executions and prevented all incorrect commits, with ablation studies showing the importance of each component.

By Haoran Zhang, Hengtong Zhang, Zhiyu Liang, Yu Yan, Decheng Zuo, Hongzhi Wang
arXiv Machine Learning
4d ago

Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation

The paper investigates how language‑model judges can make version‑dependent errors when evaluating upgraded agents. Using 35 public coding‑agent submissions, two customer‑service agents, and over a thousand expert‑labeled trajectories, the authors show that fixed judges often reject task‑conditioned error invariance and can incorrectly approve failed patches, especially as agent capability increases. Paired audits of current outputs reduce interval width only marginally, and the study concludes that independent human patch review is still necessary.

By Jiapeng Li
arXiv AI
Sep 23

Self-Healing Harness for Runtime Oversight of Agent Self-Modification

The paper introduces a self‑healing harness that enforces admission control over language‑model agents’ self‑modifications. The harness runs a Detect‑Notice‑Heal‑Validate loop, allowing agents to propose rule changes that are only granted persistent authority after demonstrating improvement on a failure case without regressing on protected cases. Across 16 benchmark runs, the harness rejected many locally beneficial proposals that caused collateral regressions, while improving task‑completion scores and reliability.

By Sina Tayebati, Divake Kumar, Nastaran Darabi, Ranganath Krishnan, Amit Ranjan Trivedi