arXiv AI By Avyay M. Casheekar

When Is an Agent Evaluation Over? Outcome Finality and Cross-Unit Separation

Read the original on arXiv AI →

arXiv:2608. 14940v1 Announce Type: new Abstract: Current agent evaluations score models on the state visible at the end of a stopped run which they count as one trial.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 7

FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflicting Financial Finality

FinalityBench is an executable benchmark that tests how agents decide on shipping, re‑capturing, refunding, or waiting when a merchant’s payment processor, ledger, ERP, and bank feed receive delayed, duplicated, dropped, or reordered messages, causing contradictory beliefs about an order. The benchmark uses a hidden canonical event log and faulted delivery streams to generate system views, scoring each episode by the merchant’s terminal economic position relative to a privileged reference. It contains 321 tasks, including 45 twin pairs where all four views are identical yet the correct disposition differs, and evaluates nine programmatic policies, revealing that a ship‑on‑first‑sign policy performs best by accuracy but worst by paired loss, while a runtime‑gated irreversible‑action policy achieves 85.4% accuracy without losing money.

By Abhishek Sharma
arXiv AI
2d ago

Global Coherence: When Every Agent Is Right and the Team Is Still Wrong - A Local-to-Global Semantic Foundation for Multi-Agent Collaboration

The paper introduces the global coherence problem, where AI agents make locally valid decisions that collectively lead to an invalid outcome due to shared state failures. It presents the Observation‑Aliasing Impossibility Theorem, establishing that a policy can guarantee a valid action only when all indistinguishable worlds share an admissible action, and shows that even with additional reasoning, roles, messages, or samples, the missing distinction cannot be recovered. The authors propose a local‑to‑global runtime semantics framework and conduct nine studies demonstrating that missing global state cannot be substituted by local intelligence.

By Xin Heng
arXiv AI
Sep 2

trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories

The paper "trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories" examines the limitations of outcome-only evaluation for large language model agents. Using a deterministic tool‑using support‑desk environment with a scripted oracle policy and a fault injector, the authors compare five different judging approaches—programmatic rules, outcome‑only, step‑rubric at two model sizes, and a self‑consistency ensemble—on metrics such as detection, step localisation, fault typing, calibration, and cost across 400 trajectories. The study finds that outcome‑only judges miss many silent faults and generate false positives, while step‑rubric judges achieve higher recall with no false alarms but at greater cost, and that none of the judges read the final reply, allowing fabricated promises to evade detection. "whyItMatters":"The findings highlight that current production‑default outcome‑only evaluations can overlook critical failures in agent behavior, underscoring the need for more nuanced, step‑level judging methods to ensure reliable LLM agent performance."

By Hadi Mohammadi
arXiv AI
Sep 25

Evaluation of Multi-Turn Consistency in LLM Agents: Survival Analysis and Failure-Rationale Taxonomy

The study evaluates how large language model agents maintain consistency over extended interactions by simulating a 20‑step delayed‑gratification task. Researchers ran 84,540 trajectories across eight model families, using survival analysis to track when agents first claim a reward and discrete‑time hazard regression to assess how factors like social visibility, persona stressors, and deliberation policy affect failure risk. They also developed a seven‑category taxonomy from 13,780 deliberation traces, revealing that early failures are impulse‑driven, later ones are fatigue‑ or cost‑benefit‑framed, and public settings elicit norm‑oriented justifications; longer deliberation correlates with higher intra‑rationale contradictions, challenging assumptions about reasoning depth and consistency.

By Igor Bogdanov, Olga Manakina, Chung-Horng Lung