arXiv AI

Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation

arXiv:2607. 24054v1 Announce Type: new Abstract: A correct answer can conceal why an agent succeeded.

arXiv Machine Learning
Jul 31

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

arXiv:2607. 09306v3 Announce Type: replace-cross Abstract: Behavioural auditing asks whether a language model behaves as it claims, but detection scores are reported without separating two targets: whether a reply was produced under a behaviour-inducing condition (exposure) and whether the behaviour surfaced in it (manifestation).

By Kwan Soo Shin
arXiv AI
Aug 28

Invocation-Level Reliability of Tool-Using Agents

The paper investigates the reliability of tool‑using agents, focusing on two failure modes: selecting the wrong tool and constructing incorrect arguments. It introduces a correct‑invocation rate metric to distinguish these errors and evaluates five open‑weight models on multi‑step tasks up to depth 8, finding that by depth 6 about 70% of a model’s clean‑context capability is lost due to earlier mistakes. The study reveals that exact‑match scoring against a fixed gold trajectory forces severity and recovery parameters to extreme values, and proposes a conditional‑on‑state scoring remedy that yields more realistic severity estimates.

By Afiya Noorain, Subhranshu Mohanty, Amritesh Banerjee, Abhijit Dasgupta
arXiv AI
Sep 3

ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent Evaluations

The paper introduces ClaimReceipt, a specification and verifier that checks whether a claim in an agent evaluation can be recomputed from retained evidence (sufficiency) and whether the evidence covers the entire experiment set (coverage). Using the CR‑2 verifier on 1,392 historical records, the authors demonstrate accurate reproduction of audit verdicts, non‑redundant field groups, and zero false positives on semantic faults. In a prospective CR‑3 run, the system correctly flags missing receipts and preserves coverage when private evidence is withheld, while adding minimal overhead to inference time and transaction size.

By Peiying Zhu, Sidi Chang
arXiv AI
4d ago

When Tools Silently Lie: Evaluating and Mitigating Blind Compliance in Tool-Augmented Data Agents

The paper introduces ToxicBench, a benchmark designed to evaluate how tool‑augmented data agents handle incorrect tool outputs. By pairing clean and poisoned observations across numerical, label, schema, and retrieval errors, the authors assess both the checking process and the final answer adoption. In a 118‑task GPT evaluation, poisoning reduces task success by 26–39 percentage points, revealing that repeated poisoning leads to wrong-answer adoption even after checking, while ordinary retries help only under one‑shot poisoning. Human annotations on 200 trajectories confirm the scoring system’s reliability, showing 96% agreement with task success and supporting the benefits of retries and audit‑based adoption.

By Zifu Tao, Changqing Yin
arXiv Computation and Language
6d ago

Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline

The paper evaluates a production text‑to‑SQL pipeline that uses an LLM as a judge, finding that the deployed gpt‑4o‑mini judge agrees with human annotators only weakly (Cohen’s kappa 0.04 on a disagreement‑enriched set and 0.42 on a random spot‑check). The authors identify a specific failure mode, GRADE‑HALLUCINATION, responsible for most over‑flags, and demonstrate that a self‑hosted Qwen3.6‑27B model achieves substantially higher agreement (kappa 0.72) at a much lower cost. They also show that ensembling judges does not improve performance, and that their audit method flags a significant portion of out‑of‑domain SQLs as potential issues.

By Haowei Liu, Hsin-Tai Wu, Yi Fang