arXiv AI

Agent Step Value: Auditing Evaluator-Channel Reversals in Black-Box Agent Traces

arXiv:2607. 04419v3 Announce Type: replace Abstract: When evaluator-derived step rewards are pooled or compared across scoring channels, their sign is treated as transportable.

arXiv AI
Aug 21

Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay

arXiv:2608. 19760v1 Announce Type: cross Abstract: Audited against causal ground truth from executed replay in a single-agent tool environment (ALFWorld), none of the step-level credit signals used to train LLM agents -- LLM-judge scores, outcome-conditioned logprob ratios, or the policy's own confidence -- identifies which steps causally matter better than chance.

By Haiyue Zhang
arXiv Machine Learning
4d ago

Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation

The paper investigates how language‑model judges can make version‑dependent errors when evaluating upgraded agents. Using 35 public coding‑agent submissions, two customer‑service agents, and over a thousand expert‑labeled trajectories, the authors show that fixed judges often reject task‑conditioned error invariance and can incorrectly approve failed patches, especially as agent capability increases. Paired audits of current outputs reduce interval width only marginally, and the study concludes that independent human patch review is still necessary.

By Jiapeng Li
arXiv AI
Jun 30

SEVA: Self-Evolving Verification Agent with Process Reward for Fact Attribution

arXiv:2606. 29713v1 Announce Type: cross Abstract: Hallucination is the reliability bottleneck for LLM-based agents, and fact attribution verifiers are the last line of defense -- yet today's verifiers emit only opaque binary labels, leaving agents unable to self-correct and operators unable to audit.

By Aojie Yuan, Yi Nian, Haiyue Zhang, Zijian Su, Yue Zhao
arXiv Machine Learning
Jul 31

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

arXiv:2607. 09306v3 Announce Type: replace-cross Abstract: Behavioural auditing asks whether a language model behaves as it claims, but detection scores are reported without separating two targets: whether a reply was produced under a behaviour-inducing condition (exposure) and whether the behaviour surfaced in it (manifestation).

By Kwan Soo Shin
arXiv Computation and Language
Sep 24

Guides That Cause Actions: An Offline Study of Guide-Action Mutual Reinforcement in Multimodal Web Agents

The paper introduces WebMRE, an offline benchmark comprising 541 tasks and 5,293 steps extracted from WebArena trajectories, designed to provide deterministic scoring for web agents without live environments. It enables the first systematic study of how guide sentences and grounded actions reinforce each other, showing that jointly decoding a guide improves element selection accuracy and that the guide acts as a causal instruction channel. The authors fine‑tune models that outperform leading zero‑shot baselines on all offline metrics.

By Chengguang Gan, Yunhao Liang, QingHao Zhang, Shiwen Ni
arXiv Machine Learning
Sep 15

GRADE: Graph Representation of LLM Agent Dependency and Execution

The paper introduces GRADE, a graph-based representation of large language model (LLM) agent executions that captures both execution steps and their dependencies. By adding graded dependency edges—observed, declared, or inferred—to the trace, the authors evaluate how this dependency layer affects failure prediction across six corpora involving tool use, coding, and web tasks. Experiments show that the dependency block can improve prediction in some settings, but its effectiveness varies with the evaluation probe and corpus, and controlled experiments demonstrate that the observed structure is not merely a degree-matched artifact.

By Yue Zhao
arXiv AI
Sep 24

Same Outcome, Different Readout: What Does a Steerable Valence Direction in LLMs Represent?

The paper investigates what a steerable valence direction in large language models (LLMs) actually represents, focusing on a good‑bad outcome direction in a maze task. By using controlled interventions that separate the realized outcome from the informational history that led to it, the authors find that directions trained on one explicit outcome encoding transfer well to another, suggesting the readout is not tied to surface form. However, when the same outcome is achieved through announced versus unannounced histories, transfer performance drops sharply, indicating that the post‑event readout remains strongly conditioned on the earlier announcement. In a matched maze‑reinforcement‑learning run, the post‑RL direction becomes more predictive of reference‑MDP return and the policy depends more on it, yet the history dependence persists. These findings support a functional, value‑related interpretation of the direction but argue against identifying it with a history‑invariant scalar valence state.

By Weihan Li, Xinlei Chen, Yuhan Song, Xiaofeng Lin, Tianshi Zheng