arXiv AI By Jaineet Shah

Causal Agent Replay: Counterfactual Attribution for LLM-Agent Failures

Read the original on arXiv AI →

arXiv:2606. 08275v1 Announce Type: cross Abstract: When an LLM agent fails -- issues a refund it should not have, calls the wrong tool, leaks data -- existing tooling answers what happened (observability) or whether it passed (evaluation), but not which step caused the failure.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 21

Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay

arXiv:2608. 19760v1 Announce Type: cross Abstract: Audited against causal ground truth from executed replay in a single-agent tool environment (ALFWorld), none of the step-level credit signals used to train LLM agents -- LLM-judge scores, outcome-conditioned logprob ratios, or the policy's own confidence -- identifies which steps causally matter better than chance.

By Haiyue Zhang
arXiv AI
Sep 2

trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories

The paper "trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories" examines the limitations of outcome-only evaluation for large language model agents. Using a deterministic tool‑using support‑desk environment with a scripted oracle policy and a fault injector, the authors compare five different judging approaches—programmatic rules, outcome‑only, step‑rubric at two model sizes, and a self‑consistency ensemble—on metrics such as detection, step localisation, fault typing, calibration, and cost across 400 trajectories. The study finds that outcome‑only judges miss many silent faults and generate false positives, while step‑rubric judges achieve higher recall with no false alarms but at greater cost, and that none of the judges read the final reply, allowing fabricated promises to evade detection. "whyItMatters":"The findings highlight that current production‑default outcome‑only evaluations can overlook critical failures in agent behavior, underscoring the need for more nuanced, step‑level judging methods to ensure reliable LLM agent performance."

By Hadi Mohammadi
arXiv AI
2d ago

When Do Causal World Models Help Modular LLM Agents

The paper introduces FedCausalCompose, a causal world‑model framework designed for modular large‑language‑model agents that interact with distinct services such as order, payment, inventory, and shipment. It demonstrates that standard observational world models suffer from irreducible interventional errors when unblocked back‑door paths exist, whereas incorporating intervention‑response evidence improves interface recovery and can outperform non‑causal baselines when coverage and local mechanism errors are controlled. Experiments show that causal interfaces are most beneficial in structured tool environments with clear API signatures, while they provide little advantage in dialogue or narrative settings unless the causal information becomes directly relevant to the agent’s decision making.

By Xinyuan Song, Zekun Cai
arXiv AI
2d ago

Verify Claims, Not Scores: Evidence-Based Verification of Modular Agents

The paper proposes a claim‑specific verification audit for modular agents that replaces aggregate task scores with evidence‑based evaluations. Each agent conclusion is recorded with supporting evidence and classified as supported, unsupported, unresolved, or not evaluated, along with the boundary of validity. The audit employs three tools—oracle policies, perfect component replacements, and verifier‑score tests—to trace value changes, locate lost value, and assess verifier effectiveness, demonstrated on a portfolio‑allocation agent in a synthetic market.

By Ali Atiah Alzahrani