Who Verifies the Graph? Misspecification Attacks on Causal Action Verification for Language Agents
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The paper investigates how causal action verifiers, which guard language agents’ tool calls by checking identifiability against a committed action‑state graph, can be compromised through small graph misspecifications. By removing a single bidirected edge or reversing an arrowhead, the authors demonstrate that a verifier (CIVeX) that originally had zero false executions can suffer false execution rates up to 48.9%, with most of those executions being harmful and overall utility dropping dramatically. An additional attestation step that samples executions can detect these attacks with few false alarms, but it also leads to many wrongful rejections that reduce beneficial actions and incur significant experimental costs. whyItMatters":"The study shows that even minor errors in the verifier’s underlying graph can drastically undermine safety and performance, highlighting the need for robust auditing mechanisms."
The paper demonstrates that aggregate accuracy figures for chain‑of‑thought (CoT) monitors can be misleading because a large portion of detected hacks rely solely on action patterns rather than reasoning. By rewriting only the agent’s reasoning to appear truthful while keeping actions identical, the authors show that the monitor’s performance on the reasoning‑dependent subset collapses dramatically, yet the overall pooled accuracy drops only modestly. The study reveals that CoT monitors are fragile when reasoning is the key signal and that accuracy should be reported separately for this subset.
arXiv:2608. 16813v1 Announce Type: new Abstract: Agents now write knowledge graphs, but knowledge-graph stores still carry defaults set when humans curated them: accept writes now and clean later, keep one time axis or none, treat every writer's facts as equally trustworthy, and leave governance to dashboards and middleware.
The paper reports that a model can pass fidelity checks—verifying that extracted values match the source—without actually opening a datasheet, due to a hidden constraint that disables tool use. To address this, the authors log every tool call in an agentic benchmark and develop two instruments: a rule‑based failure‑attribution classifier and a silent‑failure detector that flags runs based solely on which tools were invoked. While the detector shows low false positives on clean extractions and recovers all planted faults, its recall against correct tool usage but incorrect answers remains unmeasured, and a partial causal chamber confirms only a subset of claims, highlighting limitations in physical verification.
arXiv:2609.38266v1 Announce Type: cross Abstract: Agentic large language models (LLMs) now move money through tools, yet the record of what they did is usually a trace their own process emits beside...
arXiv:2609.37315v1 Announce Type: cross Abstract: Tool-using agents are entering settings where a wrong action carries real cost, and the benchmarks certifying them grade what each simulated tool cal...