arXiv:2607. 11751v1 Announce Type: cross Abstract: As multi-agent, tool-using LLM systems are deployed, a common safety net is a runtime monitor that checks each message, tool call, or step on its own.
By Yibo Hu, Ren Wang
The paper demonstrates that aggregate accuracy figures for chain‑of‑thought (CoT) monitors can be misleading because a large portion of detected hacks rely solely on action patterns rather than reasoning. By rewriting only the agent’s reasoning to appear truthful while keeping actions identical, the authors show that the monitor’s performance on the reasoning‑dependent subset collapses dramatically, yet the overall pooled accuracy drops only modestly. The study reveals that CoT monitors are fragile when reasoning is the key signal and that accuracy should be reported separately for this subset.
By Shikhar Shiromani, Leo Richter
arXiv:2607. 07474v1 Announce Type: cross Abstract: Agentic red-teaming benchmarks report whether an injected agent was compromised as a single bit: the attack succeeded, or it did not.
By Harry Owiredu-Ashley
The paper reports that a model can pass fidelity checks—verifying that extracted values match the source—without actually opening a datasheet, due to a hidden constraint that disables tool use. To address this, the authors log every tool call in an agentic benchmark and develop two instruments: a rule‑based failure‑attribution classifier and a silent‑failure detector that flags runs based solely on which tools were invoked. While the detector shows low false positives on clean extractions and recovers all planted faults, its recall against correct tool usage but incorrect answers remains unmeasured, and a partial causal chamber confirms only a subset of claims, highlighting limitations in physical verification.
By Qing Ye, Meng-Hsuan Lin
Causal action verifiers gate an agent's state-changing tool calls by checking whether each proposed intervention is identifiable against a committed action-state graph, and they issue a certificate th...
The paper investigates how causal action verifiers, which guard language agents’ tool calls by checking identifiability against a committed action‑state graph, can be compromised through small graph misspecifications. By removing a single bidirected edge or reversing an arrowhead, the authors demonstrate that a verifier (CIVeX) that originally had zero false executions can suffer false execution rates up to 48.9%, with most of those executions being harmful and overall utility dropping dramatically. An additional attestation step that samples executions can detect these attacks with few false alarms, but it also leads to many wrongful rejections that reduce beneficial actions and incur significant experimental costs.
whyItMatters":"The study shows that even minor errors in the verifier’s underlying graph can drastically undermine safety and performance, highlighting the need for robust auditing mechanisms."
By Fabio Rovai
arXiv:2606. 09315v1 Announce Type: cross Abstract: BCI-to-agent pipelines turn decoded neural activity into an authorization channel for tool-use agents, exposing a new attack surface we call \emph{brain-prompt injection}: signal-side perturbations, context-only injections, and adaptive dual-decoder attacks can all change the routed action while EEG-side or text-side monitors remain blind.
By Jianwei Tai
arXiv:2607. 19321v1 Announce Type: new Abstract: As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted.
By Lena Libon, Ben Rank, Jehyeok Yeon, David Schmotz, Jeremy Qin, Daniel Donnelly, Derck Prinzhorn, Maksym Andriushchenko
arXiv:2608.29942v1 Announce Type: cross
Abstract: The key limitation of current state-of-the-art influence-based guardrails is that they do not reliably distinguish a legitimate, user-authorized acti...
By Tanzim Ahad, Ismail Hossain, Md Jahangir Alam, Sai Puppala, Syed Bahauddin Alam, Sajedul Talukder
arXiv:2607. 24893v1 Announce Type: cross Abstract: Multi-agent LLM systems can be attacked by a payload that no single agent ever holds in full: a poisoned tool hides encrypted fragments in its observations, spreads them across several agents, and an external step reassembles and executes them after the run.
By Diego Fernandez Arias, Dev Prashant Mistry, Ren Wang, Yibo Hu
The paper investigates a critical flaw in AI coding-agent systems such as Claude Code, Codex CLI, and Cursor, where the action approved by a human is not the same as the action executed by the harness. It introduces the concept of Approval Laundering, categorizing six systematic failure modes—Scope, Argument, Temporal, Tool, Delegation, and Semantic laundering—and demonstrates these failures through controlled experiments. The authors propose an Approval Token mechanism that mitigates some laundering types but leaves others unaffected, highlighting the limitations of current enforcement strategies.
By Yang Wang
arXiv:2607. 06596v1 Announce Type: cross Abstract: Trusted monitoring is a central defense in AI control: a cheaper trusted model scores an untrusted model's actions for sabotage, and the most suspicious are audited or deferred.
By Lucas Pinto