arXiv AI

The Distributed Detectability Band Against Marginal-Preserving Attacks

arXiv:2606. 10456v1 Announce Type: cross Abstract: AI-control monitors score individual agent actions to detect misbehavior, but real harm can be distributed across many benign-looking steps, each individually below any per-step alarm.

arXiv Machine Learning
Sep 14

A False Average: Pooled CoT-Monitor Accuracy Conceals a Reasoning-Dependent Fragility

The paper demonstrates that aggregate accuracy figures for chain‑of‑thought (CoT) monitors can be misleading because a large portion of detected hacks rely solely on action patterns rather than reasoning. By rewriting only the agent’s reasoning to appear truthful while keeping actions identical, the authors show that the monitor’s performance on the reasoning‑dependent subset collapses dramatically, yet the overall pooled accuracy drops only modestly. The study reveals that CoT monitors are fragile when reasoning is the key signal and that accuracy should be reported separately for this subset.

By Shikhar Shiromani, Leo Richter
arXiv Computation and Language
Aug 31

Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction

The paper reports that a model can pass fidelity checks—verifying that extracted values match the source—without actually opening a datasheet, due to a hidden constraint that disables tool use. To address this, the authors log every tool call in an agentic benchmark and develop two instruments: a rule‑based failure‑attribution classifier and a silent‑failure detector that flags runs based solely on which tools were invoked. While the detector shows low false positives on clean extractions and recovers all planted faults, its recall against correct tool usage but incorrect answers remains unmeasured, and a partial causal chamber confirms only a subset of claims, highlighting limitations in physical verification.

By Qing Ye, Meng-Hsuan Lin
arXiv AI
3d ago

Who Verifies the Graph? Misspecification Attacks on Causal Action Verification for Language Agents

The paper investigates how causal action verifiers, which guard language agents’ tool calls by checking identifiability against a committed action‑state graph, can be compromised through small graph misspecifications. By removing a single bidirected edge or reversing an arrowhead, the authors demonstrate that a verifier (CIVeX) that originally had zero false executions can suffer false execution rates up to 48.9%, with most of those executions being harmful and overall utility dropping dramatically. An additional attestation step that samples executions can detect these attacks with few false alarms, but it also leads to many wrongful rejections that reduce beneficial actions and incur significant experimental costs. whyItMatters":"The study shows that even minor errors in the verifier’s underlying graph can drastically undermine safety and performance, highlighting the need for robust auditing mechanisms."

By Fabio Rovai
arXiv AI
Jun 9

Brain-Prompt Injection: A Route-Safety Audit for BCI-LLM Agents

arXiv:2606. 09315v1 Announce Type: cross Abstract: BCI-to-agent pipelines turn decoded neural activity into an authorization channel for tool-use agents, exposing a new attack surface we call \emph{brain-prompt injection}: signal-side perturbations, context-only injections, and adaptive dual-decoder attacks can all change the routed action while EEG-side or text-side monitors remain blind.

By Jianwei Tai
arXiv AI
Jul 29

Early Detection of Distributed Backdoors in Multi-Agent LLM Systems: A Characterization Study

arXiv:2607. 24893v1 Announce Type: cross Abstract: Multi-agent LLM systems can be attacked by a payload that no single agent ever holds in full: a poisoned tool hides encrypted fragments in its observations, spreads them across several agents, and an external step reassembles and executes them after the run.

By Diego Fernandez Arias, Dev Prashant Mistry, Ren Wang, Yibo Hu
arXiv AI
3d ago

Approval Laundering: Systematizing Approval--Execution Binding Failures in AI Coding-Agent Harnesses

The paper investigates a critical flaw in AI coding-agent systems such as Claude Code, Codex CLI, and Cursor, where the action approved by a human is not the same as the action executed by the harness. It introduces the concept of Approval Laundering, categorizing six systematic failure modes—Scope, Argument, Temporal, Tool, Delegation, and Semantic laundering—and demonstrates these failures through controlled experiments. The authors propose an Approval Token mechanism that mitigates some laundering types but leaves others unaffected, highlighting the limitations of current enforcement strategies.

By Yang Wang