The paper investigates a critical flaw in AI coding-agent systems such as Claude Code, Codex CLI, and Cursor, where the action approved by a human is not the same as the action executed by the harness. It introduces the concept of Approval Laundering, categorizing six systematic failure modes—Scope, Argument, Temporal, Tool, Delegation, and Semantic laundering—and demonstrates these failures through controlled experiments. The authors propose an Approval Token mechanism that mitigates some laundering types but leaves others unaffected, highlighting the limitations of current enforcement strategies.
By Yang Wang
arXiv:2609.08789v1 Announce Type: cross
Abstract: Frontier AI developers publish safety frameworks that commit them to evidencing whether their models are dangerous. The European Union and California...
By Louis Yiven Zhu
arXiv:2606. 18021v1 Announce Type: new Abstract: AI systems deployed in legal workflows hallucinate at rates that aggregate metrics report at ~52%, but this average conceals where errors concentrate and in which direction they run, leaving compliance officers without an actionable signal for trustworthy deployment.
By Lalit Yadav, Akshaj Gurugubelli
The paper introduces Publication Authority, a single-use, non-transferable capability that ensures AI-assisted claims can be independently challenged by providing a machine-readable, falsifiable publication record. It presents the PAC-2026 protocol, evaluates its fourth bounded semantic freeze (SF-4), and demonstrates through extensive modeling that the system enforces strict obligations on evidence, authorization, and lifecycle continuity. The study confirms internal coherence, bounded safety, and fault sensitivity, though it does not address factual truth or field efficacy.
By Torsten Olivi Tiltack, Yifei Dong, Kun Yu, Xu Wang, Wei Liu, Jianlong Zhou, Ren Ping Liu, Fang Chen
arXiv:2607. 25364v1 Announce Type: new Abstract: Tool-using agents expose structured calls but commonly attach free-form rationales.
By Genliang Zhu (Accentrust, Georgia Institute of Technology), Chu Wang (Accentrust, University of Illinois Urbana-Champaign)
BodhiPromptShield is a policy‑aware mediation layer for LLM agent pipelines that detects sensitive text spans before they propagate, replacing them with typed placeholders, semantic abstractions, or secure tokens and restoring them only at authorized execution boundaries. In evaluations on AI4Privacy, PrivacyLens, and AgentDojo datasets, the system reduces identifier exposure to 7.4% and 1.8% respectively, and limits exact identifier leakage in final actions to 2.1–3.1%. While mediation preserves factual content according to automated metrics, human annotations show a significant drop in inferability from 100% to 24–53%, indicating the need for human validation of semantic‑leakage measures.
By Bo Ma, Jinsong Wu, Weiqi Yan
arXiv:2608. 11274v1 Announce Type: cross Abstract: The dominant paradigm treats AI safety as a property to be instilled during model training via RLHF, DPO, or Constitutional AI.
By Albus W. Ng, Yi Han, Jusheng Zhang, Wenhao Wang
arXiv:2606. 04602v1 Announce Type: new Abstract: As agents grow more capable, legal-domain LLM agents promise to turn document-heavy matters into reviewable work products -- yet reliable deployment faces three obstacles: no large-scale evidence on how today's strongest model-and-harness combinations behave on end-to-end legal matters; no agent architecture adapted to the legal vertical, only general-purpose harnesses; and, in a setting that keeps shifting with new facts, authorities, and deadlines, no mechanism for systems to learn from their own outcomes.
By Hejia Geng, Leo Liu
arXiv:2608.28596v1 Announce Type: new
Abstract: Large language model (LLM) agents are increasingly embedded in scientific workflows for literature analysis, drafting, and review. Existing systems adv...
By Nidhi Jha, Siddharth Chaudhary, Ajinkya Kulkarni
arXiv:2607. 10487v1 Announce Type: cross Abstract: LLM agents can commit durable effects from authority evidence that was valid earlier in execution: a DOM snapshot, approval epoch, version witness, branch token, or worker result.
By Igor Santos-Grueiro
arXiv:2606. 12320v1 Announce Type: new Abstract: Enterprise security was built to govern data boundaries: the protected surface was data at rest and in transit, and the controls -- access control, data-loss prevention, perimeter inspection -- governed crossings of that boundary.
By Krti Tallam
The paper introduces Aegis, a runtime governance system for agentic AI that treats model outputs as action proposals and mediates them through a trusted decision layer before tool execution. Aegis evaluates proposals against active policy, resolves provenance server‑side, fails closed under uncertainty, and routes selected cases through a Senate‑style settlement process. In a sandbox evaluation across 6,300 rows, Aegis prevented all governed mock‑tool applications and risky side‑effect completions, preserving provenance and quorum evidence for all settled cases.
By Adam Mazzocchetti