arXiv AI By Lu Yan, Xuan Chen, Xiangyu Zhang

WIRE: Profiling Witnessed Within-Policy Instruction Collisions in LLM Agents

Read the original on arXiv AI →

arXiv:2605. 27784v2 Announce Type: replace Abstract: LLM agents are governed by long-lived prompt policies, where individually reasonable stand- ing rules can jointly govern the same pre- generation state.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 24

Calibrating Criterion Revision in LLM Agents: Failure Modes and a Trace-Anchored Protocol

The paper introduces a framework for evaluating how large language model agents revise their success criteria after failures, defining five non‑compensatory conditions that must be met for a criterion revision to be considered valid. Using the CMB‑0.1 protocol, the authors test twelve cross‑domain scenarios across four system configurations, finding that no model trial satisfies all five conditions and highlighting specific failure modes such as zero‑state reconstruction and inadequate intervention sensitivity. They propose a more stringent trace‑anchored CMB‑0.4 protocol to better isolate and measure criterion revision in future studies.

By Guodong Xu
arXiv AI
Sep 3

Public-Sharing Labels and Verbatim Field Egress in an MCP-to-A2A Agent Configuration: A Controlled Multi-Model Study

The study evaluates safety properties of a controlled MCP-to-A2A agent configuration by measuring verbatim field egress across ten record scenarios under three labeling conditions (CONFIDENTIAL, no header, PUBLIC – OK TO SHARE). Using four models repeated four times each, 480 trials were conducted, and the results show that adding a PUBLIC header is descriptively linked to higher verbatim egress, with the effect varying strongly by model. The study releases code, byte‑pinned traces, and an offline analysis pipeline as a public artifact.

By Arpan Kumar Mahapatra
arXiv AI
Aug 20

One Gate Is Not Enough: Composing Stateful Pre-Action Controls for Agentic AI

The paper investigates how multiple pre‑action controls—authority, resource, and evidence gates—interact in agentic AI systems. It formalizes remediation‑induced control coupling, showing that remediation can invalidate earlier judgments and that the order of remediation matters. The authors propose a remediate‑and‑regate protocol to restore soundness, analyze non‑commuting remediation operators, and demonstrate the approach on a deterministic open‑data artifact with three published engines.

By Gaston Besanson