arXiv AI By Joongho Ahn, Moonsoo Kim

From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents

Read the original on arXiv AI →

arXiv:2607. 08028v1 Announce Type: new Abstract: Enterprise large language model (LLM) applications often begin as prototypes whose behavior is carried by prompts and retrieval context.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 28

Persona-Execution Separation: An Architecture Pattern for Evolving LLM Agents under Execution Audit

The paper introduces Persona‑Execution Separation (PES), an architecture pattern that splits an LLM agent’s persona—its instructions, tone, and self‑presentation—from its execution—stateful, auditable work—by placing them in distinct trust domains linked through a governed contract bridge. PES allows the persona to evolve freely while keeping execution traceable, using mechanisms such as an approval matrix, data‑loss‑prevention exceptions, and continuous identity. A pilot implementation on a regulated digital‑employee platform demonstrated that PES successfully decouples persona drift from execution audit, preventing re‑validation or fingerprinting of hard‑asserted fields.

By Yisen Xi
arXiv AI
2d ago

Actions with Receipts: Jointly Binding Claims, Evidence, and Execution for Replayable Tool-Agent Auditing

The paper introduces a claim‑anchored execution contract that binds a tool‑using agent’s emitted claim to its exact source span, the ordered execution prefix that produced it, and the source version and access state observed. Each receipt contains deterministic anchors, source identifiers, offsets, hashes, quotes, and a domain‑separated execution commitment, allowing a verifier to reconstruct these bindings before semantic or task labels are joined. The contract defines seven independently testable properties and demonstrates high detection rates against cross‑object attacks, with strong performance on conflict‑aware support guard evaluations.

By Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang, Yina Sa, Daren Zha, Jun Xiao
arXiv AI
Sep 23

Trustworthy Agentic AI: Failure Modes, Mitigation Strategies, and a Lifecycle Framework for Autonomous LLM Systems

The paper discusses the trustworthiness of agentic AI systems built on large language models, highlighting new security and operational risks such as indirect prompt injection, memory contamination, and cross‑session data leakage. It categorizes failure modes, reviews mitigation strategies—including instruction hierarchies, context isolation, and constrained tool use—and introduces the Trustworthy Agent Development Lifecycle (TADL), a six‑phase framework for specification, design, training, evaluation, deployment, and monitoring. The authors note that TADL has not yet been empirically validated but offers a structured foundation for developing more secure and accountable agentic systems, and they call for improved benchmarks and future research priorities.

By Fayeq Jeelani Syed, Rehan Ahmad, Ali Al Bataineh, Aakriti Adhikari
arXiv AI
Aug 24

ClawSentry: A Progressive Multi-Tier Security Monitor for Safeguarding Autonomous LLM Agents

ClawSentry is an open‑source, framework‑agnostic security supervision gateway designed to protect autonomous large language model (LLM) agents from progressive risks that can arise at four points in the agent control loop: skill admission, invocation‑time intent, execution‑time effect, and post‑action consequence. It introduces a multi‑tier decision engine—deterministic L1, rule‑anchored L2, and read‑only L3—alongside a First‑Use Skill Package Review (FSPR) and an Agent Harness Protocol (AHP) that applies a single policy across multiple agent runtimes without modifying their internals. Evaluation on SkillInject and the SkillsSafety benchmark shows that ClawSentry significantly reduces contextual adversarial skill risk (ASR) while maintaining high task success rates (TSR).

By Kai Wang, Zeming Wei, BiaoJie Zeng, Chang Jin, An Wang, Xiaokun Luan, Zhixiao Lin, Jingjing Qu, Xia Hu, Xingcheng Xu