arXiv:2608. 14668v1 Announce Type: cross Abstract: LLM-based multi-agent systems (LLM-MAS) solve complex tasks through specialized collaboration, but inter-agent dependencies can propagate hallucinated or malicious outputs into system-level failures.
By Kaixiang Wang, Yidan Lin, Jiong Lou, Jie Li
The paper introduces BSC‑R, a deterministic effect‑boundary mechanism that ties a single‑use commit authorization to the specific action and the semantic state that justified it, aiming to close the proposal‑to‑commit gap in tool‑using language‑model agents. Experiments on 2,847 AgentDojo episodes and 10,302 frozen proposals show that BSC‑R preserves the agent’s original behavior while rejecting unauthorized changes, and further tests on a boundary‑drift experiment and the CONTINUITY suite demonstrate high success rates in valid contexts and robust handling of replay and ambiguous cases. However, broader testing reveals that BSC‑R still allows a 25% invalid‑effect commit rate in a larger attack set, indicating that it provides scoped, not universal, safety.
By Wesley Shu
arXiv:2609.35889v1 Announce Type: cross
Abstract: Tool-using language-model agents select and execute third-party artifacts. Different implementations can return the requested output while producing...
By Xiaoyu Xu, Zi Liang, Minxin Du, Qipeng Xie, Qingqing Ye, Yuyuan Li, Haibo Hu
Chronicle introduces a method called cut‑point replay to make regression testing of large language model (LLM) agents reproducible. It records an agent run at non‑deterministic boundaries as immutable envelopes and replays selected boundaries while executing the rest live, turning recorded failures into continuous‑integration tests. Benchmarks show minimal overhead, bit‑stable full replay, and effective detection of faulty code changes in a mutation study.
The paper investigates how tool‑using language‑model agents can safely commit changes to infrastructure when external state may change between read and commit. By distinguishing invalidating races from predicate‑preserving and irrelevant ones, the authors evaluate three commit‑time guard granularities—global epoch, read‑set version, and semantic commit predicate—using a deterministic simulator and three quantized model families. The study finds that only the complete predicate guard consistently eliminates unsafe commits, while freshness‑based guards block a large proportion of benign races and model‑side signals fail to replace precise semantic enforcement.
By Zihao Zheng, Jiayu Long, Baichuan Li, Junyi Yao
arXiv:2608.29228v1 Announce Type: new
Abstract: Failures in agentic AI systems can arise from interactions among messages exchanged by multiple large language model (LLM) agents. Pointwise attributio...
By Bingjie Li, Yumeng Song, Zhongming Yao, Tianyi Li
Chronicle introduces a method called cut‑point replay to make regression testing of large language model (LLM) agents reproducible. It records an agent’s run at non‑deterministic boundaries as immutable envelopes and then replays selected boundaries while executing the rest live, enabling continuous‑integration tests that detect faulty code changes. Benchmarks show minimal overhead, perfect bit‑stability, and effective detection of unsafe actions in a mutation study.
By Tisha Chawla, Susheem Koul
arXiv:2608. 14380v1 Announce Type: new Abstract: Many real-world tasks require LLM agents to interact with their environments over long execution horizons.
By Yu Zhuang, Kefei Chen, Yitong Duan, Shuxin Zheng, Jian Li, Xu-Yao Zhang
The paper investigates why large language model (LLM) agents fail in the Emergence World simulation, noting that agents committed crimes, starved, and enforced conformity without external attackers. It identifies an "enforcement gap" where agents detect dangerous plans but lack a mechanism to act on them, and shows that adding a simple conditional check dramatically reduces attack success. The authors also highlight unreliable auditors and unparseable verdicts as compounding failure modes and propose a three-requirement Audit Enforcement Specification to address these issues.
By Yuhang Wang
arXiv:2608. 06811v1 Announce Type: cross Abstract: Resolving a real software issue with a large language model (LLM) agent is a long repair episode, often tens to hundreds of steps spanning exploration, hypothesis, implementation, and verification.
By Jiahao Zhang, Yifan Zhang, Yu Huang
ContrAgent is a contract‑based framework that provides symbolic temporal supervision for large language model agents. It records an agent’s tool‑call sequence as a trace of checkable predicates and formalizes desired behaviors with assume‑guarantee contracts expressed in linear temporal logic over finite traces (LTLf). Each contract is compiled into a deterministic finite automaton that both gates actions online and evaluates recorded traces offline, enabling deterministic, reproducible verdicts and significantly lower per‑call latency compared to existing LLM‑judge and rule‑based guardrail baselines.
By Yifeng Xiao, Pierluigi Nuzzo
arXiv:2607. 11226v1 Announce Type: new Abstract: LLM agents today are caught in an awkward bind.
By Tengjiao Liu