arXiv AI By Haotian Chen, Bowen Ye, Yuning Zhang, Jingkun Yu

Auditing Action Settlement in LLM Agent Environments: Order, Progress, and Replay

Read the original on arXiv AI →

The paper introduces a typed snapshot‑settlement contract for auditing concurrent actions in large language model agent environments. It evaluates three properties—order sensitivity, useful progress, and replay consistency—across five settlement policies, using 28,800 exhaustive permutation trials and 2,160 scripted multistep episodes. Results show that joint policies are spatially order‑invariant with fixed priorities, but conservative rejection only completes 31.25% of agents in a six‑agent doorway task compared to 90.28% for random tickets, while a full‑state journal audit successfully replays 156 checkpoints and rejects 1,332 constructed corruptions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
4d ago

Boundary-State Control for Tool-Using Language-Model Agents: Commit-Time Consistency under State Drift

The paper introduces BSC‑R, a deterministic effect‑boundary mechanism that ties a single‑use commit authorization to the specific action and the semantic state that justified it, aiming to close the proposal‑to‑commit gap in tool‑using language‑model agents. Experiments on 2,847 AgentDojo episodes and 10,302 frozen proposals show that BSC‑R preserves the agent’s original behavior while rejecting unauthorized changes, and further tests on a boundary‑drift experiment and the CONTINUITY suite demonstrate high success rates in valid contexts and robust handling of replay and ambiguous cases. However, broader testing reveals that BSC‑R still allows a 25% invalid‑effect commit rate in a larger attack set, indicating that it provides scoped, not universal, safety.

By Wesley Shu
Hugging Face Trending Papers
Sep 17

Chronicle: Cut-Point Replay for Regression Testing of LLM Agents

Chronicle introduces a method called cut‑point replay to make regression testing of large language model (LLM) agents reproducible. It records an agent run at non‑deterministic boundaries as immutable envelopes and replays selected boundaries while executing the rest live, turning recorded failures into continuous‑integration tests. Benchmarks show minimal overhead, bit‑stable full replay, and effective detection of faulty code changes in a mutation study.

arXiv AI
Sep 25

Stale Does Not Mean Unsafe: Guard Precision for Tool-Using LLM Agents under Infrastructure State Races

The paper investigates how tool‑using language‑model agents can safely commit changes to infrastructure when external state may change between read and commit. By distinguishing invalidating races from predicate‑preserving and irrelevant ones, the authors evaluate three commit‑time guard granularities—global epoch, read‑set version, and semantic commit predicate—using a deterministic simulator and three quantized model families. The study finds that only the complete predicate guard consistently eliminates unsafe commits, while freshness‑based guards block a large proportion of benign races and model‑side signals fail to replace precise semantic enforcement.

By Zihao Zheng, Jiayu Long, Baichuan Li, Junyi Yao