arXiv AI

Where Does Exactly-Once Live? Model, Harness, and Tool-Contract Effects on Duplicate Side Effects in LLM Agents

The paper investigates where exactly‑once semantics should be enforced for tool‑using agents—within the model, the agent harness, or the tool contract—by evaluating 25,930 episodes across nine models, three harnesses, two contract variants, and fifteen recovery conditions. Using the LIMBO sandbox, the study shows that when an immediate read‑back is available, frontier models rarely duplicate lost acknowledgements, whereas weaker models do; when read‑back is unavailable, the contract’s idempotency keys explain most duplicate behavior. The authors prove that verification‑only policies cannot guarantee exactly‑once under late commits without bounded in‑flight time, and that waiting only helps when delays are short and predictable. whyItMatters":"The findings clarify that enforcing exactly‑once semantics largely depends on the tool contract and fault type, guiding designers on where to focus reliability mechanisms for LLM agents."

arXiv AI
Sep 2

Invalidation Contracts for Cross-Episode Agent Memory

The paper proposes invalidation contracts to manage cached recovery suggestions in LLM agents, attaching version stamps and cacheability hints to each suggestion so stale entries can be evicted without trial and error. The protocol separates realized savings into validity (protocol‑dependent) and compliance (planner‑dependent), showing that row‑level invalidation can significantly improve first‑try compliance and recover a substantial portion of token costs across multiple models, while table‑level invalidation can be detrimental. The study evaluates the approach across seven models, three serving paths, two domains, and about 9,400 episodes, demonstrating deterministic validity and high eviction precision.

By Michael Wu, Arquimedes Canedo
arXiv AI
Aug 28

Agent Mesh: Reliability Primitives for Non-Idempotent Agent Delegation - Identity Adequacy and Evidence Adequacy

The paper reports a failure study of a production agentic software‑delivery platform, analyzing 147 incidents across 81 runs. It shows that the standard reliability primitives—retry, timeout, and error‑rate circuit breaking—fail in practice, leading to costly loops, false trips, and blocked work. The authors identify two cross‑cutting causes—identity adequacy and evidence adequacy—and propose seven new reliability primitives that enforce reliability at the delegation level.

By Mazhar Shaikh, Anurag Rajkumar Bombarde, Harshal Pathak
arXiv AI
6d ago

Stale Does Not Mean Unsafe: Guard Precision for Tool-Using LLM Agents under Infrastructure State Races

The paper investigates how tool‑using language‑model agents can safely commit changes to infrastructure when external state may change between read and commit. By distinguishing invalidating races from predicate‑preserving and irrelevant ones, the authors evaluate three commit‑time guard granularities—global epoch, read‑set version, and semantic commit predicate—using a deterministic simulator and three quantized model families. The study finds that only the complete predicate guard consistently eliminates unsafe commits, while freshness‑based guards block a large proportion of benign races and model‑side signals fail to replace precise semantic enforcement.

By Zihao Zheng, Jiayu Long, Baichuan Li, Junyi Yao
arXiv AI
1d ago

Boundary-State Control for Tool-Using Language-Model Agents: Commit-Time Consistency under State Drift

The paper introduces BSC‑R, a deterministic effect‑boundary mechanism that ties a single‑use commit authorization to the specific action and the semantic state that justified it, aiming to close the proposal‑to‑commit gap in tool‑using language‑model agents. Experiments on 2,847 AgentDojo episodes and 10,302 frozen proposals show that BSC‑R preserves the agent’s original behavior while rejecting unauthorized changes, and further tests on a boundary‑drift experiment and the CONTINUITY suite demonstrate high success rates in valid contexts and robust handling of replay and ambiguous cases. However, broader testing reveals that BSC‑R still allows a 25% invalid‑effect commit rate in a larger attack set, indicating that it provides scoped, not universal, safety.

By Wesley Shu
arXiv AI
Aug 26

Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watched Fail

The study investigates why small language model agents tend to repeat a tool call that just failed. By recording the failed call and its error message in the transcript, the authors measure a negative corrective gain—agents are more likely to repeat the failed action, with a drop of about 1.03 nats per token. The problem is traced to the harness design rather than the model’s understanding of errors, and the authors show that replacing the verbatim call with a runtime-generated description of the failure can reduce this backfiring effect by 76%.

By Esmail Gumaan