The paper investigates where exactly‑once semantics should be enforced for tool‑using agents—within the model, the agent harness, or the tool contract—by evaluating 25,930 episodes across nine models, three harnesses, two contract variants, and fifteen recovery conditions. Using the LIMBO sandbox, the study shows that when an immediate read‑back is available, frontier models rarely duplicate lost acknowledgements, whereas weaker models do; when read‑back is unavailable, the contract’s idempotency keys explain most duplicate behavior. The authors prove that verification‑only policies cannot guarantee exactly‑once under late commits without bounded in‑flight time, and that waiting only helps when delays are short and predictable.
whyItMatters":"The findings clarify that enforcing exactly‑once semantics largely depends on the tool contract and fault type, guiding designers on where to focus reliability mechanisms for LLM agents."
By Jiapeng Li
arXiv:2607. 20972v1 Announce Type: new Abstract: Coding agents ship with one kind of memory: documents.
By Swapnanil Saha
arXiv:2607. 09691v1 Announce Type: cross Abstract: A modern coding agent can hold an entire repository in its context window.
By Brian Sam-Bodden
The paper investigates how agents that inherit consolidated memory can mistakenly rely on constraints that were once true but have since been superseded by newer records. By modeling supersession explicitly and limiting verification to two records, the study shows that native allocation often leads to stale-consistent decisions, while reallocating one verification slot to the critical provenance path markedly improves consistency. The results suggest that memory systems may need separate freshness or supersession signals beyond relevance to avoid such errors.
By Kazuki Nakayashiki
arXiv:2608. 04278v1 Announce Type: cross Abstract: Coding agents increasingly work across sessions, but prose notes can preserve a conclusion without the program state that supported it.
By Hwai-Jung Hsu, Cheng-Jan Chi, Hanna Everett
StateTape introduces a new framework for long‑horizon coding agents that rewrites the agent’s context as the code repository changes, rather than letting the context grow with every observation. It models the repository as a symbol‑level code graph, using a tape to mark symbols altered by each write and a manager model to resolve stale records. The authors provide theoretical analysis, a new benchmark called TraceBench, and empirical results showing higher resolve rates across six agents and three edit‑heavy benchmarks with minimal computational overhead.
By Ziyang Yu, Liang Zhao, Bowen Zhu, Hasibul Haque
arXiv:2607. 10569v1 Announce Type: cross Abstract: Modern coding agents expose multiple tool surfaces -- IDE primitives, bash, and Model Context Protocol (MCP) code-execution -- and the field has shipped three contradictory claims about which one matters.
By Hong Yang, Qi Yu, Travis Desell
The paper introduces a formal framework for agent harnesses that guarantees termination, prevents drift, and enforces spend limits through bounded loops, gates, and repair relations. It proves that these guarantees hold even with repair budgets and demonstrates the effectiveness of the system by identifying vacuous gates and achieving low false‑accept rates in a 69‑loop catalogue. The authors provide an instrumented implementation and a held‑out mutant corpus to validate gate correctness.
By Varun Pratap Bhardwaj, Garima Singh, Arun Pratap Bhardwaj
arXiv:2608. 12599v1 Announce Type: new Abstract: Multi-turn dialogues let users revoke constraints as easily as impose them, but revocation does not reliably take effect: models keep enacting withdrawn requirements (occasionally beneath comments asserting their removal), a failure we call \emph{behavioral relapse}, or revocation inertia.
By Haoyuan Zhu
arXiv:2607. 28545v2 Announce Type: replace-cross Abstract: Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics, logs, traces, and source code, starting from ambiguous user-facing reports, often hours after the incident began.
By Albert Gong, Kyuseong Choi, Abhineet Agarwal, Jason Schechner, Ryan Huang, Raj Agrawal, Anish Agarwal, Raaz Dwivedi
arXiv:2608. 13228v1 Announce Type: new Abstract: Agent harnesses combine retrieval, routing, state, provenance, and verification, but locally successful components may disagree on shared state.
By Saveliy Batruin
The paper investigates how altering the harness—specifically the way a coding agent manages context and tool outputs—affects performance when the underlying model and task remain unchanged. Two harness configurations were compared on three coding benchmarks: a control that preserves the full conversation in order, and a treatment that mechanically shortens older tool results to keep the context tight. Across all benchmarks, the treatment increased the mean per‑task fail‑to‑pass fraction and, in some cases, the number of complete solutions, demonstrating that the harness itself can significantly influence a frozen model’s effectiveness.
By Sydney Lewis