arXiv Machine Learning By Bardia Mohammadi, Lars Klein, Aman Chadha, Akhil Arora, Laurent Bindschaedler

The Working Set of a Coding Agent: Coherence Debt in Repository-Scale Tasks

Read the original on arXiv Machine Learning →

arXiv:2608. 16630v1 Announce Type: cross Abstract: Repository-scale coding requires an agent to keep tests, imports, configuration, and migration rules consistent within a bounded context window.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 25

Where Does Exactly-Once Live? Model, Harness, and Tool-Contract Effects on Duplicate Side Effects in LLM Agents

The paper investigates where exactly‑once semantics should be enforced for tool‑using agents—within the model, the agent harness, or the tool contract—by evaluating 25,930 episodes across nine models, three harnesses, two contract variants, and fifteen recovery conditions. Using the LIMBO sandbox, the study shows that when an immediate read‑back is available, frontier models rarely duplicate lost acknowledgements, whereas weaker models do; when read‑back is unavailable, the contract’s idempotency keys explain most duplicate behavior. The authors prove that verification‑only policies cannot guarantee exactly‑once under late commits without bounded in‑flight time, and that waiting only helps when delays are short and predictable. whyItMatters":"The findings clarify that enforcing exactly‑once semantics largely depends on the tool contract and fault type, guiding designers on where to focus reliability mechanisms for LLM agents."

By Jiapeng Li
arXiv Computation and Language
Aug 27

When Stale Constraints Go Unchecked: Budgeted Verification Failures in Inherited Agent Memory

The paper investigates how agents that inherit consolidated memory can mistakenly rely on constraints that were once true but have since been superseded by newer records. By modeling supersession explicitly and limiting verification to two records, the study shows that native allocation often leads to stale-consistent decisions, while reallocating one verification slot to the critical provenance path markedly improves consistency. The results suggest that memory systems may need separate freshness or supersession signals beyond relevance to avoid such errors.

By Kazuki Nakayashiki
arXiv AI
4d ago

StateTape: Action-Conditioned Evidence Lifecycle Modeling for Long-Horizon Coding Agents

StateTape introduces a new framework for long‑horizon coding agents that rewrites the agent’s context as the code repository changes, rather than letting the context grow with every observation. It models the repository as a symbol‑level code graph, using a tape to mark symbols altered by each write and a manager model to resolve stale records. The authors provide theoretical analysis, a new benchmark called TraceBench, and empirical results showing higher resolve rates across six agents and three edit‑heavy benchmarks with minimal computational overhead.

By Ziyang Yu, Liang Zhao, Bowen Zhu, Hasibul Haque