arXiv AI By Hwai-Jung Hsu, Cheng-Jan Chi, Hanna Everett

EA-Graph: Artifact-Anchored Verification Memory for Coding Agents under Upstream Drift

Read the original on arXiv AI →

arXiv:2608. 04278v1 Announce Type: cross Abstract: Coding agents increasingly work across sessions, but prose notes can preserve a conclusion without the program state that supported it.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 27

When Stale Constraints Go Unchecked: Budgeted Verification Failures in Inherited Agent Memory

The paper investigates how agents that inherit consolidated memory can mistakenly rely on constraints that were once true but have since been superseded by newer records. By modeling supersession explicitly and limiting verification to two records, the study shows that native allocation often leads to stale-consistent decisions, while reallocating one verification slot to the critical provenance path markedly improves consistency. The results suggest that memory systems may need separate freshness or supersession signals beyond relevance to avoid such errors.

By Kazuki Nakayashiki
arXiv AI
Sep 3

The Memory Trust Gap: Capability-Dependent Failures in Persistent-Memory Agents

The paper investigates how persistent memory in AI agents can lead to over‑trust in stale facts, creating a "Memory Trust Gap" that worsens as model capability increases. Using a benchmark with Benefit and Safety suites across Qwen3 models of varying sizes, the authors show that larger models are more prone to harmful over‑trust, especially when metadata is absent or misleading. They also demonstrate that mitigation strategies such as exposing metadata or pre‑resolving conflicts improve accuracy, but the effectiveness depends on model size and dataset.

By Jundong Hu, Shekar Ramachandran
arXiv Machine Learning
Sep 23

Impact Is Not Invalidation: Ask About the Claim, Not the Diff

The paper investigates how machine‑learning models can determine whether a claim (a test assertion) remains valid after a code change. It compares two questioning strategies: asking whether a diff preserves behavior versus asking whether a specific claim still holds. The authors find that the latter approach yields far higher precision (up to 0.974) across models of varying cost, while the former performs poorly (precision 0.291–0.329). They also benchmark against a regression‑test selector, showing that even near‑complete knowledge of a change’s reach does not reliably identify falsified claims. The study is grounded in 10,369 mined claims with 184 execution‑verified flips from 23 Python libraries.

By Atul Anand