arXiv AI

ERRAND: Budgeted Maintenance of Agent Memory

ERRAND is a new method for budgeted maintenance of agent memory that treats revalidation of stored knowledge as a priced errand competing for scarce actions. It uses an errand index that is single‑peaked, allowing certainty in either direction to cost nothing, and repairs by writing new versions rather than deleting old ones. In experiments across two drifting tool‑use worlds, ERRAND outperforms non‑oracle policies, achieving up to 10.0 percentage points improvement over eager revalidation while using only 11.0% of steps, and it self‑terminates when no budget is imposed.

arXiv AI
Sep 4

Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory

The paper introduces PlanFence, a dependency-scoped action‑validation protocol for distributed large language model (LLM) agent teams. PlanFence requires plans to cite the exact public records they rely on, and executors validate only those records that could affect the pending action, replanning or blocking if validation is incomplete. In 30 controlled live workflows, a freshness‑only executor always acted on obsolete plans, whereas PlanFence completed all tasks without invalid actions, demonstrating controlled safety and system‑cost benefits.

By Evan Chen, Shiqiang Wang, Christopher G. Brinton
arXiv AI
Sep 24

Bounded Loops: Pre-Run Spend Bounds, Proved Termination, and Verified Completion for Agent Harnesses

The paper introduces a formal framework for agent harnesses that guarantees termination, prevents drift, and enforces spend limits through bounded loops, gates, and repair relations. It proves that these guarantees hold even with repair budgets and demonstrates the effectiveness of the system by identifying vacuous gates and achieving low false‑accept rates in a 69‑loop catalogue. The authors provide an instrumented implementation and a held‑out mutant corpus to validate gate correctness.

By Varun Pratap Bhardwaj, Garima Singh, Arun Pratap Bhardwaj
arXiv AI
Sep 7

FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflicting Financial Finality

FinalityBench is an executable benchmark that tests how agents decide on shipping, re‑capturing, refunding, or waiting when a merchant’s payment processor, ledger, ERP, and bank feed receive delayed, duplicated, dropped, or reordered messages, causing contradictory beliefs about an order. The benchmark uses a hidden canonical event log and faulted delivery streams to generate system views, scoring each episode by the merchant’s terminal economic position relative to a privileged reference. It contains 321 tasks, including 45 twin pairs where all four views are identical yet the correct disposition differs, and evaluates nine programmatic policies, revealing that a ship‑on‑first‑sign policy performs best by accuracy but worst by paired loss, while a runtime‑gated irreversible‑action policy achieves 85.4% accuracy without losing money.

By Abhishek Sharma
arXiv AI
6d ago

A General Framework for Budgeted Threshold Incentives on Request

The paper introduces a request-driven framework for designing budgeted threshold incentives on on-demand delivery platforms. It decomposes the process into four stages—conditional prediction, population reduction, trajectory integration, and budget allocation—using seven interchangeable modules that share conditional trajectory laws. The framework includes a response-correction step that reweights abundant no-offer data to match short pilot moments, and the authors prove that the end-to-end value loss is bounded by the sum of stage errors, with empirical results showing significant speedups and reduced regret compared to traditional trials.

By Zhuolin Wu, Chengrui Zhu, Wenhua Nie, Kenny Ye Liang, Junming Lin, Haiyang Li, Zhilin Li, Wenjia Geng, Zeyu Wu, Yinan Wu, Jinghua Hao, Renqing He
arXiv AI
Aug 14

Dead text or binding clause? Measuring and restoring constraint influence in black-box LLM dialogues

arXiv:2608. 12599v1 Announce Type: new Abstract: Multi-turn dialogues let users revoke constraints as easily as impose them, but revocation does not reliably take effect: models keep enacting withdrawn requirements (occasionally beneath comments asserting their removal), a failure we call \emph{behavioral relapse}, or revocation inertia.

By Haoyuan Zhu