arXiv Machine Learning By Sajjad Khan

Resume Means Resume: A Machine-Checked Conformance Contract for Checkpoint, Interrupt, and Resume Semantics in Workflow Persistence Layers

Read the original on arXiv Machine Learning →

arXiv:2608. 03836v1 Announce Type: new Abstract: A framework that persists execution state so a run can be interrupted, survive a crash, and continue must decide what a resume means for effects that already fired.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 25

Stale Does Not Mean Unsafe: Guard Precision for Tool-Using LLM Agents under Infrastructure State Races

The paper investigates how tool‑using language‑model agents can safely commit changes to infrastructure when external state may change between read and commit. By distinguishing invalidating races from predicate‑preserving and irrelevant ones, the authors evaluate three commit‑time guard granularities—global epoch, read‑set version, and semantic commit predicate—using a deterministic simulator and three quantized model families. The study finds that only the complete predicate guard consistently eliminates unsafe commits, while freshness‑based guards block a large proportion of benign races and model‑side signals fail to replace precise semantic enforcement.

By Zihao Zheng, Jiayu Long, Baichuan Li, Junyi Yao
arXiv AI
Sep 25

Where Does Exactly-Once Live? Model, Harness, and Tool-Contract Effects on Duplicate Side Effects in LLM Agents

The paper investigates where exactly‑once semantics should be enforced for tool‑using agents—within the model, the agent harness, or the tool contract—by evaluating 25,930 episodes across nine models, three harnesses, two contract variants, and fifteen recovery conditions. Using the LIMBO sandbox, the study shows that when an immediate read‑back is available, frontier models rarely duplicate lost acknowledgements, whereas weaker models do; when read‑back is unavailable, the contract’s idempotency keys explain most duplicate behavior. The authors prove that verification‑only policies cannot guarantee exactly‑once under late commits without bounded in‑flight time, and that waiting only helps when delays are short and predictable. whyItMatters":"The findings clarify that enforcing exactly‑once semantics largely depends on the tool contract and fault type, guiding designers on where to focus reliability mechanisms for LLM agents."

By Jiapeng Li
arXiv AI
Sep 2

REVISE: Validity-Guided Recovery for Online Revisions in Agent Workflows

The paper introduces “Revise”, a runtime system that performs validity-guided, fine-grained recovery for online revisions in structured agent workflows. When a revision arrives, Revise intersects the change with recorded data and control dependencies, propagates the impact through the partially executed DAG, stops invalid work, preserves unaffected progress, and recomputes only the affected region. Experiments on real coding‑agent traces and LangGraph/LLMCompiler applications show that Revise matches a latest‑version oracle, reduces model calls by up to 56%, and improves service‑level objective goodput under load.

By Ruoling Qi, Xuaner Wu, Penghang Liu, Jian Chen, Yirui Liu