arXiv:2606. 07316v2 Announce Type: replace-cross Abstract: Can a committee of LLM agents reach agreement that is certifiable at the level of meaning, not only at the level of a label?
By Haoran Xu, Lei Zhang, Iadh Ounis, Xianbin Wang
The paper introduces a framework for evaluating how large language model agents revise their success criteria after failures, defining five non‑compensatory conditions that must be met for a criterion revision to be considered valid. Using the CMB‑0.1 protocol, the authors test twelve cross‑domain scenarios across four system configurations, finding that no model trial satisfies all five conditions and highlighting specific failure modes such as zero‑state reconstruction and inadequate intervention sensitivity. They propose a more stringent trace‑anchored CMB‑0.4 protocol to better isolate and measure criterion revision in future studies.
By Guodong Xu
The paper introduces BSC‑R, a deterministic effect‑boundary mechanism that ties a single‑use commit authorization to the specific action and the semantic state that justified it, aiming to close the proposal‑to‑commit gap in tool‑using language‑model agents. Experiments on 2,847 AgentDojo episodes and 10,302 frozen proposals show that BSC‑R preserves the agent’s original behavior while rejecting unauthorized changes, and further tests on a boundary‑drift experiment and the CONTINUITY suite demonstrate high success rates in valid contexts and robust handling of replay and ambiguous cases. However, broader testing reveals that BSC‑R still allows a 25% invalid‑effect commit rate in a larger attack set, indicating that it provides scoped, not universal, safety.
By Wesley Shu
The paper investigates how tool‑using language‑model agents can safely commit changes to infrastructure when external state may change between read and commit. By distinguishing invalidating races from predicate‑preserving and irrelevant ones, the authors evaluate three commit‑time guard granularities—global epoch, read‑set version, and semantic commit predicate—using a deterministic simulator and three quantized model families. The study finds that only the complete predicate guard consistently eliminates unsafe commits, while freshness‑based guards block a large proportion of benign races and model‑side signals fail to replace precise semantic enforcement.
By Zihao Zheng, Jiayu Long, Baichuan Li, Junyi Yao
arXiv:2607. 16109v1 Announce Type: new Abstract: State machine replication (SMR) and Byzantine fault-tolerant (BFT) consensus guarantee agreement despite a bounded number of arbitrary, colluding faulty participants.
By Jun He, Deying Yu
arXiv:2606. 07834v1 Announce Type: cross Abstract: LLM judges increasingly turn verdicts into system commitments.
By Haoran Xu
arXiv:2606. 19356v1 Announce Type: cross Abstract: When multi-agent LLM systems produce bad answers, not all failures are equal: some answers are grounded in the right material but incomplete, while others are simply ungrounded and should be stopped.
By Anantha Sharma
The paper investigates how machine‑learning models can determine whether a claim (a test assertion) remains valid after a code change. It compares two questioning strategies: asking whether a diff preserves behavior versus asking whether a specific claim still holds. The authors find that the latter approach yields far higher precision (up to 0.974) across models of varying cost, while the former performs poorly (precision 0.291–0.329). They also benchmark against a regression‑test selector, showing that even near‑complete knowledge of a change’s reach does not reliably identify falsified claims. The study is grounded in 10,369 mined claims with 184 execution‑verified flips from 23 Python libraries.
By Atul Anand
arXiv:2607. 12650v1 Announce Type: cross Abstract: Tool access alone does not make LLM empirical reasoning governable: accepted outputs need not descend from attested evidence, and accepted deductions need not hold up under formal scrutiny.
By Junyu Ren
arXiv:2606. 12703v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) agents increasingly run with persistent memory that accumulates across user sessions.
By Tarun Sharma
arXiv:2607. 05397v1 Announce Type: cross Abstract: Agent systems increasingly execute rather than advise.
By James Rhodes, George Kang
arXiv:2606. 31023v1 Announce Type: cross Abstract: Hard-constrained sequential decision systems have no certified way to spend the test-time compute of modern AI: executing the multi-step drafts of a learned policy or a frozen LLM forfeits the feasibility guarantee a trusted solver provides, while invoking the solver at every step forfeits the speed the AI offers.
By Chenyu Zhou, Qiliang Jiang, Shuning Wu, Xu Zhou