FinalityBench is an executable benchmark that tests how agents decide on shipping, re‑capturing, refunding, or waiting when a merchant’s payment processor, ledger, ERP, and bank feed receive delayed, duplicated, dropped, or reordered messages, causing contradictory beliefs about an order. The benchmark uses a hidden canonical event log and faulted delivery streams to generate system views, scoring each episode by the merchant’s terminal economic position relative to a privileged reference. It contains 321 tasks, including 45 twin pairs where all four views are identical yet the correct disposition differs, and evaluates nine programmatic policies, revealing that a ship‑on‑first‑sign policy performs best by accuracy but worst by paired loss, while a runtime‑gated irreversible‑action policy achieves 85.4% accuracy without losing money.
By Abhishek Sharma
arXiv:2606. 07316v1 Announce Type: cross Abstract: Byzantine collaboration among large-language-model agents requires a finality-control primitive: given delivered stochastic, structured natural-language proposals, the protocol must decide whether the round supports a commit, what kind of commit, or a typed safe abort.
By Haoran Xu, Lei Zhang, Iadh Ounis, Xianbin Wang
The paper introduces a claim‑anchored execution contract that binds a tool‑using agent’s emitted claim to its exact source span, the ordered execution prefix that produced it, and the source version and access state observed. Each receipt contains deterministic anchors, source identifiers, offsets, hashes, quotes, and a domain‑separated execution commitment, allowing a verifier to reconstruct these bindings before semantic or task labels are joined. The contract defines seven independently testable properties and demonstrates high detection rates against cross‑object attacks, with strong performance on conflict‑aware support guard evaluations.
By Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang, Yina Sa, Daren Zha, Jun Xiao
arXiv:2608. 05373v1 Announce Type: cross Abstract: Intraday market manipulation is hard to detect because its footprint is brief, buried in millions of quotes, and statistically similar to ordinary volatility.
By Alex Chen, Maria Hybinette
The paper evaluates CASCADE, a fully local layered defense for Model Context Protocol (MCP)-based systems, by conducting a component ablation and corpus audit on a fixed 5,000-sample dataset. It demonstrates that the choice of aggregation convention heavily influences reported metrics, that detection performance varies with provenance, and that the released configuration does not fully disclose the operating point. The study also shows that a local review model invoked for a third of requests does not alter classification outcomes, highlighting the importance of reproducibility and transparency in defense evaluations.
By \.Ipek Abas{\i}kele\c{s} Turgut, Edip G\"um\"u\c{s}
arXiv:2609.37819v1 Announce Type: cross
Abstract: Electronic invoices are replacing paper invoices worldwide, but today's centralized architectures leave three problems unsolved on the consumption si...
By Jia Cai