The paper investigates how tool‑using language‑model agents can safely commit changes to infrastructure when external state may change between read and commit. By distinguishing invalidating races from predicate‑preserving and irrelevant ones, the authors evaluate three commit‑time guard granularities—global epoch, read‑set version, and semantic commit predicate—using a deterministic simulator and three quantized model families. The study finds that only the complete predicate guard consistently eliminates unsafe commits, while freshness‑based guards block a large proportion of benign races and model‑side signals fail to replace precise semantic enforcement.
By Zihao Zheng, Jiayu Long, Baichuan Li, Junyi Yao
arXiv:2605. 27784v2 Announce Type: replace Abstract: LLM agents are governed by long-lived prompt policies, where individually reasonable stand- ing rules can jointly govern the same pre- generation state.
By Lu Yan, Xuan Chen, Xiangyu Zhang
arXiv:2609.24165v1 Announce Type: new
Abstract: Synchrotron data reduction, detector calibration followed by azimuthal integration of terabyte-scale diffraction series, is a multi-step, expert-bound...
By Pawan K. Tripathi, Hemant Sharma, Andrew Chuang, Mathew J. Cherukara
Autonomous agents that author building information models need more than access to a host API. They need a representation of what they intended, what a compiler decided on their behalf, what was refus...
arXiv:2609.14578v1 Announce Type: new
Abstract: Autonomous agents that author building information models need more than access to a host API. They need a representation of what they intended, what a...
By Dmitry Kuklev
arXiv:2609.37603v1 Announce Type: cross
Abstract: A common safeguard for a data pipeline is redundant computation: derive each published number by two routes built on different technology and refuse...
By Fabio Rovai
arXiv:2606. 07808v1 Announce Type: new Abstract: Reasoning language models deployed in agentic workflows must follow an instruction hierarchy: when instructions from different sources conflict, the model should obey the highest-privilege applicable instruction.
By Sanjay Kariyappa, G. Edward Suh
VeriPhy is an auditable physical‑verification system that transforms a text prompt into typed physical obligations and a statically validated execution plan before any video frames are generated. During execution, it gates calls to frozen low‑level experts (segmentation, tracking, counting, depth, OCR, audio‑event detection, etc.) and records provenance‑carrying evidence for each action. The system maps these records to a three‑valued state—supported, contradicted, or unknown—providing traceable verdicts that can be used to refine generation models.
By Wenzhuo Xu, Yuchen Zhu, Chongjian Ge, Xuan Shen, Jing Shi, Jason Kuen, Yongxin Chen, Molei Tao, Christopher McComb, Noelia Grande Guti\'errez, Jiuxiang Gu
arXiv:2606. 22504v1 Announce Type: cross Abstract: Coding agents often receive broad tool access for an entire task, even when a resource is needed only for one subgoal.
By Igor Santos-Grueiro
The paper introduces the first benchmark suite for formal verification of PLC programs, covering both Structured Text and Ladder Diagram encodings of IEC 61131‑3. It contains 50 programs in 83 variants across ten industrial domains, each paired with a formal property, a machine‑checkable expected verdict, and a violation witness in SV‑COMP format. The suite’s ground‑truth methodology uses construction, fault injection, and cross‑tool consensus to ensure reliable verdicts, and the authors validate it with ESBMC and nuXmv, revealing format and semantics fragmentation in existing tools.
By Pierre Dantas, Lucas Cordeiro, Waldir Junior
arXiv:2606. 06240v1 Announce Type: cross Abstract: Persistent memory for an LLM agent is a write-heavy substrate: every belief update is a versioned write, and a new claim may contradict a stored one.
By Ziming Wang
arXiv:2607. 28545v2 Announce Type: replace-cross Abstract: Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics, logs, traces, and source code, starting from ambiguous user-facing reports, often hours after the incident began.
By Albert Gong, Kyuseong Choi, Abhineet Agarwal, Jason Schechner, Ryan Huang, Raj Agrawal, Anish Agarwal, Raaz Dwivedi