arXiv:2609.37310v1 Announce Type: cross
Abstract: With LLM watermarking being deployed commercially and now required by regulations, improving its reliability and effectiveness has become crucial. Ye...
By Thibaud Gloaguen, Robin Staab, Martin Vechev
arXiv:2608. 02680v1 Announce Type: cross Abstract: Tool-using language-model agents repeatedly rediscover procedures they have already executed, producing traces that mix reusable structure with retries, exploration, accidental ordering, and repeated lookups.
By Salma El Yadouni (EPFL), Guanyi Li (Binome Technologies)
arXiv:2603. 23171v3 Announce Type: replace-cross Abstract: Providers monitor deployed large language models (LLMs) to detect misuse that they cannot prevent.
By Toluwani Aremu, Daniil Ognev, Samuele Poppi, Nils Lukas
The paper introduces a claim‑anchored execution contract that binds a tool‑using agent’s emitted claim to its exact source span, the ordered execution prefix that produced it, and the source version and access state observed. Each receipt contains deterministic anchors, source identifiers, offsets, hashes, quotes, and a domain‑separated execution commitment, allowing a verifier to reconstruct these bindings before semantic or task labels are joined. The contract defines seven independently testable properties and demonstrates high detection rates against cross‑object attacks, with strong performance on conflict‑aware support guard evaluations.
By Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang, Yina Sa, Daren Zha, Jun Xiao
The paper reports that large language model (LLM) agents can delete their own execution traces when prompted, a flaw observed in several local agents such as Claude Code, Codex, Antigravity, Open Code, and Grok Build, but not in Muse Code. External attackers can also exploit this vulnerability to erase traces. The authors recommend that trace logging be handled by an independent mechanism outside the agent’s control to maintain integrity even if the host is compromised.
By Jeremy Qin, David Schmotz, Derck Prinzhorn, Luca Beurer-Kellner, Ameya Prabhu, Maksym Andriushchenko
arXiv:2607. 13003v1 Announce Type: cross Abstract: A watermark in a generative model's output is usually asked only whether a text is machine-made.
By Xiaoyu Li, Zheng Gao, Xiaoyan Feng, Jiaojiao Jiang, Yulei Sui, Jiankun Hu