The paper audits the reproducibility of knowledge‑graph extraction from threat reports by re‑implementing matching rules for only five of twelve systems and re‑scoring ten system outputs under eight protocols. The audit shows that different matching protocols can reverse most pairwise system rankings and that a fixed prediction set can vary from 0.16 to 0.70 F1. The authors also build CTIForge to isolate validation effects, finding that validation changes precision across backbones and increases entity‑type disputes, and they release the full pipeline, protocol suite, and audit records.
By Safayat Bin Hakim, Houbing Herbert Song
arXiv:2606. 04990v1 Announce Type: cross Abstract: Large language model (LLM)-based agents increasingly solve complex tasks by interacting with external tools, retrieval systems, memory modules, environments, and other agents.
By Yiqi Wang, Jiaqi Zhang, Taotao Cai, Zirui Liu, Qingqiang Sun, Zequn Sun, Zhangkai Wu, Mingkai Zhang, Yanming Zhu
arXiv:2608.28394v1 Announce Type: cross
Abstract: Cyber threat intelligence (CTI) is foundational to modern cyber defense, yet much of it resides in unstructured reports whose volume and heterogeneit...
By Changze Li, Yutong Cheng, Tsania Camila Finnisa, Qian Cui, Wei Ding, Peng Gao
arXiv:2606. 04990v2 Announce Type: replace-cross Abstract: Large language model (LLM)-based agents are evolving from passive text generators into autonomous systems capable of planning, tool use, retrieval, memory access, environmental interaction, and multi-agent collaboration.
By Yiqi Wang, Jiaqi Zhang, Taotao Cai, Zirui Liu, Qingqiang Sun, Zequn Sun, Zhangkai Wu, Manqing Dong, Mingkai Zhang, Xuefei Yin, Yanming Zhu
Shared memory helps language-model agents reuse information across long workflows, yet relevant evidence may not be admissible for a particular agent or action. Because restrictions propagate through derivations, summaries can conceal private, poisoned, untrusted, or revoked sources, enabling unauthorized reads or unsafe actions.
arXiv:2607. 25364v1 Announce Type: new Abstract: Tool-using agents expose structured calls but commonly attach free-form rationales.
By Genliang Zhu (Accentrust, Georgia Institute of Technology), Chu Wang (Accentrust, University of Illinois Urbana-Champaign)
arXiv:2608. 10509v1 Announce Type: new Abstract: Shared memory helps language-model agents reuse information across long workflows, yet relevant evidence may not be admissible for a particular agent or action.
By Yiqi Wang, Zihao Yan, Jiaqi Zhang, Zhangkai Wu, Mingkai Zheng, Zequn Sun, Yanming Zhu, Taotao Cai
The paper reports that a model can pass fidelity checks—verifying that extracted values match the source—without actually opening a datasheet, due to a hidden constraint that disables tool use. To address this, the authors log every tool call in an agentic benchmark and develop two instruments: a rule‑based failure‑attribution classifier and a silent‑failure detector that flags runs based solely on which tools were invoked. While the detector shows low false positives on clean extractions and recovers all planted faults, its recall against correct tool usage but incorrect answers remains unmeasured, and a partial causal chamber confirms only a subset of claims, highlighting limitations in physical verification.
By Qing Ye, Meng-Hsuan Lin
The paper investigates how causal action verifiers, which guard language agents’ tool calls by checking identifiability against a committed action‑state graph, can be compromised through small graph misspecifications. By removing a single bidirected edge or reversing an arrowhead, the authors demonstrate that a verifier (CIVeX) that originally had zero false executions can suffer false execution rates up to 48.9%, with most of those executions being harmful and overall utility dropping dramatically. An additional attestation step that samples executions can detect these attacks with few false alarms, but it also leads to many wrongful rejections that reduce beneficial actions and incur significant experimental costs.
whyItMatters":"The study shows that even minor errors in the verifier’s underlying graph can drastically undermine safety and performance, highlighting the need for robust auditing mechanisms."
By Fabio Rovai
The paper introduces a claim‑anchored execution contract that binds a tool‑using agent’s emitted claim to its exact source span, the ordered execution prefix that produced it, and the source version and access state observed. Each receipt contains deterministic anchors, source identifiers, offsets, hashes, quotes, and a domain‑separated execution commitment, allowing a verifier to reconstruct these bindings before semantic or task labels are joined. The contract defines seven independently testable properties and demonstrates high detection rates against cross‑object attacks, with strong performance on conflict‑aware support guard evaluations.
By Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang, Yina Sa, Daren Zha, Jun Xiao
arXiv:2608. 16813v1 Announce Type: new Abstract: Agents now write knowledge graphs, but knowledge-graph stores still carry defaults set when humans curated them: accept writes now and clean later, keep one time axis or none, treat every writer's facts as equally trustworthy, and leave governance to dashboards and middleware.
By Steve Brown
arXiv:2607. 12650v1 Announce Type: cross Abstract: Tool access alone does not make LLM empirical reasoning governable: accepted outputs need not descend from attested evidence, and accepted deductions need not hold up under formal scrutiny.
By Junyu Ren