arXiv:2606. 07054v1 Announce Type: cross Abstract: Autonomous LLM agents can pursue hidden malicious objectives through sequences of individually benign actions, making sabotage difficult to detect using standard trajectory-level monitoring.
By Vijitha Mittapalli, Shreyaa Jayant Dani, Satya Srujana Pilli, Snigdha Ansu, Mohammadreza Teymoorianfard, Franck Dernoncourt, Hongjie Chen, Yu Wang, Ryan A. Rossi, Nesreen K. Ahmed
arXiv:2602. 16346v4 Announce Type: replace-cross Abstract: LLM-based agents execute real-world workflows via tools and memory.
By Nivya Talokar, Ayush K Tarun, Murari Mandal, Maksym Andriushchenko, Antoine Bosselut
arXiv:2605. 01143v2 Announce Type: replace Abstract: Large Language Model (LLM)-powered agents demonstrate strong capabilities in autonomous task execution, tool use, and multi-step reasoning.
By Sheldon Yu, Yingcheng Sun, Hanqing Guo, Qianqian Tong
The paper introduces Attnlocate, a runtime framework that localizes behavior‑guiding instructions within the attention matrix of large language model agents. By treating this localization as an object detection task, Attnlocate uses a multi‑head, multi‑layer attention aggregation scheme and a 1‑D U‑Net to identify spans that influence tool‑calling decisions. The system then adjudicates potential malicious invocations based on the authority of the source, achieving high detection metrics across diverse LLM families and demonstrating transferability to unseen models.
By Yichao Gao, Yumo Zhang, Yunhao Yao, Haohua Du, Puhan Luo, Ruiqi Li, Zhiqiang Wang
arXiv:2603. 19423v2 Announce Type: replace-cross Abstract: Large language model (LLM) agents increasingly rely on external tools (file operations, API calls, database transactions) to autonomously complete complex multi-step tasks.
By Shawn Li, Yue Zhao
arXiv:2606. 11063v1 Announce Type: new Abstract: AI control protocols oversee untrusted models by monitoring their actions and modifying potentially unsafe steps, often using a trusted model.
By Joachim Schaeffer, Thomas Jiralerspong, Alexander Panfilov, Guillaume Lajoie, Jonas Geiping, Yoshua Bengio, Roland S. Zimmermann
arXiv:2607. 26998v1 Announce Type: cross Abstract: Large language model (LLM) agents automate penetration testing through an observation-action loop, selecting actions based on observations returned by tools.
By Ruoyu Wang, Heng Zhao, Renjie Wu, Mengnan Zhao, Zhixuan Chu, Wanyu Lin, Tianhang Zheng
arXiv:2605. 00994v2 Announce Type: replace-cross Abstract: Finetuning can significantly modify the behavior of large language models, including introducing harmful or unsafe behaviors.
By Mohammed Abu Baker, Luca Baroni, Dan Wilhelm
The paper introduces INTENT-AS-A-TOOL, a method that equips large language models with intent-targeted tools to provide a fine-grained, judge‑free signal of their commitment to specific behaviors during reasoning. By monitoring the probability of calling these intent tools, the authors can track how intent evolves throughout generation, complementing chain‑of‑thought monitoring and expanding post‑hoc labels into dense trajectories. The approach identifies critical steps for online intervention, demonstrating that action preferences are useful for detecting agentic misalignment in autonomous agents.
By Yutong Zhang, Jianshuo Dong, Peng Xu, Long Wang, Jie Zhang, Tianwei Zhang, Xiaoping Zhang, Han Qiu
The paper titled "Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents" demonstrates how an attacker can bypass current skill‑scanning defenses by crafting malicious skills that evade both static analysis and LLM‑based semantic checks. By moving malicious payloads into natural language and distributing instructions across files, the white‑box attacker named Pretext achieves high evasion rates—up to 97% against a frozen detector and 77% against a co‑adaptive one—across three open‑source models.
By Tobias Kaisar, Aritra Dhar
Skills extend an agent's capabilities by injecting instructions and information into the context, and are widely used by agents such as OpenClaw and Claude Code. Prior work shows third-party marketpla...
KC-Bench is a dynamic interactive benchmark designed to evaluate how large language model agents reconcile user instructions, internal knowledge, and real‑time environmental observations. It consists of 238 manually curated multi‑turn tasks that test world‑knowledge conflicts, input inconsistencies, and multi‑source temporal conflicts, using a user simulator, stateful tools, deterministic environment assertions, an open‑source natural‑language evaluator, and human trajectory verification. Evaluation of nine models—including DeepSeek‑V4‑Flash, GLM‑5.2, and MiniMax‑M3—reveals significant cross‑domain variation, with no model reliably handling factual correction, identity consistency, and temporal conflict resolution across all settings, and shows that missed conflicts can propagate to tool calls or synthetic protected‑data flows.