arXiv:2606. 07963v1 Announce Type: new Abstract: Backdoor attacks in large language models (LLMs) are often treated as isolated trigger-response failures, motivating defenses tailored to specific triggers or behaviors.
By Omar Mahmoud, Aly M. Kassem, Thommen George Karimpanal, Buddhika Laknath Semage, Negar Rostamzadeh, Golnoosh Farnadi, Santu Rana
arXiv:2605. 01133v3 Announce Type: replace-cross Abstract: Large language model (LLM)-powered multi-agent systems (MAS) enable agents to communicate and share information, achieving strong performance on complex tasks.
By Lingxi Zhang, Guangtao Zheng, Hanjie Chen
The paper investigates whether hidden representations in latent-based multi‑agent systems can carry attack information that remains effective during normal operation. A latent attack framework is introduced, reactivating attack effects through latent interventions without using adversarial text. Experiments show that these latent attacks can significantly degrade task performance, especially when targeting inter‑agent KV‑cache handoffs, and that the degradation cannot be explained by simple perturbations or invalid generation.
By Chenxi Wang, Ruiyang Huang, Jiayan Sun, Lei Wei, Yifan Wu
arXiv:2603. 21194v2 Announce Type: replace-cross Abstract: Multi-agent discussions have been widely adopted, motivating growing efforts to develop attacks that expose their vulnerabilities.
By Qiuchi Xiang, Haoxuan Qu, Hossein Rahmani, Jun Liu
arXiv:2607. 06807v1 Announce Type: cross Abstract: While enabling effective collaboration on complex tasks, LLM-based Multi-Agent Systems (MAS) face critical security challenges due to vulnerabilities at the agent and interaction levels.
By Haowen Xu, Xue Tan, Lei Ma, Zhihao Zhang, Chao Wang, Qingze Wang, Ping Chen, Jun Dai, Xiaoyan Sun
arXiv:2608. 02657v1 Announce Type: cross Abstract: Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.
By Jianshuo Dong, Yiming Liu, Maosen Zhang, Nan Deng, Xu Peng, Xiaoping Zhang, Tianwei Zhang, Jie Zhang, Han Qiu