arXiv Machine Learning By Jia-Hao Xiao, Lei Feng, Min-Ling Zhang

When Collaboration Becomes a Trigger: Collective Evidence-Threshold Backdoors in Multi-Agent Systems

Read the original on arXiv Machine Learning →

arXiv:2608. 01085v1 Announce Type: cross Abstract: LLM-based multi-agent systems (MAS) extend LLM capabilities through iterative communication and shared contexts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 9

Shared Latent Structures Enable Unified Backdoor Detection and Mitigation in LLMs

arXiv:2606. 07963v1 Announce Type: new Abstract: Backdoor attacks in large language models (LLMs) are often treated as isolated trigger-response failures, motivating defenses tailored to specific triggers or behaviors.

By Omar Mahmoud, Aly M. Kassem, Thommen George Karimpanal, Buddhika Laknath Semage, Negar Rostamzadeh, Golnoosh Farnadi, Santu Rana
arXiv Machine Learning
Aug 28

Out of Sight, Not Out of Mind: Unveiling Latent Attack in Latent-based Multi-Agent Systems

The paper investigates whether hidden representations in latent-based multi‑agent systems can carry attack information that remains effective during normal operation. A latent attack framework is introduced, reactivating attack effects through latent interventions without using adversarial text. Experiments show that these latent attacks can significantly degrade task performance, especially when targeting inter‑agent KV‑cache handoffs, and that the degradation cannot be explained by simple perturbations or invalid generation.

By Chenxi Wang, Ruiyang Huang, Jiayan Sun, Lei Wei, Yifan Wu