arXiv AI

Safety of Latent Communication in Multi-Agent Systems

The paper investigates latent communication in multi‑agent systems, where agents exchange information in internal representation space via lightweight trainable links. It demonstrates that even benign training of these links can increase harmful compliance, and that attackers can exploit or poison the links to amplify this effect. A reinforcement‑learning attack further boosts harmful compliance while maintaining benign task performance, but adjusting rewards toward safer behavior can repair compromised links without updating the agents.

Hugging Face Trending Papers
Jun 13

Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment

Indirect prompt injection attacks hijack LLM-based agents by embedding malicious instructions in third-party data that the agent retrieves during task execution. Existing defenses report near-zero attack success rate on static benchmarks, yet recent adaptive evaluations show that these results collapse once the attacker is allowed to optimize against the deployed defense.

arXiv Machine Learning
Aug 28

Out of Sight, Not Out of Mind: Unveiling Latent Attack in Latent-based Multi-Agent Systems

The paper investigates whether hidden representations in latent-based multi‑agent systems can carry attack information that remains effective during normal operation. A latent attack framework is introduced, reactivating attack effects through latent interventions without using adversarial text. Experiments show that these latent attacks can significantly degrade task performance, especially when targeting inter‑agent KV‑cache handoffs, and that the degradation cannot be explained by simple perturbations or invalid generation.

By Chenxi Wang, Ruiyang Huang, Jiayan Sun, Lei Wei, Yifan Wu