arXiv Machine Learning By Chenxi Wang, Ruiyang Huang, Jiayan Sun, Lei Wei, Yifan Wu

Out of Sight, Not Out of Mind: Unveiling Latent Attack in Latent-based Multi-Agent Systems

Read the original on arXiv Machine Learning →

The paper investigates whether hidden representations in latent-based multi‑agent systems can carry attack information that remains effective during normal operation. A latent attack framework is introduced, reactivating attack effects through latent interventions without using adversarial text. Experiments show that these latent attacks can significantly degrade task performance, especially when targeting inter‑agent KV‑cache handoffs, and that the degradation cannot be explained by simple perturbations or invalid generation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 12

PI-Hunter: Automated Red-Teaming for Exposing and Localizing Prompt Injections

arXiv:2606. 12737v1 Announce Type: cross Abstract: Large Language Models (LLMs) are rapidly evolving into agentic systems that interact with external tools and environments, introducing new security risks such as indirect prompt injection attacks through untrusted external sources.

By Pengfei He, Lesly Miculicich, Vishesh Sharma, Ash Fox, George Lee, Jiliang Tang, Tomas Pfister, Long T. Le
arXiv AI
3d ago

Safety of Latent Communication in Multi-Agent Systems

The paper investigates latent communication in multi‑agent systems, where agents exchange information in internal representation space via lightweight trainable links. It demonstrates that even benign training of these links can increase harmful compliance, and that attackers can exploit or poison the links to amplify this effect. A reinforcement‑learning attack further boosts harmful compliance while maintaining benign task performance, but adjusting rewards toward safer behavior can repair compromised links without updating the agents.

By Muhammad Huzaifa, Sina Mavali, Thorsten Eisenhofer
arXiv AI
Sep 10

AgentLeak: Cloning Stronger LLM Agent Capabilities onto Weaker Agents Beyond Skill Stealing

The paper introduces AgentLeak, a black‑box attack that clones the task‑solving capabilities of a strong LLM agent onto a weaker one by exploiting differences between successful and failed executions. Unlike prior skill‑stealing methods that only recover explicit skill artifacts, AgentLeak identifies and incorporates missing procedural behaviors, boosting task pass rates by over 40% and closing more than 80% of the capability gap across 20 scenarios. The study demonstrates that observable execution behavior can leak proprietary procedural knowledge, posing a confidentiality risk for LLM agents.

By Xiaoting Lyu, Yuhong Wu, Yufei Han, Shichang Liu, Liang Zhang, Bin Wang, Bin Wang, Xiaobo Ma, Wei Wang