Detecting Multi-Agent Collusion Through Multi-Agent Interpretability
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2608. 02698v1 Announce Type: cross Abstract: Tool-using agents built on large language models (LLMs) are increasingly deployed not by a single operator but by many, side by side on shared infrastructure.
The paper introduces Verifiable Latent Alignments (VLA), a framework that monitors and steers hidden communication channels between language‑model agents. VLA links private latent states to public actions via event identifiers, enabling causal analysis. Experiments on a multi‑agent auction benchmark show high detection accuracy and effective mitigation of collusion, even without training on attack examples.
arXiv:2602. 16346v4 Announce Type: replace-cross Abstract: LLM-based agents execute real-world workflows via tools and memory.
arXiv:2608. 02657v1 Announce Type: cross Abstract: Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.
arXiv:2607. 24893v1 Announce Type: cross Abstract: Multi-agent LLM systems can be attacked by a payload that no single agent ever holds in full: a poisoned tool hides encrypted fragments in its observations, spreads them across several agents, and an external step reassembles and executes them after the run.
arXiv:2605. 01143v2 Announce Type: replace Abstract: Large Language Model (LLM)-powered agents demonstrate strong capabilities in autonomous task execution, tool use, and multi-step reasoning.