arXiv AI By Aaron Rose, Carissa Cullen, Sahar Abdelnabi, Philip Torr, Brandon Gary Kaplowitz, Christian Schroeder de Witt

Detecting Multi-Agent Collusion Through Multi-Agent Interpretability

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Aug 20

Beyond the Transcript: Detecting Covert Co ordination in Latent Multi-Agent Communication

The paper introduces Verifiable Latent Alignments (VLA), a framework that monitors and steers hidden communication channels between language‑model agents. VLA links private latent states to public actions via event identifiers, enabling causal analysis. Experiments on a multi‑agent auction benchmark show high detection accuracy and effective mitigation of collusion, even without training on attack examples.

By Ramneet Kaur, Pradyumna Chari, Ramesh Raskar, Jugad Singh, Sumit Kumar Jha, Anirban Roy
arXiv AI
Jul 29

Early Detection of Distributed Backdoors in Multi-Agent LLM Systems: A Characterization Study

arXiv:2607. 24893v1 Announce Type: cross Abstract: Multi-agent LLM systems can be attacked by a payload that no single agent ever holds in full: a poisoned tool hides encrypted fragments in its observations, spreads them across several agents, and an external step reassembles and executes them after the run.

By Diego Fernandez Arias, Dev Prashant Mistry, Ren Wang, Yibo Hu