arXiv AI By Qiuchi Xiang, Haoxuan Qu, Hossein Rahmani, Jun Liu

Is Monitoring Enough? Strategic Agent Selection For Stealthy Attack in Multi-Agent Discussions

Read the original on arXiv AI →

arXiv:2603. 21194v2 Announce Type: replace-cross Abstract: Multi-agent discussions have been widely adopted, motivating growing efforts to develop attacks that expose their vulnerabilities.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 12

PI-Hunter: Automated Red-Teaming for Exposing and Localizing Prompt Injections

arXiv:2606. 12737v1 Announce Type: cross Abstract: Large Language Models (LLMs) are rapidly evolving into agentic systems that interact with external tools and environments, introducing new security risks such as indirect prompt injection attacks through untrusted external sources.

By Pengfei He, Lesly Miculicich, Vishesh Sharma, Ash Fox, George Lee, Jiliang Tang, Tomas Pfister, Long T. Le
arXiv AI
Sep 18

MAS-Shield: A Defense Framework for Secure and Efficient LLM MAS

MAS-Shield is a defense framework for Large Language Model–based Multi-Agent Systems that uses a coarse‑to‑fine filtering pipeline. It first selects critical agents, then applies lightweight auditing to most cases, and finally escalates only suspicious signals to a heavyweight committee. Experiments show a 92.5% recovery rate against adversarial attacks and a latency reduction of over 70% compared to existing methods.

By Kaixiang Wang, Zhaojiacheng Zhou, Bunyod Suvonov, Jiong Lou, Zihan Wang, Yuxiang Zheng, Yidan Lin, Wutong Zhang, Xianghan Kong, Chentao Wu, Jie Li
arXiv AI
3d ago

MiniRep: Robust Reputation-Based Aggregation for Multi-Agent Debate

MiniRep is a reputation‑based aggregation system designed for multi‑agent debate (MAD) that remains robust even when malicious agents are present. It evaluates agents on both their current task performance and historical reputation, while preventing groups of agents with highly similar responses from dominating the final decision. Experiments on the MATH benchmark show that MiniRep consistently outperforms conventional MAD aggregation and other reputation‑based approaches across a wide range of attack scenarios.

By Jiaming Zhang, Yuwan Liu, Yue Huang, Sisi Duan