arXiv AI By Jiaming Zhang, Yuwan Liu, Yue Huang, Sisi Duan

MiniRep: Robust Reputation-Based Aggregation for Multi-Agent Debate

Read the original on arXiv AI →

MiniRep is a reputation‑based aggregation system designed for multi‑agent debate (MAD) that remains robust even when malicious agents are present. It evaluates agents on both their current task performance and historical reputation, while preventing groups of agents with highly similar responses from dominating the final decision. Experiments on the MATH benchmark show that MiniRep consistently outperforms conventional MAD aggregation and other reputation‑based approaches across a wide range of attack scenarios.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
1d ago

MADBench: Benchmarking the Security of Multi-Agent Debate

MADBench is a benchmark that evaluates the security of Multi-Agent Debate (MAD) systems, which allow large language models to exchange and critique answers. The study categorizes attacks into a layered taxonomy aligned with the MAD workflow and tests six attack families across 356 source tasks and 3,958 test cases. Results indicate that while MAD can reduce answer accuracy attacks compared to single-agent baselines, it may amplify unauthorized reads or writes, and even with collusion among agents, the final answer changes from correct to wrong only 28.30% of the time.

By Yuwan Liu, Jiaming Zhang, Yue Huang, Sisi Duan
arXiv AI
Jun 12

Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents

arXiv:2606. 13385v1 Announce Type: cross Abstract: Web agents driven by large language models (LLMs) are increasingly deployed in real-world environments, where they operate over untrusted web content and execute actions with direct consequences.

By Zihao Wang, Yiming Li, Yutong Wu, Zheyu Liu, Kangjie Chen, Fok Kar Wai, Pin-Yu Chen, Vrizlynn L. L. Thing, Bo Li, Dacheng Tao, Tianwei Zhang