MADBench is a benchmark that evaluates the security of Multi-Agent Debate (MAD) systems, which allow large language models to exchange and critique answers. The study categorizes attacks into a layered taxonomy aligned with the MAD workflow and tests six attack families across 356 source tasks and 3,958 test cases. Results indicate that while MAD can reduce answer accuracy attacks compared to single-agent baselines, it may amplify unauthorized reads or writes, and even with collusion among agents, the final answer changes from correct to wrong only 28.30% of the time.
By Yuwan Liu, Jiaming Zhang, Yue Huang, Sisi Duan
arXiv:2606. 13385v1 Announce Type: cross Abstract: Web agents driven by large language models (LLMs) are increasingly deployed in real-world environments, where they operate over untrusted web content and execute actions with direct consequences.
By Zihao Wang, Yiming Li, Yutong Wu, Zheyu Liu, Kangjie Chen, Fok Kar Wai, Pin-Yu Chen, Vrizlynn L. L. Thing, Bo Li, Dacheng Tao, Tianwei Zhang
arXiv:2510. 20963v2 Announce Type: replace Abstract: Multi-agent debate (MAD) was proposed as a promising approach for ensembling the wisdom of multiple large language models (LLMs) to improve reasoning and provide effective supervision to superhuman LLMs.
By Yongqiang Chen, Gang Niu, James Cheng, Bo Han, Masashi Sugiyama
arXiv:2501. 00745v3 Announce Type: replace-cross Abstract: The increasing integration of Large Language Model (LLM) based search engines has transformed the landscape of information retrieval.
By Xiyang Hu
arXiv:2605. 01133v3 Announce Type: replace-cross Abstract: Large language model (LLM)-powered multi-agent systems (MAS) enable agents to communicate and share information, achieving strong performance on complex tasks.
By Lingxi Zhang, Guangtao Zheng, Hanjie Chen
arXiv:2508. 16481v3 Announce Type: replace Abstract: Ensuring the safe use of agentic systems requires a thorough understanding of the range of malicious behaviors these systems may exhibit.
By Jonathan N\"other, Adish Singla, Goran Radanovic
arXiv:2607. 07474v1 Announce Type: cross Abstract: Agentic red-teaming benchmarks report whether an injected agent was compromised as a single bit: the attack succeeded, or it did not.
By Harry Owiredu-Ashley
MITRE‑SAGE is a multi‑agent retrieval‑augmented generation framework that combines semantic and structural cybersecurity knowledge to enhance large language model question‑answering. It decomposes tasks into query interpretation, evidence retrieval, and answer synthesis, supporting vulnerability assessment, threat profiling, and relationship extraction. The authors also introduce MITRE‑QA, a benchmark of 3,000 question‑answer pairs, and show that MITRE‑SAGE outperforms standalone LLMs and conventional RAG methods, with a lightweight configuration achieving top performance on most tasks.
By Ali Habibzadeh, Farid Feyzi, Reza Ebrahimi Atani
arXiv:2606. 12918v1 Announce Type: cross Abstract: Hierarchical multi-agent systems (MAS) are rapidly being deployed in high-stakes workflows across domains such as finance and software engineering.
By Chejian Xu, Zhaorun Chen, Jingyang Zhang, Freddy Lecue, Avni Kothari, Sarah Tan, Wenbo Guo, Bo Li
arXiv:2602. 09222v2 Announce Type: replace-cross Abstract: Large language model (LLM) based web agents are increasingly deployed to automate complex online tasks by directly interacting with web sites and performing actions on users' behalf.
By Georgios Syros, Evan Rose, Brian Grinstead, Christoph Kerschbaumer, William Robertson, Cristina Nita-Rotaru, Alina Oprea
MITRE‑SAGE is a multi‑agent retrieval‑augmented generation framework that combines semantic and structural cybersecurity knowledge to enhance large language model question‑answering. It decomposes tasks into query interpretation, evidence retrieval, and answer synthesis, supporting vulnerability assessment, threat profiling, and relationship extraction. Experiments show that MITRE‑SAGE outperforms standalone LLMs and conventional RAG methods, with a lightweight Qwen2.5‑based configuration excelling on most benchmark tasks.
By Ali Habibzadeh, Farid Feyzi, Reza Ebrahimi Atani
arXiv:2606. 25476v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have demonstrated remarkable performance across natural language processing tasks, yet their deployment in high-stakes applications raises critical concerns regarding reliability, safety, and trustworthiness.
By Abrar Alotaibi, Raed Mughus, Moataz Ahmed