arXiv:2602.18998v2 Announce Type: replace
Abstract: LLM agents are increasingly expected to operate as general-purpose systems that resolve real-world user requests, yet their dynamic scaling behavio...
By Xiaochuan Li, Ryan Ming, Pranav Setlur, Abhijay Paladugu, Andy Tang, Hao Kang, Shuai Shao, Rong Jin, Chenyan Xiong
arXiv:2604. 02863v2 Announce Type: replace Abstract: Majority voting is the standard for aggregating multi-agent responses into a final decision.
By Yiqing Liu, Hantao Yao, Wu Liu, Yongdong Zhang
arXiv:2602. 16745v2 Announce Type: replace-cross Abstract: Test-time scaling can improve model performance by aggregating stochastic reasoning trajectories.
By Zhangyi Liu, Huaizhi Qu, Xiaowei Yin, He Sun, Yanjun Han, Tianlong Chen, Zhun Deng
MiniRep is a reputation‑based aggregation system designed for multi‑agent debate (MAD) that remains robust even when malicious agents are present. It evaluates agents on both their current task performance and historical reputation, while preventing groups of agents with highly similar responses from dominating the final decision. Experiments on the MATH benchmark show that MiniRep consistently outperforms conventional MAD aggregation and other reputation‑based approaches across a wide range of attack scenarios.
By Jiaming Zhang, Yuwan Liu, Yue Huang, Sisi Duan
arXiv:2601. 19921v2 Announce Type: replace-cross Abstract: Multi-agent debate (MAD) is widely used to improve large language model (LLM) performance through test-time scaling, yet recent work shows that vanilla MAD often underperforms simple majority vote despite higher computational cost.
By Xiaochen Zhu, Caiqi Zhang, Yizhou Chi, Tom Stafford, Nigel Collier, Andreas Vlachos
arXiv:2511. 00802v2 Announce Type: replace-cross Abstract: With data-driven development now widely adopted, online A/B testing is an established method for measuring the effects of new technologies.
By Jie JW Wu, Ayanda Patrick Herlihy, Ahmad Saleem Mirza, Ali Afoud, Fatemeh Fard