arXiv AI By Nicholas Lee, Lutfi Eren Erdogan, Chris Joseph John, Surya Krishnapillai, Michael W. Mahoney, Kurt Keutzer, Amir Gholami

Agentic Test-Time Scaling for WebAgents

Read the original on arXiv AI →

arXiv:2602. 12276v2 Announce Type: replace Abstract: Test-time scaling has become a standard way to improve performance and boost reliability of neural network models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
3d ago

MiniRep: Robust Reputation-Based Aggregation for Multi-Agent Debate

MiniRep is a reputation‑based aggregation system designed for multi‑agent debate (MAD) that remains robust even when malicious agents are present. It evaluates agents on both their current task performance and historical reputation, while preventing groups of agents with highly similar responses from dominating the final decision. Experiments on the MATH benchmark show that MiniRep consistently outperforms conventional MAD aggregation and other reputation‑based approaches across a wide range of attack scenarios.

By Jiaming Zhang, Yuwan Liu, Yue Huang, Sisi Duan
arXiv AI
Jun 2

Demystifying Multi-Agent Debate: The Role of Confidence and Diversity

arXiv:2601. 19921v2 Announce Type: replace-cross Abstract: Multi-agent debate (MAD) is widely used to improve large language model (LLM) performance through test-time scaling, yet recent work shows that vanilla MAD often underperforms simple majority vote despite higher computational cost.

By Xiaochen Zhu, Caiqi Zhang, Yizhou Chi, Tom Stafford, Nigel Collier, Andreas Vlachos