arXiv AI

ProtocolBench: Which LLM MultiAgent Protocol to Choose?

arXiv:2510. 17149v3 Announce Type: replace Abstract: As large-scale multi-agent systems evolve, the communication protocol layer has become a critical yet under-evaluated factor shaping performance and reliability.

arXiv AI
4d ago

TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing

arXiv:2605.18859v3 Announce Type: replace-cross Abstract: LLM routing matters most in long-horizon applications such as coding agents, deep research systems, and computer-use agents, where a single u...

By Pei Yang, Wanyi Chen, Tongyun Yang, Pengbin Feng, Jiarong Xing, Wentao Guo, Yuhang Yao, Yuhang Han, Hanchen Li, Xu Wang, Zeyu Wang, Jie Xiao, Anjie Yang, Liang Tian, Lynn Ai, Eric Yang, Tianyu Shi
arXiv AI
Sep 25

MeshHeal: Two-Timescale Self-Healing for Gray Failures in Decentralized LLM Agent Networks

MeshHeal is a fully decentralized self‑healing framework for decentralized LLM‑based multi‑agent systems that addresses gray failures—situations where an agent remains responsive but its task‑solving quality degrades. It operates on two timescales: a fast adaptive hierarchy that escalates uncertain or low‑scoring outputs to committee review and correction, and a slow peer‑relative detector that aggregates scores to distinguish persistent degradation from normal variation, triggering mandatory review and eventual exclusion of degraded agents while allowing recovered agents to rejoin. MeshHeal’s evaluation, using Model‑Backed MAS Evaluation, shows it achieves higher degraded‑phase accuracy (0.839) on BBH, MATH, and MMLU‑Pro with fewer tokens per task compared to the baseline Symphony.

By Keru Chen, Sen Lin, Yingbin Liang, Nathaniel D. Bastian, Shaofeng Zou
arXiv AI
Sep 2

UniACE: A Unified Framework for Evaluating LLM Agentic Capabilities

UniACE is a unified framework that standardizes the evaluation of large language model (LLM) agents by representing each benchmark as an instruction–tool–environment triplet and running models through a shared, task‑agnostic harness in isolated runtimes. It preserves native success criteria, offers an offline mode for dynamic‑resource tasks, and standardizes efficiency metrics, execution records, and failure attribution. Applying UniACE to 7 benchmarks across 24 domains and 15 models revealed significant score shifts, ranking reversals, and sensitivity to evidence representation, highlighting the impact of evaluation configuration on reported agent performance.

By Pengyu Zhu, Lijun Li, Yaxing Lyu, Qianxin Luo, Jingyi Yang, Yi Liu, Tingfeng Hui, Xinyu Yuan, Li Sun, Sen Su, Jing Shao