UniACE is a unified framework that standardizes the evaluation of large language model (LLM) agents by representing each benchmark as an instruction–tool–environment triplet and running models through a shared, task‑agnostic harness in isolated runtimes. It preserves native success criteria, offers an offline mode for dynamic‑resource tasks, and standardizes efficiency metrics, execution records, and failure attribution. Applying UniACE to 7 benchmarks across 24 domains and 15 models revealed significant score shifts, ranking reversals, and sensitivity to evidence representation, highlighting the impact of evaluation configuration on reported agent performance.
By Pengyu Zhu, Lijun Li, Yaxing Lyu, Qianxin Luo, Jingyi Yang, Yi Liu, Tingfeng Hui, Xinyu Yuan, Li Sun, Sen Su, Jing Shao
The paper investigates how scaling a team of small language‑model agents affects performance across different orchestration architectures. By testing eight architectures on five short‑answer benchmarks and an executable‑code benchmark, it finds that team scaling yields large gains on arithmetic word‑problem tasks but only modest improvements on multiple‑choice and code generation tasks, with no single architecture dominating all tasks. The authors explain these patterns using a generate‑transform decomposition that separates coverage and transformation effects, showing that arithmetic tasks benefit from both coverage and critic‑guided transformation, while other tasks are limited by saturation or poor conversion.
By Blaz Bertalanic, Carolina Fortuna
arXiv:2606. 10662v1 Announce Type: cross Abstract: Multi-agent systems (MAS) can scale large language model reasoning at test time by decomposing complex problems into parallel subtasks.
By Yuzhen Mao, Azalia Mirhoseini
The paper investigates how coordination among AI agents serving different users degrades performance compared to a single coordinating agent. Across five advanced models and 77 scenarios in four shared-resource environments—API key budgets, clinic calendars, personal assistant bookings, and merge queues—the study finds that multi‑agent teams consistently underperform, sometimes collapsing entirely, and that even with communication channels coordination overhead remains significant. The authors identify specific failure modes such as stalling, action overriding, and claim fabrication, and propose environment‑specific mitigations like team leads and procedural instructions, while releasing the MAMUBench benchmark for future research.
By Sahan Paliskara, Nattaput Namchittai, Andrew Lampinen
arXiv:2608.22510v1 Announce Type: new
Abstract: Agent benchmarks often evaluate only final answers even when agents run on stateful runtimes. We argue this under-specifies what is being evaluated: th...
By YuanHang Xiao
arXiv:2606. 08340v1 Announce Type: new Abstract: As language models are increasingly deployed as autonomous agents, they must coordinate with others over long horizons in open-ended interactive tasks.
By Kale-ab Abebe Tessera, Andras Szecsenyi, Cameron Barker, Alexander Rutherford, Davide Paglieri, Aidan Scannell, Henry Gouk, Elliot J. Crowley, Tim Rockt\"aschel, Amos Storkey