arXiv AI By Zhenzhen Ren, Jiyan He, Xinpeng Zhang, Zhenxing Qian, Ke Han, Shuxin Zheng, GuoBiao Li, Xiaoqing Zhang

OrchBench: Evaluating Multi-Agent Orchestration Plans in Isolation via Deterministic Simulation

Read the original on arXiv AI →

arXiv:2607. 25656v1 Announce Type: new Abstract: Complex tasks often decompose into parallelizable yet interdependent subtasks, making orchestration critical to the performance of multi-agent systems (MAS).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 2

UniACE: A Unified Framework for Evaluating LLM Agentic Capabilities

UniACE is a unified framework that standardizes the evaluation of large language model (LLM) agents by representing each benchmark as an instruction–tool–environment triplet and running models through a shared, task‑agnostic harness in isolated runtimes. It preserves native success criteria, offers an offline mode for dynamic‑resource tasks, and standardizes efficiency metrics, execution records, and failure attribution. Applying UniACE to 7 benchmarks across 24 domains and 15 models revealed significant score shifts, ranking reversals, and sensitivity to evidence representation, highlighting the impact of evaluation configuration on reported agent performance.

By Pengyu Zhu, Lijun Li, Yaxing Lyu, Qianxin Luo, Jingyi Yang, Yi Liu, Tingfeng Hui, Xinyu Yuan, Li Sun, Sen Su, Jing Shao
arXiv AI
4d ago

An Exact Generate - Transform Decomposition of Small-LLM Team Scaling Across Orchestration Architectures

The paper investigates how scaling a team of small language‑model agents affects performance across different orchestration architectures. By testing eight architectures on five short‑answer benchmarks and an executable‑code benchmark, it finds that team scaling yields large gains on arithmetic word‑problem tasks but only modest improvements on multiple‑choice and code generation tasks, with no single architecture dominating all tasks. The authors explain these patterns using a generate‑transform decomposition that separates coverage and transformation effects, showing that arithmetic tasks benefit from both coverage and critic‑guided transformation, while other tasks are limited by saturation or poor conversion.

By Blaz Bertalanic, Carolina Fortuna
arXiv Computation and Language
Sep 23

Agensh: Scaling Organizational Intelligence to 1,024 Agents

Agensh is a new multi‑agent harness that eliminates a central orchestrator by letting workers self‑organize through a continuous cooperation loop. The system uses a shared workspace, message interface, and shared context to coordinate tasks, verify results, and merge progress asynchronously. Experiments on ProgramBench and pandoc show that scaling from 1 to 1,024 agents improves test‑pass rates by up to 49% relative, demonstrating that agent count is a viable scaling dimension for complex tasks.

By Zhihao Zhan, Ting Song, Li Dong, Shaohan Huang, Jianxun Lian, Yan Xia, Furu Wei
arXiv AI
Aug 26

PeakBench: Benchmarking Resource-Aware Tool Invocation in LLM Agents

PeakBench is a new benchmark designed to evaluate how large language model agents invoke multiple tools while respecting resource constraints and parallel execution. It provides executable multi‑tool workflows with dependency annotations and measured resource profiles, and introduces a two‑part evaluation framework that separates logical planning from physical scheduling. The study shows that strong logical planning alone does not guarantee safe or efficient execution, and that providing resource information can reduce overflows and improve utilization.

By Zhi-Kai Chen, Xu-Xiang Zhong, Song-Yan Li, De-Chuan Zhan, Han-Jia Ye