arXiv AI By Sara Vera Marjanovi\'c, Jiacheng Xu, Aleksandr Laptev, Grigor Nalbandyan, Erik Arakelyan, Evelina Bakhaturina

Mo' Models, Mo' Problems: How to best select model pools when designing Multi-Agent Systems

Read the original on arXiv AI →

The paper investigates how to choose model pools for Multi-Agent Systems (MAS) that combine multiple model outputs to tackle complex reasoning tasks. It evaluates eight selection strategies—such as model size, accuracy, and answer diversity—across both before-generation (routing) and after-generation (majority-voting, LLM-as-a-judge) MAS architectures on scientific benchmarks. The study finds that expanding the candidate pool often harms performance, that selecting candidates within a single model family yields the best relative gains, and that indiscriminate addition of heterogeneous models can destabilize the system.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
22h ago

You're Hired: Strategic Model Selection for LLM Collaboration

arXiv:2609.38816v1 Announce Type: new Abstract: While multi-agent and model collaboration algorithms gain traction to combine the strengths of diverse Large Language Models (LLMs), existing systems r...

By Zongwan Cao, Ziyuan Yang, Shangbin Feng, Michael Duan, Skyler Hallinan, Bingbing Wen, Lucy Lu Wang, Yulia Tsvetkov
arXiv AI
Aug 28

DIANOIA: Diagnostic Decomposition and Joint Optimization for Multi-Agent Reasoning

DIANOIA introduces a diagnostic framework for multi‑agent large language model systems, decomposing reasoning gain into three measurable channels—coverage, fidelity, and synthesis. The protocol identifies bottleneck channels for a given task and implements a corresponding multi‑agent system with role‑diverse proposers, execution‑grounded verification, and iterative synthesis. Experiments on GSM8K, AIME‑2025, MBPP, and BFCL‑SP show that DIANOIA outperforms strong baselines, achieving significant token savings and accuracy gains while accurately pinpointing the critical channels.

By Yiming Yang, Zhuoyuan Li, Fanxiang Zeng, Hao Fu, Yue Liu
arXiv AI
Sep 18

Architectural Design, Not Only Model Intelligence, Governs Multi-Agent LLM Performance

The paper argues that the architecture of multi‑agent large language model (LLM) frameworks, rather than just the intelligence of the underlying models, largely determines system performance. It introduces a taxonomy of architectural dimensions—such as orchestration, memory, planning interfaces, specialization, and communication topology—and presents MAFBench, a unified evaluation suite. An empirical study across nine frameworks, keeping the LLM constant, reveals six design principles and shows that choices like orchestration and communication topology can dramatically affect latency, accuracy, and coordination success.

By Abdelghny Orogat, Ana Rostam, Essam Mansour
arXiv AI
Jul 3

PACE: A Proxy for Agentic Capability Evaluation

arXiv:2607. 02032v1 Announce Type: new Abstract: Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure.

By Yueqi Song, Lintang Sutawika, Jiarui Liu, Lindia Tjuatja, Jiayi Geng, Yunze Xiao, Daniel Lee, Aditya Bharat Soni, Vincent Lo, Xiang Yue, Graham Neubig