arXiv Machine Learning By Yi Xie, Zhanke Zhou, Chentao Cao, Bo Liu, Bo Han

DICE: Entropy-Regularized Equilibrium Selection for Stable Multi-Agent LLM Coordination

Read the original on arXiv Machine Learning →

arXiv:2606. 08068v1 Announce Type: new Abstract: Multi-agent large language model (LLM) systems often fail to reliably outperform a single strong model equipped with best-of-N sampling.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 3

Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems

The paper introduces a game-theoretic framework for coordinating multi-agent large language model (LLM) systems, modeling orchestrator–worker interactions as a bilevel coordination game. It analyzes textual reflection as stochastic movement over semantic memory states, providing finite-time bounds, lower bounds, and an information-theoretic impossibility result. Building on these insights, the authors propose Stochastic Reflective Memory Ascent (SRMA), a grounded evaluation mechanism that guarantees convergence under certain conditions and demonstrates improved performance on 500 SWE‑bench instances.

By Yihang Chen, Yuxiang Chen, Yuxuan Huang, Meng Fang, Weilin Luo, Jun Wang
arXiv AI
Jun 2

S-SPPO: Semantic-Calibrated Self-Play Preference Optimization

arXiv:2606. 01561v1 Announce Type: new Abstract: Aligning Large Language Models (LLMs) with human preferences is often formulated via Direct Preference Optimization (DPO).

By Xiwen Chen, Wenhui Zhu, Jingjing Wang, Peijie Qiu, Zhipeng Wang, Huayu Li, ZhengXiao He, Xuanzhao Dong, Prayag Tiwari, Mingkun Xu, Yujian Xiong, Feng Luo, Abolfazl Razi, Brendan Hogan Rappazzo, Anderson Schneider, Yuriy Nevmyvaka
arXiv AI
Jul 7

Regime-Conditional Stabilisation of LLM-Augmented Cooperative Multi-Agent Reinforcement Learning

arXiv:2607. 04470v1 Announce Type: cross Abstract: Large Language Models (LLMs) offer a natural interface for translating human objectives into reward signals for cooperative multi-agent reinforcement learning (MARL), yet the training-time dynamics of this integration remain poorly understood.

By Faid Keddouri, Sohaib Houhou, Aissa Boulmerka, Nadir Farhi