arXiv AI By Sumanyu Muku

Verifying Coordination in Parallel Coding Agents: NP-Bench and a Scheduling Planner

Read the original on arXiv AI →

The paper introduces NP‑Bench, a benchmark and a proactive scheduling planner that coordinates parallel large‑language‑model coding agents. By partitioning work scopes and ordering merges ahead of time, the planner improves clean‑integration rates from 1/9 to 9/9 and eliminates merge conflicts, outperforming both no‑coordination and reactive‑detection baselines. It also demonstrates that cross‑session memory can eliminate repeated mistakes and that routing facts to agents does not improve long‑context accuracy at scale.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
6d ago

Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams

The paper investigates how coordination among AI agents serving different users degrades performance compared to a single coordinating agent. Across five advanced models and 77 scenarios in four shared-resource environments—API key budgets, clinic calendars, personal assistant bookings, and merge queues—the study finds that multi‑agent teams consistently underperform, sometimes collapsing entirely, and that even with communication channels coordination overhead remains significant. The authors identify specific failure modes such as stalling, action overriding, and claim fabrication, and propose environment‑specific mitigations like team leads and procedural instructions, while releasing the MAMUBench benchmark for future research.

By Sahan Paliskara, Nattaput Namchittai, Andrew Lampinen
arXiv AI
Jun 9

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents

arXiv:2606. 08340v1 Announce Type: new Abstract: As language models are increasingly deployed as autonomous agents, they must coordinate with others over long horizons in open-ended interactive tasks.

By Kale-ab Abebe Tessera, Andras Szecsenyi, Cameron Barker, Alexander Rutherford, Davide Paglieri, Aidan Scannell, Henry Gouk, Elliot J. Crowley, Tim Rockt\"aschel, Amos Storkey