arXiv:2607. 22917v2 Announce Type: replace-cross Abstract: Large Language Model (LLM) agents have significantly improved coding and programming workflows.
By Shouren Wang
arXiv:2608. 16801v1 Announce Type: new Abstract: We study how teams of AI coding agents coordinate while solving programming tasks.
By Giuseppe Destefanis, Tomaso Aste
AgentRoom introduces a real‑time collaborative editing protocol that enables concurrent coding by multiple large language model agents within a CRDT‑backed shared workspace. By providing file‑level claim, status, and broadcast tools, it allows agents to coordinate directly rather than relying on serial phase handoffs or independent sampling. Experiments with five frontier coding‑CLI models show that AgentRoom reduces task abandonment and run‑to‑run variation compared to solo or parallel‑merge approaches, highlighting the importance of coordination over mere parallelism.
By Seonglae Cho, Donghyun Lee
arXiv:2606. 09751v1 Announce Type: new Abstract: Foundation models are moving from response generation into operational roles.
By Arsalan Shahid, Gordon Suttie, Philip Black
arXiv:2607. 26637v1 Announce Type: cross Abstract: Deployed LLM agents increasingly keep their long-term memory as a filesystem: a directory tree of markdown files that the agent itself reads, writes, and reorganizes through generic file tools.
By Sizhe Zhou, Sheldon Yu, Hui Wei, Junda Wu, Siru Ouyang, Yizhu Jiao, Shijia Pan, Julian McAuley, Yu Zhang, Tong Yu, Jiawei Han
The paper investigates how coordination among AI agents serving different users degrades performance compared to a single coordinating agent. Across five advanced models and 77 scenarios in four shared-resource environments—API key budgets, clinic calendars, personal assistant bookings, and merge queues—the study finds that multi‑agent teams consistently underperform, sometimes collapsing entirely, and that even with communication channels coordination overhead remains significant. The authors identify specific failure modes such as stalling, action overriding, and claim fabrication, and propose environment‑specific mitigations like team leads and procedural instructions, while releasing the MAMUBench benchmark for future research.
By Sahan Paliskara, Nattaput Namchittai, Andrew Lampinen