OpenCollab is a multi‑agent coding framework that unifies organization design, enforces experimental control on a shared runtime, and tracks execution via fine‑grained event streams. It introduces the metric Adherence to measure whether the declared organization is actually realized, showing that small configuration changes can shift adherence from 47.2% to 97.2%. Experiments demonstrate that a two‑coder workflow built on OpenCollab achieves new state‑of‑the‑art performance against mainstream harnesses while using the fewest tokens, and that a well‑designed organization can outperform strong existing harnesses.
By Chun-Wah Hsu, Kai Gong, Yu Wu, Xianhe Chen, Mengyang Liu, Jie Li, Hanyu Li, Zhixuan Liu, Naisheng Tang, Jiaying Chi, Ziheng Fan, Xuning He, Xiaokang Yang, Xue Jiang, Yihong Dong
arXiv:2606. 00953v1 Announce Type: new Abstract: Multi-agent Large Language Model (LLM) systems offer a way to decompose complex tasks, such as coding, through parallelization and context isolation.
By Xu Yang, Lunyiu Nie, Ethan Chandra, Stanislav Gannutin, Fangru Lin, Swarat Chaudhuri
arXiv:2609.32490v2 Announce Type: replace
Abstract: LLM-based multi-agent systems (MASs) have shown strong potential for solving complex tasks, but most assume that task requirements are sufficiently...
By Yuchen Song, Andong Chen, Wenxin Zhu, Muyun Yang, Tiejun Zhao
arXiv:2606. 01533v1 Announce Type: cross Abstract: Computer use agents (CUAs) today are primarily deployed as single serial agents.
By Jing Yu Koh, Ruslan Salakhutdinov, Daniel Fried
arXiv:2608.29204v1 Announce Type: cross
Abstract: Generative AI-based software engineering agents are becoming routine contributors to real-world software projects. On GitHub, developers can assign t...
By Jonan Richards, Kosei Horikawa, Youmei Fan, Yutaro Kashiwa, Mairieli Wessel
The paper introduces Harness Primitives—reusable agent harness mechanisms mined from failed task trajectories—and a framework called STITCH that selects and composes these primitives into task‑specific harnesses at test time. This approach avoids generating or debugging harness code for each task, achieving up to 12‑point gains in task success over fixed harness baselines and outperforming human‑designed harnesses like Codex CLI. STITCH also demonstrates minimal test‑time overhead (2.7%) and scales efficiently with the size of the primitive library.
By Peng Kuang, Haibo Jin, Dehao Wu, Feiyang Deng, Xiaopeng Yuan, Jerry Wang, Haohan Wang