OpenCollab is a multi‑agent coding framework that unifies organization design, enforces experimental control on a shared runtime, and tracks execution via fine‑grained event streams. It introduces the metric Adherence to measure whether the declared organization is actually realized, showing that small configuration changes can shift adherence from 47.2% to 97.2%. Experiments demonstrate that a two‑coder workflow built on OpenCollab achieves new state‑of‑the‑art performance against mainstream harnesses while using the fewest tokens, and that a well‑designed organization can outperform strong existing harnesses.
By Chun-Wah Hsu, Kai Gong, Yu Wu, Xianhe Chen, Mengyang Liu, Jie Li, Hanyu Li, Zhixuan Liu, Naisheng Tang, Jiaying Chi, Ziheng Fan, Xuning He, Xiaokang Yang, Xue Jiang, Yihong Dong
arXiv:2606. 00953v1 Announce Type: new Abstract: Multi-agent Large Language Model (LLM) systems offer a way to decompose complex tasks, such as coding, through parallelization and context isolation.
By Xu Yang, Lunyiu Nie, Ethan Chandra, Stanislav Gannutin, Fangru Lin, Swarat Chaudhuri
arXiv:2609.32490v2 Announce Type: replace
Abstract: LLM-based multi-agent systems (MASs) have shown strong potential for solving complex tasks, but most assume that task requirements are sufficiently...
By Yuchen Song, Andong Chen, Wenxin Zhu, Muyun Yang, Tiejun Zhao
arXiv:2606. 01533v1 Announce Type: cross Abstract: Computer use agents (CUAs) today are primarily deployed as single serial agents.
By Jing Yu Koh, Ruslan Salakhutdinov, Daniel Fried
arXiv:2608.29204v1 Announce Type: cross
Abstract: Generative AI-based software engineering agents are becoming routine contributors to real-world software projects. On GitHub, developers can assign t...
By Jonan Richards, Kosei Horikawa, Youmei Fan, Yutaro Kashiwa, Mairieli Wessel
The paper introduces Harness Primitives—reusable agent harness mechanisms mined from failed task trajectories—and a framework called STITCH that selects and composes these primitives into task‑specific harnesses at test time. This approach avoids generating or debugging harness code for each task, achieving up to 12‑point gains in task success over fixed harness baselines and outperforming human‑designed harnesses like Codex CLI. STITCH also demonstrates minimal test‑time overhead (2.7%) and scales efficiently with the size of the primitive library.
By Peng Kuang, Haibo Jin, Dehao Wu, Feiyang Deng, Xiaopeng Yuan, Jerry Wang, Haohan Wang
arXiv:2606. 03103v1 Announce Type: new Abstract: Real-world professional desktop workflows in specialized creative and engineering software unfold over long horizons and often require human-in-the-loop coordination, where agents proactively seek necessary information and users provide additional instructions, clarifications, feedback, or corrections as the task progresses.
By Wenkai Wang, Tao Xiong, Jingchen Ni, Yunpeng Bao, Xiyun Li, Tianqi Liu, Hongcan Guo, Zilong Huang, Shengyu Zhang
arXiv:2607. 13034v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly automate multi-step engineering and informatics workflows, yet they rarely ask how much effort a task actually requires.
By Junjie Yin, Xinyu Feng
arXiv:2609.37143v1 Announce Type: cross
Abstract: Modern coding agents can deliver increasingly large repository-level changes, and recent benchmarks reflect this by emphasizing long-horizon tasks wi...
By Yun Peng, Zihan Wu, Zeyang Zhuang, Xin Zhou, Rui Shu, Xu Han, Chun Yong Chong, Yuan Wang, Jiakun Liu
arXiv:2605.27995v3 Announce Type: replace
Abstract: Large language model (LLM)-based agents have shown strong capabilities in using external tools to solve complex tasks. However, existing evaluation...
By Kou Shi, Ziao Zhang, Shiting Huang, Avery Nie, Zhen Fang, Qiuchen Wang, Lin Chen, Huaian Chen, Zehui Chen, Feng Zhao
arXiv:2607. 10059v1 Announce Type: new Abstract: Agent systems based on large language models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents know when to abstain.
By Xun Liu, Yi Evie Zhang, Vira Kasprova, Parisa Rabbani, Pardis Sadat Zahraei, Tianyu Zhang, Ali Ebrahimpour-Boroojeny, Varun Chandrasekaran
arXiv:2606. 10662v1 Announce Type: cross Abstract: Multi-agent systems (MAS) can scale large language model reasoning at test time by decomposing complex problems into parallel subtasks.
By Yuzhen Mao, Azalia Mirhoseini