Collaborative Human-Agent Protocol (CHAP)
arXiv:2606. 09751v1 Announce Type: new Abstract: Foundation models are moving from response generation into operational roles.
arXiv:2606. 14445v1 Announce Type: cross Abstract: Existing multi-agent software development systems have proposed many forms of agent collaboration, including role-based collaboration and automated code review.
arXiv:2606. 09751v1 Announce Type: new Abstract: Foundation models are moving from response generation into operational roles.
The paper investigates challenges in open‑source large‑language‑model (LLM) based multi‑agent systems (MAS). By analyzing 944 issues extracted from 21 projects, it finds that orchestration and execution problems are most common, with workflow, tool integration, and memory issues as primary causes. The predominant remedy identified is optimizing workflow, and the study offers empirically grounded implications for improving orchestration, tool integration, and memory mechanisms in LLM‑based MAS.
OpenCollab is a multi‑agent coding framework that unifies organization design, enforces experimental control on a shared runtime, and tracks execution via fine‑grained event streams. It introduces the metric Adherence to measure whether the declared organization is actually realized, showing that small configuration changes can shift adherence from 47.2% to 97.2%. Experiments demonstrate that a two‑coder workflow built on OpenCollab achieves new state‑of‑the‑art performance against mainstream harnesses while using the fewest tokens, and that a well‑designed organization can outperform strong existing harnesses.
AgentRoom introduces a real‑time collaborative editing protocol that enables concurrent coding by multiple large language model agents within a CRDT‑backed shared workspace. By providing file‑level claim, status, and broadcast tools, it allows agents to coordinate directly rather than relying on serial phase handoffs or independent sampling. Experiments with five frontier coding‑CLI models show that AgentRoom reduces task abandonment and run‑to‑run variation compared to solo or parallel‑merge approaches, highlighting the importance of coordination over mere parallelism.
arXiv:2606. 26924v1 Announce Type: cross Abstract: LLM coding harnesses grant agents broad file and shell access, yet the configuration layer that steers them -- rules files, agent definitions, IDE-specific markdown -- is largely unmanaged.
arXiv:2606. 19616v1 Announce Type: cross Abstract: Autonomous coding agents now open millions of pull requests, yet large-scale studies find their PRs are produced faster but accepted less often - a coordination and trust gap that pull-request-level telemetry cannot explain.
arXiv:2605.29313v2 Announce Type: replace Abstract: LLM multi-agent systems often coordinate through natural-language dialogue or loosely structured shared memory, making intermediate state difficult...
Dr. Claw is an open‑source AI scientist workspace that integrates existing command‑line coding agents into a single, auditable, human‑in‑the‑loop workflow. It uses persistent state objects, a reusable skill library, and multi‑executor coordination to link human decisions with AI execution, creating a traceable and recoverable loop for planning, execution, and writing. The authors demonstrate the system with an interactive scenario and a failure‑recovery walkthrough, and show that, when the underlying executor is held constant, Dr. Claw achieves higher research completeness while preserving an auditable process trail.
arXiv:2608. 20195v1 Announce Type: cross Abstract: Technical documentation is written for human developers, but an increasing share of software changes is now authored by autonomous coding agents.
arXiv:2609.32965v2 Announce Type: replace Abstract: Multiple agents may often conflict in an organization: for example, one coding agent changes an interface in a repository, but another continues to...
The paper investigates how third‑party API routers, which sit between coding agents and large language model providers, can introduce a control gap by inspecting and modifying requests and responses. Through an empirical study using the SIDEL framework, the authors evaluate four levels of router‑side injection (Response Substitution, Response Append, LLM‑Polished Injection, and LLM‑Polished with Distribution Alignment Injection) across 400 curated samples and four representative coding agents. The results show that router‑side interventions significantly alter repository‑level actions and evade existing client‑side safeguards, achieving a 0% defense success rate without additional mitigations.
arXiv:2606. 24429v1 Announce Type: cross Abstract: Generative AI coding agents are entering the open-source supply chain, yet their diverse and often invisible traces leave their prevalence poorly understood.