arXiv AI By Sen Zhao, Ruiqi Kong, Zuyu Zhang, Lifeng Shen, Xinyu He, Ding Zou, Xu Zhang, Qinghua Zhang

CoMemBench: Benchmarking Collaborative Memory Boundaries across Multi-Agent Workflow Topologies

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
1d ago

Topological Coherence for Self-evolving Multi-agent Systems

The paper introduces TOCOMAS, a Topology‑Coherent Multi‑Agent System that enforces topological coherence—consistent responsibility, handoff, and memory boundaries—within self‑evolving multi‑agent systems. TOCOMAS grounds task graphs in tool interfaces, groups compatible task nodes into reusable responsibility domains, and derives collaboration and memory visibility rules that respect task dependencies. In experiments on BBEH, WorkBench, SWE‑Bench‑Verified, and CoMemBench, TOCOMAS outperforms baseline methods in task success, verified progress, handoffs, and memory isolation.

By Sen Zhao, Ruiqi Kong, Zuyu Zhang, Lifeng Shen, Xinyu He, Xu Zhang, Qinghua Zhang
arXiv AI
Sep 2

UniACE: A Unified Framework for Evaluating LLM Agentic Capabilities

UniACE is a unified framework that standardizes the evaluation of large language model (LLM) agents by representing each benchmark as an instruction–tool–environment triplet and running models through a shared, task‑agnostic harness in isolated runtimes. It preserves native success criteria, offers an offline mode for dynamic‑resource tasks, and standardizes efficiency metrics, execution records, and failure attribution. Applying UniACE to 7 benchmarks across 24 domains and 15 models revealed significant score shifts, ranking reversals, and sensitivity to evidence representation, highlighting the impact of evaluation configuration on reported agent performance.

By Pengyu Zhu, Lijun Li, Yaxing Lyu, Qianxin Luo, Jingyi Yang, Yi Liu, Tingfeng Hui, Xinyu Yuan, Li Sun, Sen Su, Jing Shao
arXiv AI
Jun 2

CoMIC: Collaborative Memory and Insights Circulation for Long-Horizon LLM Agents in Cloud-Edge Systems

arXiv:2606. 00756v1 Announce Type: new Abstract: Deploying lightweight Large Language Model (LLM) agents on edge servers can reduce latency and move agentic services closer to users, but resource-constrained edge models often struggle with long-horizon tasks that require persistent memory, subgoal tracking, and reflection.

By Yannan Wang, Longli Yang, Zhen Liu, Abhishek Kumar, Carsten Maple