arXiv AI

CalBench: Evaluating Coordination-Privacy Trade-offs in Multi-Agent LLMs

arXiv:2605. 09823v3 Announce Type: replace-cross Abstract: Personal AI assistants are beginning to act as delegates with access to calendars, inboxes, and user preferences.

arXiv AI
Jun 9

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents

arXiv:2606. 08340v1 Announce Type: new Abstract: As language models are increasingly deployed as autonomous agents, they must coordinate with others over long horizons in open-ended interactive tasks.

By Kale-ab Abebe Tessera, Andras Szecsenyi, Cameron Barker, Alexander Rutherford, Davide Paglieri, Aidan Scannell, Henry Gouk, Elliot J. Crowley, Tim Rockt\"aschel, Amos Storkey