The paper introduces ROTATE, a regret-driven open‑ended training framework that jointly improves an Ad Hoc Teamwork (AHT) agent and an adversarial teammate generator. Unlike traditional two‑stage pipelines, ROTATE alternates between enhancing the agent and generating teammates that specifically probe its collaboration weaknesses. Experiments on Overcooked and Level‑Based Foraging show that ROTATE outperforms existing baselines on unseen teammates, setting a new benchmark for robust, generalizable teamwork.
By Caroline Wang, Arrasy Rahman, Benjamin Nativi, Jiaxun Cui, Yoonchang Sung, Peter Stone
arXiv:2602. 17737v2 Announce Type: replace-cross Abstract: Mutual adaptation is a central challenge in human-AI teaming, as humans naturally adjust their strategies in response to an AI agent's behavior.
By Upasana Biswas, Durgesh Kalwar, Subbarao Kambhampati, Sarath Sreedharan
JaxAHT is a new open‑source library built with JAX that speeds up and standardizes research in Ad Hoc Teamwork (AHT). It offers a unified framework for generating teammates, training ego agents, and evaluating performance against unseen partners, delivering roughly 95× faster wall‑clock times than comparable PyTorch implementations. The library also supplies a diverse set of evaluation teammates for Level‑Based Foraging, Overcooked, and Hanabi, and is used to run a large‑scale benchmark that shows no single algorithm dominates and that agent modeling mainly helps in role‑based, diverse teammate settings.
By Caroline Wang, Rolando Fernandez, Zelal Su Mustafaoglu, Montek Kundan, Jiaxun Cui, Lingyun Xiao, Zhihan Wang, Di Yang Shi, Aditya Madhan, Johnny Liu, Arrasy Rahman, Peter Stone
arXiv:2604. 00830v3 Announce Type: replace-cross Abstract: Test-Time Learning (TTL) enables language agents to iteratively refine their performance through repeated interactions with the environment at inference time.
By Zhanzhi Lou, Hui Chen, Yibo Li, Qian Wang, Bryan Hooi
arXiv:2607. 27177v1 Announce Type: new Abstract: Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents.
By Peter Tisnikar, Maja Swieczkowska, Benteng Ma, Gerard Canal, Matteo Leonetti
arXiv:2601. 19810v2 Announce Type: replace-cross Abstract: Unsupervised pre-training can equip reinforcement learning agents with prior knowledge and accelerate learning in downstream tasks.
By Octavio Pappalardo
arXiv:2606. 03698v1 Announce Type: new Abstract: A central goal of large language model (LLM) research is to build agentic systems that can plan, act, and adapt through sustained interaction with dynamic environments.
By Sangeun Park, Minhae Kwon
arXiv:2605.01457v4 Announce Type: replace
Abstract: How can generative offline multi-agent reinforcement learning achieve both fast joint trajectory generation and effective cooperation? Multi-agent...
By Guowei Zou, Haitao Wang, Beiwen Zhang, Boning Zhang, Hejun Wu
arXiv:2609.40048v1 Announce Type: new
Abstract: Ultra-long video temporal grounding requires balancing long-range evidence search with fine-grained event understanding under a limited visual budget,...
By Yiduo Jia, Muzhi Zhu, Jinchuan Shi, Hao Zhong, Yuling Xi, Ke Liu, Hao Chen
UnifiedPlayers is a cooperative framework that jointly adapts planning, execution, and evaluation for tool-integrated reinforcement learning agents. It consists of a Planning Player that generates tasks, an Execution Player that creates multi-turn trajectories with Python tool calls, and an Evaluation Player that builds executable verifiers, all coordinated by role‑specific rewards under GRPO. The approach outperforms prior baselines on mathematical and general reasoning benchmarks and yields a verifier with high adversarial detection accuracy and more discriminative reward signals.
By Wenjie Liao, Liangjie Zhao, Zehong Cao
arXiv:2608. 07959v1 Announce Type: new Abstract: Ultra-long egocentric video understanding requires reasoning over temporally sparse evidence distributed across hours or days, challenging current multimodal models with limited context and the grounding of key video segments.
By Keyang Zhong, Kuo Wang, Peng Liu, Quanlong Zheng, Junlin Xie, Zhijia Liang, Yanhao Zhang, Guanbin Li
arXiv:2608. 06381v1 Announce Type: cross Abstract: Explainable AI (XAI) has shown promise for human-agent collaboration, yet results rely on hand-crafted policies in custom environments, limiting generalizability to state-of-the-art teaming research.
By Mateus Levi Sim\~oes Fernandes, Alberto Sardinha