arXiv:2601. 07376v2 Announce Type: replace Abstract: We introduce \textsc{OpenTinker}, an open infrastructure for training large language model (LLM) agents with many LoRA-backed policies over shared execution resources.
By Siqi Zhu, Jiaxuan You
arXiv:2608. 16798v1 Announce Type: cross Abstract: Agent harnesses have substantially improved performance on long-horizon tasks by coordinating agent interactions with the environment.
By Huatong Song, Fei Bai, Ming Yang, Renyuan Li, Jia Deng, Jujie He, Zhange Zhang, Daixuan Cheng, Yan Xing, Qi Yun, Xuxing Chen, Danyang Li, Feng Chang, Chuan Hao, Ran Tao, Jian Yang, Bryan Dai, Wayne Xin Zhao, Mingjie Tang, Ji-Rong Wen
arXiv:2512. 09706v2 Announce Type: replace Abstract: The paradigm of agentic AI is shifting from engineered complex workflows to post-training native models.
By Kaichen He, Zihao Wang, Muyao Li, Anji Liu, Yitao Liang
arXiv:2607. 19117v1 Announce Type: new Abstract: Parameterized action reinforcement learning has shown strong performance in environments requiring both discrete action selection and continuous parameterization.
By Ubayd Ali Bapoo, Clement N Nyirenda
arXiv:2606. 30072v1 Announce Type: new Abstract: Cooperative tasks in Multi-Agent Reinforcement Learning (MARL) require agents to collectively maximize a shared return.
By Daiki E. Matsunaga, Junho Na, Tri Wahyu Guntara, Scott Sanner, Pascal Poupart, Jongmin Lee, Kee-Eung Kim
arXiv:2609.39777v1 Announce Type: new
Abstract: LLM-based multi-agent systems coordinate specialized reasoning through aggregation, interaction, and adaptive control, yet their potential for graph le...
By Jiayi Yang, Yifang Chen, Yuanfu Sun, Xinyan Ge, Qiaoyu Tan
The paper introduces SUN (Semantically UNified) Programs, typed executables that translate grounded relations into optimal control objectives, satisfaction predicates, and learning rewards. Using the Kuafu harness, a foundation model orchestrates scene preparation, verification, residual reinforcement learning, and data generation, repairing candidate programs and calibrating reward weights. Across nine multi‑stage manipulation tasks, Kuafu achieves an 82.03% success rate, outperforms learned baselines, generates demonstrations 10.57× faster than human teleoperation, and transfers zero‑shot to physical Franka and Kinova robots.
By Weiqi Wang, Zhi Li, Yudong Lei, David Martinez, Xiaofeng Gao, Yuxin Jiang, Chenfanfu Jiang, Yingnian Wu, Demetri Terzopoulos, Ran Gong
NashDreamer is a new model-based reinforcement learning framework designed for two-player zero-sum imperfect-information games. It introduces a centralized Multi-Agent Recurrent State-Space Model that separates environment dynamics from player strategy effects, enabling the use of any policy gradient algorithm while preserving convergence guarantees to Nash equilibria. Experiments on four benchmark games show that NashDreamer achieves significantly better sample efficiency than model-free baselines early in training, and the authors analyze its optimization landscape, noting a potential vulnerability to posterior collapse in stochastic settings.
By Tom\'a\v{s} Hole\v{c}ek, Viliam Lis\'y
UnifiedPlayers is a cooperative framework that jointly adapts planning, execution, and evaluation for tool-integrated reinforcement learning agents. It consists of a Planning Player that generates tasks, an Execution Player that creates multi-turn trajectories with Python tool calls, and an Evaluation Player that builds executable verifiers, all coordinated by role‑specific rewards under GRPO. The approach outperforms prior baselines on mathematical and general reasoning benchmarks and yields a verifier with high adversarial detection accuracy and more discriminative reward signals.
By Wenjie Liao, Liangjie Zhao, Zehong Cao
arXiv:2509. 24575v2 Announce Type: replace-cross Abstract: This paper presents a framework to prompt multi-robot teams with high-level tasks using natural language expressions.
By Eduardo Sebasti\'an, Nicolas Pfitzer, Ajay Shankar, Amanda Prorok
Deep Reinforcement Learning (RL) is notoriously sample inefficient. One contributing factor is that RL agents are typically initialized from scratch, forcing them to acquire task-relevant knowledge through online interaction.
arXiv:2608. 06015v1 Announce Type: cross Abstract: Deep Reinforcement Learning (RL) is notoriously sample inefficient.
By Xinwei Liu, Junyuan Liang, Jianting Zhang, Wuhui Chen