arXiv AI

Explicit Trajectory Diversity for RL-Based Post-Training of LLM Agents

The paper introduces Trajectory-guided Joint Policy Optimization (TJPO), a reinforcement‑learning framework that explicitly encourages diversity in the trajectories of large‑language‑model agents. By defining task‑specific trajectory descriptors, TJPO measures and optimizes diversity as a set‑level function, avoiding population‑based training. Experiments on Sokoban and ALFWorld demonstrate that TJPO yields diverse, interpretable behaviors while preserving strong task performance.

arXiv Machine Learning
Jun 16

DRA-GRPO: Your GRPO Needs to Know Diverse Reasoning Paths for Mathematical Reasoning

arXiv:2505. 09655v5 Announce Type: replace-cross Abstract: Post-training LLMs with Reinforcement Learning, specifically Group Relative Policy Optimization (GRPO), has emerged as a paradigm for enhancing mathematical reasoning.

By Xiwen Chen, Wenhui Zhu, Peijie Qiu, Xuanzhao Dong, Hao Wang, Haiyu Wu, Huayu Li, Aristeidis Sotiras, Yalin Wang, Abolfazl Razi
arXiv AI
Jul 24

Drive As You Like: Multi-Head Diffusion with Reinforcement Learning for Personalized Driving

arXiv:2508. 16947v2 Announce Type: replace-cross Abstract: Despite significant progress, imitation learning-based autonomous driving planners remain largely restricted to reproducing high-frequency biased behaviors, overlooking the inherent behavioral diversity of human driving.

By Fan Ding, Xuewen Luo, Fucai Ke, Hwa Hui Tew, Susilawati Susilawati, Vishnu Monn Baskaran, Junn Yong Loo
arXiv AI
Aug 19

PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMs

PlanPO introduces a group planning-aware policy optimization method for multi-turn agentic large language models, addressing the issue of advantage collapse caused by treating all successful trajectories equally. By incorporating coarse-to-fine advantage signals that reflect differences in trajectory and turn lengths, PlanPO encourages agents to learn generalizable planning and generation behaviors. Experiments show a 27.2% average improvement over GRPO on benchmarks such as ALFWorld, WebShop, and SciWorld, with minimal extra training cost.

By Dayang Liang, Liyuan He, Xuan Feng, Shuxin Li, Bo An, Yunlong Liu
arXiv AI
Sep 4

Out-of-Distribution Generalisation with Sequence Models in Offline Multi-Agent Reinforcement Learning

The paper investigates zero‑shot task generalisation in offline multi‑agent reinforcement learning by extending sequence‑modeling architectures to support multi‑task observation and action spaces and variable agent counts. It finds that increasing task diversity, rather than merely enlarging the dataset, is the key driver for robust zero‑shot transfer. Experiments on four challenging environments show a 3.2× mean improvement on held‑out tasks compared to single‑task models and outperform strong behaviour‑cloning baselines.

By Oussama Hidaoui, Omer Ebead, Ulrich Armel Mbou Sob, Siddarth Singh, Juan Claude Formanek, Felix Chalumeau, Omayma Mahjoub, Sasha Abramowitz, Ruan John de Kock, Wiem Khlifi, Louay Ben Nessir, Simon Verster Du Toit, Daniel Rajaonarivonivelomanantsoa, Asim Awad Osman, Arnol Manuel Fokam, Refiloe Shabe, Arnu Pretorius