arXiv AI By Huaiyu Fu, Heng Cao, Hao Wang, Jian Ya, Tao Chen

Explicit Trajectory Diversity for RL-Based Post-Training of LLM Agents

Read the original on arXiv AI →

The paper introduces Trajectory-guided Joint Policy Optimization (TJPO), a reinforcement‑learning framework that explicitly encourages diversity in the trajectories of large‑language‑model agents. By defining task‑specific trajectory descriptors, TJPO measures and optimizes diversity as a set‑level function, avoiding population‑based training. Experiments on Sokoban and ALFWorld demonstrate that TJPO yields diverse, interpretable behaviors while preserving strong task performance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jun 16

DRA-GRPO: Your GRPO Needs to Know Diverse Reasoning Paths for Mathematical Reasoning

arXiv:2505. 09655v5 Announce Type: replace-cross Abstract: Post-training LLMs with Reinforcement Learning, specifically Group Relative Policy Optimization (GRPO), has emerged as a paradigm for enhancing mathematical reasoning.

By Xiwen Chen, Wenhui Zhu, Peijie Qiu, Xuanzhao Dong, Hao Wang, Haiyu Wu, Huayu Li, Aristeidis Sotiras, Yalin Wang, Abolfazl Razi