arXiv Machine Learning By Cornelius V. Braun, Sayantan Auddy, Marc Toussaint

Trajectory First: A Curriculum for Discovering Diverse Policies

Read the original on arXiv Machine Learning →

arXiv:2506. 01568v4 Announce Type: replace Abstract: Being able to solve a task in diverse ways makes agents more robust to task variations and less prone to local optima.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
3d ago

Explicit Trajectory Diversity for RL-Based Post-Training of LLM Agents

The paper introduces Trajectory-guided Joint Policy Optimization (TJPO), a reinforcement‑learning framework that explicitly encourages diversity in the trajectories of large‑language‑model agents. By defining task‑specific trajectory descriptors, TJPO measures and optimizes diversity as a set‑level function, avoiding population‑based training. Experiments on Sokoban and ALFWorld demonstrate that TJPO yields diverse, interpretable behaviors while preserving strong task performance.

By Huaiyu Fu, Heng Cao, Hao Wang, Jian Ya, Tao Chen
arXiv AI
Sep 4

Out-of-Distribution Generalisation with Sequence Models in Offline Multi-Agent Reinforcement Learning

The paper investigates zero‑shot task generalisation in offline multi‑agent reinforcement learning by extending sequence‑modeling architectures to support multi‑task observation and action spaces and variable agent counts. It finds that increasing task diversity, rather than merely enlarging the dataset, is the key driver for robust zero‑shot transfer. Experiments on four challenging environments show a 3.2× mean improvement on held‑out tasks compared to single‑task models and outperform strong behaviour‑cloning baselines.

By Oussama Hidaoui, Omer Ebead, Ulrich Armel Mbou Sob, Siddarth Singh, Juan Claude Formanek, Felix Chalumeau, Omayma Mahjoub, Sasha Abramowitz, Ruan John de Kock, Wiem Khlifi, Louay Ben Nessir, Simon Verster Du Toit, Daniel Rajaonarivonivelomanantsoa, Asim Awad Osman, Arnol Manuel Fokam, Refiloe Shabe, Arnu Pretorius