arXiv:2501. 14622v5 Announce Type: replace Abstract: Learning efficient representations for decision-making policies is a challenge in imitation learning (IL).
By Aleksandar Vujinovic, Aleksandar Kovacevic
arXiv:2512. 09706v2 Announce Type: replace Abstract: The paradigm of agentic AI is shifting from engineered complex workflows to post-training native models.
By Kaichen He, Zihao Wang, Muyao Li, Anji Liu, Yitao Liang
arXiv:2607. 27973v1 Announce Type: new Abstract: Recently, Reinforcement Learning (RL) has emerged as a crucial paradigm for the post-training of Large Language Model (LLM) agents.
By Cong Li, Peixi Peng, Yisen Zhao, Xinyu Hu, Shudong Liu, Zhan Su, Zhuojian Li
The paper introduces a reinforcement learning post‑training scheme that trains robot world models on their own autoregressive rollouts, using a contrastive RL objective adapted from diffusion models. It also proposes a training protocol that compares multiple variable‑length futures, a multi‑view visual fidelity reward, and demonstrates state‑of‑the‑art rollout fidelity on the DROID dataset, outperforming baselines on LPIPS, SSIM, and human preference tests.
By Jai Bardhan, Patrik Drozdik, Josef Sivic, Vladimir Petrik
World models enable agents to reason about future outcomes and learn policies from their knowledge of state transition, but existing approaches primarily focus on reconstructing future observations or...
OneBid is a unified auto‑bidding foundation model that consolidates diverse cost‑per‑X (oCPX) advertising scenarios into a single framework. It builds on Decision Transformer by conditioning on two atomic signals—Return‑to‑Go for conversion value and Cost‑to‑Go for cost ratio—and incorporates value‑aware regularization. A sequence‑level Mixture‑of‑Experts architecture captures cross‑scenario knowledge while preserving low latency, and a Critic‑guided Relative Offline Policy optimization (CROP) aligns the backbone with scenario‑specific preferences without unsafe online exploration. In production at Kuaishou, OneBid achieved a 2.2% overall ADVV increase and up to 13.1% in the ROAS scenario.
By Yewen Li, Peng Jiang, Yitian Li, Pengfei Lv, Xialong Liu, Peng Jiang, Qingpeng Cai
arXiv:2606. 00780v1 Announce Type: cross Abstract: Offline meta-reinforcement learning leverages static datasets to enable agents to generalize to unseen environments by combining offline efficiency with meta-learning adaptability, yet it faces key challenges from context and policy distribution shifts.
By Fuyuan Qian, Menglong Zhang, Song Wang, Quanying Liu
The paper investigates zero‑shot task generalisation in offline multi‑agent reinforcement learning by extending sequence‑modeling architectures to support multi‑task observation and action spaces and variable agent counts. It finds that increasing task diversity, rather than merely enlarging the dataset, is the key driver for robust zero‑shot transfer. Experiments on four challenging environments show a 3.2× mean improvement on held‑out tasks compared to single‑task models and outperform strong behaviour‑cloning baselines.
By Oussama Hidaoui, Omer Ebead, Ulrich Armel Mbou Sob, Siddarth Singh, Juan Claude Formanek, Felix Chalumeau, Omayma Mahjoub, Sasha Abramowitz, Ruan John de Kock, Wiem Khlifi, Louay Ben Nessir, Simon Verster Du Toit, Daniel Rajaonarivonivelomanantsoa, Asim Awad Osman, Arnol Manuel Fokam, Refiloe Shabe, Arnu Pretorius
arXiv:2502. 19544v3 Announce Type: replace Abstract: Leveraging offline data is a promising way to improve the sample efficiency of online reinforcement learning (RL).
By Yi Zhao, Aidan Scannell, Wenshuai Zhao, Yuxin Hou, Tianyu Cui, Le Chen, Dieter B\"uchler, Arno Solin, Juho Kannala, Joni Pajarinen
arXiv:2606. 24133v1 Announce Type: new Abstract: The composition of training data, governed by the diversity of sources and their mixing strategy, is a cornerstone of Large Language Model (LLM) pre-training.
By Chenhao Dang, Jing Ma, Mingjie Liao
arXiv:2610.01882v1 Announce Type: cross
Abstract: Multi-agent reinforcement learning (MARL) provides a powerful framework for learning coordinated behaviors through interactions with the environment....
By Zhuoran Li, Yunzhan Li, Xun Wang, Yihan Du, Longbo Huang
arXiv:2603. 01891v2 Announce Type: replace Abstract: Action chunking improves exploration and accelerates value propagation in long-horizon reinforcement learning, but naively applying off-policy methods to the temporally extended action space at reduced decision frequency offsets these gains, leading to poor sample efficiency.
By C. F. Maximilian Nagy, Onur Celik, Emiliyan Gospodinov, Florian Seligmann, Weiran Liao, Aryan Kaushik, Gerhard Neumann