arXiv:2602. 04037v3 Announce Type: replace Abstract: Learning domain adaptive policies that can generalize to unseen transition dynamics, remains a fundamental challenge in learning-based control.
By Pengcheng Wang, Qinghang Liu, Haotian Lin, Yiheng Li, Guojian Zhan, Masayoshi Tomizuka, Yixiao Wang
The paper introduces Transfer Learning for Evolving Domains (TrED), a framework that models how data availability changes over time in real-world applications. TrED treats the entire trajectory of model updates as a single learning problem, rather than isolated snapshots, and defines a data availability process, a flexible learning protocol, and an evaluation criterion that scores the whole trajectory. The authors review existing transfer learning methods, noting that most are tailored to specific regimes and do not optimize the full trajectory, and argue that TrED is a well‑posed, unsolved research direction.
By Ricardo Ribeiro Pereira, Jacopo Bono, Hugo Ferreira, Pedro Ribeiro, Pedro Saleiro, Pedro Bizarro, Carlos Soares
arXiv:2605.13054v2 Announce Type: replace-cross
Abstract: Cross-domain offline reinforcement learning learns a target policy from pre-collected source and target datasets with different dynamics. Whe...
By Minung Kim, Jeongmo Kim, Gwanwoo Choi, Seungyul Han
The paper introduces Diffusion-Augmented Markov Decision Processes (DA‑MDPs), a framework that extends Maximum Entropy Reinforcement Learning to diffusion-based policies. DA‑MDPs treat each reverse‑diffusion step as an RL decision, deriving a tractable reverse‑KL bound that decomposes across denoising transitions and yields diffusion‑augmented soft rewards, value functions, and policy objectives. The authors implement this framework with PPO, REPPO, and a maximum‑entropy WPO variant, showing improved continuous‑control performance, higher success rates on manipulation tasks, and memory‑efficient training with action chunking.
By Sebastian Sanokowski, Kaustubh Patil, Majid Khadiv
arXiv:2603. 27044v3 Announce Type: replace-cross Abstract: Deep Reinforcement Learning (DRL) is widely recognized as sample-inefficient, a limitation attributable in part to the high dimensionality and substantial functional redundancy inherent to the policy parameter space.
By Andrea Fraschini, Davide Tenedini, Riccardo Zamboni, Mirco Mutti, Marcello Restelli
arXiv:2606. 24601v1 Announce Type: new Abstract: Multi-agent reinforcement learning (MARL) addresses the problem of training multiple agents that pursue collaborative, competitive, or mixed objectives.
By Anurag Akula, Satheesh K. Perepu, Abhishek Sarkar, Kaushik Dey