arXiv AI

DADiff: Diffusion-Driven Cross-Domain Policy Adaptation for Reinforcement Learning

arXiv:2607. 16090v1 Announce Type: cross Abstract: Transferring policies across domains poses a vital challenge in reinforcement learning, due to the dynamics mismatch between the source and target domains.

arXiv Machine Learning
Jun 19

DADP: Domain Adaptive Diffusion Policy

arXiv:2602. 04037v3 Announce Type: replace Abstract: Learning domain adaptive policies that can generalize to unseen transition dynamics, remains a fundamental challenge in learning-based control.

By Pengcheng Wang, Qinghang Liu, Haotian Lin, Yiheng Li, Guojian Zhan, Masayoshi Tomizuka, Yixiao Wang
arXiv Machine Learning
Sep 14

Transfer Learning for Evolving Domains

The paper introduces Transfer Learning for Evolving Domains (TrED), a framework that models how data availability changes over time in real-world applications. TrED treats the entire trajectory of model updates as a single learning problem, rather than isolated snapshots, and defines a data availability process, a flexible learning protocol, and an evaluation criterion that scores the whole trajectory. The authors review existing transfer learning methods, noting that most are tailored to specific regimes and do not optimize the full trajectory, and argue that TrED is a well‑posed, unsolved research direction.

By Ricardo Ribeiro Pereira, Jacopo Bono, Hugo Ferreira, Pedro Ribeiro, Pedro Saleiro, Pedro Bizarro, Carlos Soares
arXiv AI
3d ago

Diffusion-Augmented Markov Decision Processes for Maximum Entropy Reinforcement Learning

The paper introduces Diffusion-Augmented Markov Decision Processes (DA‑MDPs), a framework that extends Maximum Entropy Reinforcement Learning to diffusion-based policies. DA‑MDPs treat each reverse‑diffusion step as an RL decision, deriving a tractable reverse‑KL bound that decomposes across denoising transitions and yields diffusion‑augmented soft rewards, value functions, and policy objectives. The authors implement this framework with PPO, REPPO, and a maximum‑entropy WPO variant, showing improved continuous‑control performance, higher success rates on manipulation tasks, and memory‑efficient training with action chunking.

By Sebastian Sanokowski, Kaustubh Patil, Majid Khadiv
arXiv AI
Jul 7

Unsupervised Behavioral Compression: Learning Low-Dimensional Policy Manifolds through State-Occupancy Matching

arXiv:2603. 27044v3 Announce Type: replace-cross Abstract: Deep Reinforcement Learning (DRL) is widely recognized as sample-inefficient, a limitation attributable in part to the high dimensionality and substantial functional redundancy inherent to the policy parameter space.

By Andrea Fraschini, Davide Tenedini, Riccardo Zamboni, Mirco Mutti, Marcello Restelli
arXiv AI
Sep 25

RD-JEPA: Predictive latent pretraining for few-trajectory transfer across reaction--diffusion equations

RD‑JEPA is a joint‑embedding predictive architecture designed for self‑supervised pretraining on reaction‑diffusion trajectories. The model is pretrained on five parameterized systems and then adapted to three held‑out systems that were not seen during pretraining. Using as few as one, five, or ten complete trajectories from a held‑out system, RD‑JEPA outperforms five supervised surrogate baselines, an independently trained control that removes the trajectory‑dependent predictive latent pathway, and an architecture‑matched model trained from scratch, achieving lower mean relative discrete β field error and mean absolute spatial first‑difference error across various output resolutions, forecast horizons, and adaptation trajectory choices.

By Chenhao Si, Ming Yan
arXiv AI
Jun 10

MMD Guidance: Training-Free Distribution Adaptation for Diffusion Models via Maximum Mean Discrepancy Guidance

arXiv:2601. 08379v2 Announce Type: replace-cross Abstract: Pre-trained diffusion models have emerged as powerful generative priors for both unconditional and conditional sample generation, yet their outputs often deviate from the characteristics of user-specific target data.

By Matina Mahdizadeh Sani, Nima Jamali, Mohammad Jalali, Farzan Farnia