arXiv:2607. 02137v1 Announce Type: cross Abstract: We study timestep allocation for score-based diffusion sampling, where a learned reverse-time dynamics is discretized on a finite grid.
By Yilie Huang, Wenpin Tang, Xun Yu Zhou
We study timestep allocation for score-based diffusion sampling, where a learned reverse-time dynamics is discretized on a finite grid. Uniform and hand-crafted schedules are standard choices, but they rely on fixed prescriptions and can therefore be suboptimal.
arXiv:2607.02137v3 Announce Type: replace-cross
Abstract: We study timestep allocation for score-based diffusion sampling, where a learned reverse-time dynamics is discretized on a finite grid. Unifo...
By Yilie Huang, Wenpin Tang, Xun Yu Zhou
arXiv:2602. 08689v2 Announce Type: replace Abstract: Diffusion models generate samples through an iterative denoising process guided by a pretrained neural network.
By Constant Bourdrez, Alexandre V\'erine, Olivier Capp\'e
arXiv:2606. 15048v1 Announce Type: new Abstract: Diffusion models are typically trained with objectives that focus on local denoising targets at individual time steps (or adjacent pairs), which do not enforce consistency between predictions along the denoising trajectory.
By Qizhen Ying, Yangchen Pan, Victor Adrian Prisacariu, Junfeng Wen
arXiv:2607. 07693v1 Announce Type: cross Abstract: Reinforcement learning from human feedback (RLHF) has emerged as a powerful paradigm for aligning generative models with human preferences.
By Eric Zhu, Abhinav Shrivastava, Soumik Mukhopadhyay
arXiv:2602.12624v2 Announce Type: replace
Abstract: Diffusion-based generative models have achieved remarkable performance across various domains, yet their practical deployment is often limited by h...
By Sangwoo Jo, Sungjoon Choi
arXiv:2607. 17326v1 Announce Type: new Abstract: Transfer-oriented reinforcement learning requires evaluating algorithms along dimensions that go beyond standard sample efficiency.
By Hany Hamed, Abhishek Naik, Colin Bellinger, A. Rupam Mahmood
arXiv:2605. 12236v2 Announce Type: replace-cross Abstract: Fine-tuning pre-trained robot policies with reinforcement learning (RL) often inherits the bottlenecks introduced by pre-training with behavioral cloning (BC), which produces narrow action distributions that lack the coverage necessary for downstream exploration.
By Matthew M. Hong, Jesse Zhang, Anusha Nagabandi, Abhishek Gupta
arXiv:2606.22394v3 Announce Type: replace
Abstract: Consistency distillation has significantly accelerated diffusion-model inference, but its sampling dynamics remain underexplored. We reveal an asym...
By Songtao Tian, Guhan Chen, Bohan Li, Jingyi Ma, Zixiong Yu
The paper introduces Learned End-to-End Guidance Schedules (LEEGS) for diffusion models, which train a time‑dependent guidance schedule to balance data quality and requirement satisfaction while reducing sampling steps. LEEGS minimizes the guidance function over a small set of examples using stochastic gradient descent and employs a gradient approximation to cut training time by a factor of four. Experiments on tasks such as image inpainting, noisy image inverse problems, face‑ID‑guided generation, and PDE problems show that LEEGS outperforms baselines at the same computational budget or matches constant guidance with only 10% of the steps.
By Aneesh Barthakur, Mathias Niepert, Luiz F. O. Chamon
The paper introduces Diffusion-Augmented Markov Decision Processes (DA‑MDPs), a framework that extends Maximum Entropy Reinforcement Learning to diffusion-based policies. DA‑MDPs treat each reverse‑diffusion step as an RL decision, deriving a tractable reverse‑KL bound that decomposes across denoising transitions and yields diffusion‑augmented soft rewards, value functions, and policy objectives. The authors implement this framework with PPO, REPPO, and a maximum‑entropy WPO variant, showing improved continuous‑control performance, higher success rates on manipulation tasks, and memory‑efficient training with action chunking.
By Sebastian Sanokowski, Kaustubh Patil, Majid Khadiv