arXiv Machine Learning

Unsupervised Diffusion Solver for Combinatorial Optimization via Combinatorial Adjoint Matching

arXiv:2605. 30920v2 Announce Type: replace Abstract: Diffusion-based neural solvers have shown strong promise for combinatorial optimization (CO), but existing methods typically rely on supervised training with large collections of near-optimal solutions.

arXiv Machine Learning
Jul 17

A Continuous-Time Reinforcement Learning Framework for Fine-Tuning Discrete Diffusion Models

arXiv:2607. 14522v1 Announce Type: new Abstract: We formulate reinforcement learning (RL) in continuous time with discrete state spaces and possibly arbitrary action spaces via a stochastic control approach, where the state dynamics are modeled as a controlled continuous-time Markov chain (CTMC).

By Zikun Zhang, Jiayuan Sheng, David D. Yao, Wenpin Tang
arXiv Machine Learning
Aug 27

Bayesian Flow Networks for Offline Trajectory Planning

The paper introduces Bayesian Flow Networks for Offline Trajectory Planning (BFN-RL), a generative modeling framework that unifies discrete and continuous trajectory synthesis for offline reinforcement learning. Unlike prior diffusion models that rely on Gaussian noise, BFN-RL iteratively updates distribution parameters, enabling a categorical planner to produce future state sequences and an inverse-dynamics model to translate these states into actions. Experiments demonstrate that BFN-RL effectively generates trajectories in both discrete planning and continuous control tasks, highlighting its versatility across data modalities.

By Ludvig Killingberg, Helge Langseth
arXiv AI
4d ago

Grab a Coffee: Future-Aware Guidance for Discrete Diffusion with Compiled Objectives

COFFEE is a plug‑and‑play framework that enables future‑aware guidance for discrete diffusion models by separating sequence dependence from the objective. It uses a target‑free carrier to absorb marginal token distributions and a compiled finite‑state model to capture how token combinations affect sequence‑level preferences, allowing global preferences to be transferred to unresolved positions without retraining the diffusion model. The framework supports both hard constraints and learned soft objectives and demonstrates strong control results across symbolic, language, and biological benchmarks.

By Hua (Edward), Xu, Dongxin Li, Gwen Yidou-Weng, Guy Van den Broeck, Wei Wang, Anji Liu
arXiv AI
3d ago

Diffusion-Augmented Markov Decision Processes for Maximum Entropy Reinforcement Learning

The paper introduces Diffusion-Augmented Markov Decision Processes (DA‑MDPs), a framework that extends Maximum Entropy Reinforcement Learning to diffusion-based policies. DA‑MDPs treat each reverse‑diffusion step as an RL decision, deriving a tractable reverse‑KL bound that decomposes across denoising transitions and yields diffusion‑augmented soft rewards, value functions, and policy objectives. The authors implement this framework with PPO, REPPO, and a maximum‑entropy WPO variant, showing improved continuous‑control performance, higher success rates on manipulation tasks, and memory‑efficient training with action chunking.

By Sebastian Sanokowski, Kaustubh Patil, Majid Khadiv