arXiv AI

ReNFT: Repairing Mode Collapse in Reward Post-Training via Internal Probability-Mass Recalibration

ReNFT is a method that repairs mode collapse in diffusion generators after reward post‑training by internally recalibrating probability mass. It identifies suppressed alternatives through anti‑hub prompts and uses two policy‑dominated routes to generate counterfactual proposals, then applies reward‑based ranking and a joint‑and‑paired NFT update to restore diversity while preserving reward. Experiments on PickScore and GenEval show that ReNFT retains almost all of the original reward while significantly boosting diversity metrics.

arXiv Machine Learning
Jun 16

GD$^2$PO: Mitigating Multi-Reward Conflicts via Group-Dynamic reward-Decoupled Policy Optimization

arXiv:2606. 16771v1 Announce Type: new Abstract: As LLMs advance, post-training reinforcement learning (RL) increasingly relies on multi-dimensional rewards to cultivate comprehensive capabilities.

By Haotian Liu, Yihao Liu, Jingwei Ni, Siyuan Huang, Xinpeng Liu, Pengyu Cheng, Jiajun Song, Ruijin Ding, Junfeng Li, Zhechao Yu, Mengyu Zhou, Hongteng Xu, Xiaoxi Jiang, Guanjun Jiang
arXiv Machine Learning
Aug 5

Latent Reward Registers for Diffusion Preference Alignment

arXiv:2608. 03929v1 Announce Type: new Abstract: Aligning diffusion models with human preferences usually relies on a sparse terminal reward evaluated on the final generated samples, presenting a severe temporal credit-assignment challenge across the multi-step denoising process.

By Yuanshen Guan, Zipeng Feng, Zhiwei Xiong, Peiqin Sun
arXiv Machine Learning
Sep 25

Uncertainty-Gated Exploration Noise Suppresses Task Collapse in Online RL Fine-Tuning of a Flow-Matching Vision-Language-Action Policy

The paper investigates task collapse—a failure mode where online RL fine‑tuning of a pretrained flow‑matching vision‑language‑action policy erodes performance on individual tasks—using a 450M‑parameter SmolVLA policy on LIBERO‑10. Three exploration‑noise strategies are compared: a fixed noise scale, a learned noise network, and an uncertainty‑gated controller that reallocates exploration based on novelty and competence signals without task labels. The uncertainty‑gated controller prevents task collapse across all tested seeds, whereas the other two approaches consistently cause collapse, demonstrating its effectiveness in preserving task performance during fine‑tuning.

By Mehmet Turan Yard{\i}mc{\i}, Yunus Emre \c{C}o\u{g}urcu
arXiv AI
Sep 4

FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience

arXiv:2609. 03241v1 Announce Type: cross Abstract: A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal verifiers provide reliable yet sparse supervision, while dense same-model guidance can reinforce false confidence or overconcentrate learning on a narrow solution mode.

By Zixun Huang, Kishan Panaganti, Haitao Mi, Leowei Liang
arXiv AI
Jul 7

Regime-Conditional Stabilisation of LLM-Augmented Cooperative Multi-Agent Reinforcement Learning

arXiv:2607. 04470v1 Announce Type: cross Abstract: Large Language Models (LLMs) offer a natural interface for translating human objectives into reward signals for cooperative multi-agent reinforcement learning (MARL), yet the training-time dynamics of this integration remain poorly understood.

By Faid Keddouri, Sohaib Houhou, Aissa Boulmerka, Nadir Farhi