arXiv AI

Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation

arXiv:2607. 14962v1 Announce Type: cross Abstract: Text-to-image (T2I) models can synthesize realistic, prompt-aligned images, yet samples generated for the same prompt often cover only a small subset of visually distinct modes.

arXiv Machine Learning
1d ago

Debias Anything: Fairness with Diversity without Supervision in Diffusion Models

The paper introduces a method called Debias Anything that jointly addresses fairness and diversity in diffusion models without requiring sensitive-attribute annotations. By connecting a frozen diffusion model to a pretrained vision-language embedding space via an adapter, the approach uses pairs of text prompts to guide batch composition toward desired attribute proportions and employs a disagreement score to promote diversity. The method is applicable to both unconditional and text-conditional diffusion models and demonstrates improved quality and diversity while maintaining comparable fairness levels in experiments.

By Th\'eau d'Audiffret, Mariia Vladimirova, Jean-Yves Franceschi
arXiv Computer Vision
Sep 7

Learning to Credit the Right Steps: Objective-aware Process Optimization for Visual Generation

The paper introduces Objective-aware Trajectory Credit Assignment (OTCA), a framework that refines reinforcement learning for diffusion-based visual generation. OTCA decomposes credit across denoising steps and allocates multiple reward signals adaptively, addressing the coarse, uniform reward assignment of existing GRPO pipelines. Experiments demonstrate that OTCA consistently enhances image and video generation quality across various metrics.

By Rui Li, Ke Hao, Yuanzhi Liang, Haibin Huang, Chi Zhang, Yun Gu, Xuelong Li
arXiv Machine Learning
Aug 20

To Go Far, Go Together: Diverse Preferences Induce a Curriculum for Reward Optimization

The paper introduces CurriPO, a tree‑structured curriculum that automatically adapts to diverse user reward models in AI alignment tasks. By exploiting the natural hierarchy between easy‑ and hard‑to‑optimize reward models, CurriPO covers a broad user population in a single traversal, reusing previously incorporated reward models. Experiments on personalized continuous control show that CurriPO improves population satisfaction by 1.2–2.1× over the strongest baseline while cutting training time and better serving users traditionally underserved by conventional optimization.

By Taehyung Kim, Jongeun Choi
arXiv Computer Vision
Sep 7

Compositional Reward Models for Conditional Medical Image Generation

The paper introduces PRISM, a Compositional Reward Model framework that decomposes image quality into multiple verifier‑grounded stages for conditional medical image generation. By assigning distinct rewards for fine‑to‑coarse properties—such as intensity, texture, structural alignment, and semantic fidelity—and combining them via a Hierarchical Constrained Propagation mechanism, PRISM addresses shortcomings of single‑scalar reward approaches. Experiments on PanNuke, CeDeM, and ISIC datasets show that data generated with PRISM improves downstream model performance, achieving higher mDice, lower MRE, and increased F1 scores compared to baseline methods.

By Aayush Kumar Tyagi, Prathosh A. P., Mausam
arXiv Machine Learning
Jun 29

Qwen-Image-2.0-RL Technical Report

arXiv:2606. 27608v1 Announce Type: cross Abstract: We present Qwen-Image-2.

By Yixian Xu, Kaiyuan Gao, Yuxiang Chen, Yilei Chen, Zecheng Tang, Zihao Liu, Zikai Zhou, Deqing Li, Hao Meng, Kuan Cao, Jiahao Li, Jie Zhang, Liang Peng, Lihan Jiang, Ningyuan Tang, Shengming Yin, Tianhe Wu, Xiaoyue Chen, Yan Shu, Yanran Zhang, Yi Wang, Yu Wu, Yujia Wu, Zekai Zhang, Zhendong Wang, Xiao Xu, Kun Yan, Chenfei Wu
arXiv AI
Aug 6

Beyond the Dirac Delta: Mitigating Diversity Collapse in Reinforcement Fine-Tuning for Versatile Image Generation

arXiv:2601. 12401v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has emerged as a powerful paradigm for fine-tuning large-scale generative models, such as diffusion and flow models, to align with complex human preferences and user-specified tasks.

By Jinmei Liu, Haoru Li, Zhenhong Sun, Chaofeng Chen, Yatao Bian, Bo Wang, Daoyi Dong, Chunlin Chen, Zhi Wang
Hugging Face Trending Papers
Aug 19

To Go Far, Go Together: Diverse Preferences Induce a Curriculum for Reward Optimization

The paper introduces CurriPO, a tree‑structured curriculum that adapts to diverse user reward models in AI alignment. By automatically building a curriculum that branches and reuses reward models, it addresses the problem of users whose reward models are hard to optimize, a group often underserved by conventional methods. Experiments on personalized continuous control demonstrate that CurriPO improves population satisfaction by 1.2–2.1× over the best baseline while cutting training time.