arXiv AI

Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric Rewards

The paper presents a post‑training approach for text‑to‑image models that combines a preference reward, trained on large human preference data, with rubric‑based rewards that assess prompt faithfulness and other desirable traits. The authors show that a simple reward composition strategy outperforms a naive weighted average, leading to significant Elo gains on the Arena leaderboard for models like Flux2dev and Ideogram‑4. They also release Arena‑T2I‑Training, a 1K subset of data to aid reproducible research in post‑training.

arXiv Computer Vision
Aug 28

RubricRM: Generative Reward Modeling via Dynamic Rubrics for Image Generation and Editing

RubricRM introduces a pairwise generative reward modeling framework that generates an input‑specific rubric—comprising evaluation dimensions, weights, and scoring criteria—to score candidate images. The method is trained in two stages: supervised fine‑tuning to learn the rubric‑based scoring paradigm and GRPO to refine dimension‑level rewards. Experiments on text‑to‑image generation and instruction‑based image editing benchmarks demonstrate that RubricRM outperforms existing specialized reward models and competes with strong proprietary MLLM judges while using smaller backbones.

By Zijian Kan, Wei Wang, Long Luo, Bing Zhao, Xuan Ren, Weixu Qiao, Wenbo Li, Hu Wei, Lin Qu
arXiv Computer Vision
Sep 30

Think Before You Score: Thinking Reward Model for Visual Generation

arXiv:2609.37372v1 Announce Type: new Abstract: Visual reward models are essential for evaluating and improving visual generation models, yet existing approaches typically map task conditions and can...

By Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
arXiv Machine Learning
Sep 7

Beyond Pairwise Preferences: Listwise Reward-Aware Alignment for Diffusion Models

The paper introduces Diffusion LAIR, a listwise preference optimization technique that leverages continuous reward scores instead of binary pairwise comparisons to align text‑to‑image diffusion models. LAIR transforms reward scores into centered advantage weights and optimizes an advantage‑weighted regression objective on an implicit reward defined by denoising‑loss improvement over a reference model, with a quadratic penalty to regulate reward magnitude. Experiments demonstrate that Diffusion LAIR surpasses strong baseline methods on SD1.5 and SDXL across generation, compositional, and editing tasks.

By Austin Wang, Jiaqi Han, Stefano Ermon, Yisong Yue
arXiv AI
3d ago

Traversing the Satisfaction-Diversity Frontier in Text-to-Image Diffusion

The paper introduces SatisDive, a training‑free inference method that balances reward and diversity in text‑to‑image diffusion by enforcing a reward floor for each image and a diversity cutoff for the batch. By adjusting the reward floor, the method traces a Pareto frontier between worst‑candidate reward and batch diversity. Experiments on Pick‑a‑Pic show that SatisDive consistently outperforms FK steering, improving worst‑candidate reward by up to 0.70 in some settings and Pareto‑dominating FK steering across overlapping DreamSim ranges.

By Kevin Zhai, Siva Rajesh Kasa, Soumya Roy, Sumit Negi, Mubarak Shah