arXiv AI By Yuanhao Ban, I-Hung Hsu, Anastasios Angelopoulos, Wei-Lin Chiang, Ion Stoica, Cho-Jui Hsieh

Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric Rewards

Read the original on arXiv AI →

The paper presents a post‑training approach for text‑to‑image models that combines a preference reward, trained on large human preference data, with rubric‑based rewards that assess prompt faithfulness and other desirable traits. The authors show that a simple reward composition strategy outperforms a naive weighted average, leading to significant Elo gains on the Arena leaderboard for models like Flux2dev and Ideogram‑4. They also release Arena‑T2I‑Training, a 1K subset of data to aid reproducible research in post‑training.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Aug 28

RubricRM: Generative Reward Modeling via Dynamic Rubrics for Image Generation and Editing

RubricRM introduces a pairwise generative reward modeling framework that generates an input‑specific rubric—comprising evaluation dimensions, weights, and scoring criteria—to score candidate images. The method is trained in two stages: supervised fine‑tuning to learn the rubric‑based scoring paradigm and GRPO to refine dimension‑level rewards. Experiments on text‑to‑image generation and instruction‑based image editing benchmarks demonstrate that RubricRM outperforms existing specialized reward models and competes with strong proprietary MLLM judges while using smaller backbones.

By Zijian Kan, Wei Wang, Long Luo, Bing Zhao, Xuan Ren, Weixu Qiao, Wenbo Li, Hu Wei, Lin Qu