arXiv AI By Kevin Zhai, Siva Rajesh Kasa, Soumya Roy, Sumit Negi, Mubarak Shah

Traversing the Satisfaction-Diversity Frontier in Text-to-Image Diffusion

Read the original on arXiv AI →

The paper introduces SatisDive, a training‑free inference method that balances reward and diversity in text‑to‑image diffusion by enforcing a reward floor for each image and a diversity cutoff for the batch. By adjusting the reward floor, the method traces a Pareto frontier between worst‑candidate reward and batch diversity. Experiments on Pick‑a‑Pic show that SatisDive consistently outperforms FK steering, improving worst‑candidate reward by up to 0.70 in some settings and Pareto‑dominating FK steering across overlapping DreamSim ranges.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
3d ago

Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric Rewards

The paper presents a post‑training approach for text‑to‑image models that combines a preference reward, trained on large human preference data, with rubric‑based rewards that assess prompt faithfulness and other desirable traits. The authors show that a simple reward composition strategy outperforms a naive weighted average, leading to significant Elo gains on the Arena leaderboard for models like Flux2dev and Ideogram‑4. They also release Arena‑T2I‑Training, a 1K subset of data to aid reproducible research in post‑training.

By Yuanhao Ban, I-Hung Hsu, Anastasios Angelopoulos, Wei-Lin Chiang, Ion Stoica, Cho-Jui Hsieh
arXiv Machine Learning
Jun 18

The Reward Was in Your Data All Along: Correcting Flow Matching with Discriminator-Guided RL

arXiv:2606. 19162v1 Announce Type: new Abstract: Score- and flow-matching models often rely on preference-based reinforcement learning for two purposes: aligning with subjective preferences and, surprisingly, recovering properties such as visual realism and coherent object structure that matching-based training is intended to learn from the data itself.

By Nicolas Beltran-Velez, Felix Friedrich, Zhang Xiaofeng, Reyhane Askari-Hemmat, Xiaochuang Han, Adriana Romero-Soriano, Michal Drozdzal
arXiv Machine Learning
Sep 7

Beyond Pairwise Preferences: Listwise Reward-Aware Alignment for Diffusion Models

The paper introduces Diffusion LAIR, a listwise preference optimization technique that leverages continuous reward scores instead of binary pairwise comparisons to align text‑to‑image diffusion models. LAIR transforms reward scores into centered advantage weights and optimizes an advantage‑weighted regression objective on an implicit reward defined by denoising‑loss improvement over a reference model, with a quadratic penalty to regulate reward magnitude. Experiments demonstrate that Diffusion LAIR surpasses strong baseline methods on SD1.5 and SDXL across generation, compositional, and editing tasks.

By Austin Wang, Jiaqi Han, Stefano Ermon, Yisong Yue