arXiv Machine Learning
Sep 7

Beyond Pairwise Preferences: Listwise Reward-Aware Alignment for Diffusion Models

The paper introduces Diffusion LAIR, a listwise preference optimization technique that leverages continuous reward scores instead of binary pairwise comparisons to align text‑to‑image diffusion models. LAIR transforms reward scores into centered advantage weights and optimizes an advantage‑weighted regression objective on an implicit reward defined by denoising‑loss improvement over a reference model, with a quadratic penalty to regulate reward magnitude. Experiments demonstrate that Diffusion LAIR surpasses strong baseline methods on SD1.5 and SDXL across generation, compositional, and editing tasks.

By Austin Wang, Jiaqi Han, Stefano Ermon, Yisong Yue
arXiv AI
Oct 2

Personalized Image Generation with Reasoning and Reflection

The paper introduces a unified benchmark for personalized image generation that uses users' historical data—such as reviews, posts, images, captions, and metadata—to create images aligned with their lifestyle and aesthetic preferences. It defines two tasks: Personalized Scene Generation, which places objects in scenes reflecting user preferences for product presentation, and Personalized Creative Generation, which produces novel images faithful to a user's aesthetic for social media content. The authors also propose PEARL, a method that interleaves multimodal reasoning with a frozen image generator, achieving a 15% average improvement over baselines on personalization metrics.

By Bo Ni, Ngoc N. Tran, Qinwen Ge, Franck Dernoncourt, Seunghyun Yoon, Samyadeep Basu, Sungchul Kim, Puneet Mathur, Nedim Lipka, Tong Yu, Yu Wang, Ryan A. Rossi, Tyler Derr