Personalized Image Generation with Reasoning and Reflection
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The paper introduces a unified benchmark for personalized image generation that uses users' historical data—such as reviews, posts, images, captions, and metadata—to create images aligned with their lifestyle and aesthetic preferences. It defines two tasks: Personalized Scene Generation, which places objects in scenes reflecting user preferences for product presentation, and Personalized Creative Generation, which produces novel images faithful to a user's aesthetic for social media content. The authors also propose PEARL, a method that interleaves multimodal reasoning with a frozen image generator, achieving a 15% average improvement over baselines on personalization metrics.
arXiv:2606. 08841v1 Announce Type: new Abstract: Text-to-image diffusion models are increasingly deployed in open-ended creative contexts, yet their outputs remain impersonal, optimized for aggregate aesthetics rather than individual taste.
arXiv:2511. 00609v4 Announce Type: replace Abstract: Personalized image preference assessment aims to evaluate an individual user's image preferences by relying only on a small set of reference images as prior information.
arXiv:2610.09015v1 Announce Type: new Abstract: Diffusion models can generate high-quality images, yet aligning their outputs with individual user preferences remains challenging. A key bottleneck is...
PhotoBench is a new benchmark built from authentic personal photo albums that moves beyond simple visual matching to focus on personalized, intent-driven retrieval. It incorporates a multi-source profiling framework that combines visual semantics, spatial‑temporal metadata, social identity, and temporal events to generate complex queries reflecting users’ life trajectories. Evaluation on PhotoBench reveals two key limitations: a modality gap where unified embedding models fail on non‑visual constraints, and a source fusion paradox where agentic systems struggle with tool orchestration.
Recent text-to-image models such as DALLE-3 excel at following diverse prompts yet remain blind to individual aesthetic preferences. We study personalized image generation, where models must align outputs with a user's implicit visual preferences based on a few historically preferred images and a short prompt.