arXiv AI

ZIPP:Zero-shot Image Personalization from Personas

arXiv:2606. 08841v1 Announce Type: new Abstract: Text-to-image diffusion models are increasingly deployed in open-ended creative contexts, yet their outputs remain impersonal, optimized for aggregate aesthetics rather than individual taste.

arXiv AI
2d ago

Personalized Image Generation with Reasoning and Reflection

arXiv:2610.00737v1 Announce Type: cross Abstract: Personalized image generation has remained narrowly focused on conditional synthesis from curated visual exemplars, rather than capturing who a user...

By Bo Ni, Ngoc N. Tran, Qinwen Ge, Franck Dernoncourt, Seunghyun Yoon, Samyadeep Basu, Sungchul Kim, Puneet Mathur, Nedim Lipka, Tong Yu, Yu Wang, Ryan A. Rossi, Tyler Derr
arXiv Machine Learning
Jun 2

SurrogateSHAP: Training-Free Contributor Attribution for Text-to-Image (T2I) Models

arXiv:2601. 22276v2 Announce Type: replace Abstract: As Text-to-Image (T2I) diffusion models are increasingly used in real-world creative workflows, a principled framework for valuing contributors who provide a collection of data is essential for fair compensation and sustainable data marketplaces.

By Mingyu Lu, Soham Gadgil, Chris Lin, Chanwoo Kim, Su-In Lee
arXiv Computer Vision
Aug 24

CogCanvas: A Benchmark for Evaluating Multi-Subject Reference-Based Image Generation

CogCanvas is a new benchmark for multi-subject reference-based image generation, featuring 1,952 curated reference images of 100 celebrities, 115 objects/fashion items, and 29 real-world backgrounds. It generates 1,361 compositional prompts with 2–5 people, using a pipeline that includes DINOv2 deduplication, aesthetic filtering, and automated graph derivation for interaction and positioning. The benchmark evaluates three tasks—reference-based multi-human-object generation, text-to-image compositional generation, and reference retrieval—under a six-axis protocol, and introduces BG‑Sim and Attr‑VQA metrics to assess background fidelity and attribute binding.

By Long-Bao Nguyen, Quang-Khai Le, Tam V. Nguyen, Minh-Triet Tran, Trung-Nghia Le
Hugging Face Trending Papers
Aug 3

MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing

Recent advances in unified multimodal models have significantly improved text-guided image editing abilities. In particular, models such as Nano-Banana-Pro and GPT-Image-2 demonstrate emerging capabilities in multi-source image editing (MIE), including tasks such as object synthesis, person-background composition, and cross-image style fusion.

arXiv Computer Vision
Sep 7

SEAL: Semantic-aware Single-image Sticker Personalization with a Large-scale Sticker-tag Dataset

The paper introduces SEAL, a plug‑and‑play module that enhances single‑image sticker personalization in diffusion models by adding a semantic‑guided spatial attention loss, a split‑merge token strategy, and structure‑aware layer restriction. SEAL integrates without altering the U‑Net backbone and addresses overfitting issues such as visual entanglement and structural rigidity. Alongside SEAL, the authors release StickerBench, a large sticker dataset with six attribute tags to enable systematic evaluation of identity preservation and contextual controllability.

By Changhyun Roh, Yonghyun Jeong, Jonghyun Lee, Chanho Eom, Jihyong Oh