The task of synthesizing stylistically coherent fashion outfits from massive item libraries, known as fashion outfit generation, remains a non-trivial challenge, primarily due to the non-monotonic and implicit nature of aesthetic compatibility, coupled with the exponentially large combinatorial search space. In this paper, we formalize this task as Constrained Ensemble Generation (CEG) and model it as a finite-horizon deterministic Markov Decision Process.
arXiv:2608.13888v2 Announce Type: replace
Abstract: Fashion Outfit Composition (FOC) requires sequentially assembling fashion items into a stylistically cohesive ensemble. Existing works struggle to...
By Kaicheng Pang, Xingxing Zou, Ruohan Xu, Waikeung Wong
Fashion complementary image generation (CIG) aims to create garments that stylistically match a seed item based on user intent, making it a natural multimodal grounding problem where models must inter...
arXiv:2608.23302v1 Announce Type: new
Abstract: Fashion complementary image generation (CIG) aims to create garments that stylistically match a seed item based on user intent, making it a natural mul...
By Matteo Attimonelli, Claudio Pomo, Alessandro De Bellis, Danilo Danese, Dietmar Jannach, Tommaso Di Noia
MMTryon is a multi‑modal, multi‑reference virtual try‑on framework that generates high‑quality compositional try‑on results using text instructions and multiple garment images. It addresses three overlooked problems: supporting multiple try‑on items, allowing dressing style specification via text, and eliminating reliance on segmentation models by using a parsing‑free garment encoder and a scalable data generation pipeline. Experiments on high‑resolution benchmarks and in‑the‑wild test sets show MMTryon outperforms state‑of‑the‑art methods qualitatively and quantitatively.
By Xujie Zhang, Ente Lin, Michael Kampffmeyer, Zhenyu Xie, Jiang Li, Ting Liu, Xiaochao Qu, Luoqi Liu, Xiaodan Liang
CogCanvas is a new benchmark for multi-subject reference-based image generation, featuring 1,952 curated reference images of 100 celebrities, 115 objects/fashion items, and 29 real-world backgrounds. It generates 1,361 compositional prompts with 2–5 people, using a pipeline that includes DINOv2 deduplication, aesthetic filtering, and automated graph derivation for interaction and positioning. The benchmark evaluates three tasks—reference-based multi-human-object generation, text-to-image compositional generation, and reference retrieval—under a six-axis protocol, and introduces BG‑Sim and Attr‑VQA metrics to assess background fidelity and attribute binding.
By Long-Bao Nguyen, Quang-Khai Le, Tam V. Nguyen, Minh-Triet Tran, Trung-Nghia Le
arXiv:2609.14100v1 Announce Type: cross
Abstract: Fashion Image Captioning (FIC) plays a vital role in enhancing user experience and product search in e-commerce platforms. Unlike natural scene image...
By Abhirama Subramanyam Penamakuri, Shreya Shukla, Anand Mishra
arXiv:2607. 23189v1 Announce Type: cross Abstract: AI-generated content (AIGC) has made significant progress, with 2D generative models becoming ready-to-use tools for the digital fashion industry.
By Shenghao Yang, Hongtao Zhang, Yuhan Yi, Zhihao Tang, Zihao Cui, Lian Wen, Han Yan, Yuan Gao, Mingbo Zhao
arXiv:2608.29804v1 Announce Type: new
Abstract: Virtual try-on (VTON) requires not only realistic generation but also faithful preservation of garment characteristics. However, existing evaluation me...
By Kaidong Zhang, Yukang Ding, Xiaoyu Liu, Ying Chen
arXiv:2601. 22276v2 Announce Type: replace Abstract: As Text-to-Image (T2I) diffusion models are increasingly used in real-world creative workflows, a principled framework for valuing contributors who provide a collection of data is essential for fair compensation and sustainable data marketplaces.
By Mingyu Lu, Soham Gadgil, Chris Lin, Chanwoo Kim, Su-In Lee
arXiv:2609.13259v1 Announce Type: cross
Abstract: Virtual Try-On (VTON) aims to dress a person with the reference garment, producing visually reasonable results aligned with human preferences. Turnin...
By Xueheng Li, Yong Liu, Xiaolong Fu, Wen Xue, Chengjun Xie, Yipeng Sun, Yan Li, Simiu Gu
arXiv:2606. 08841v1 Announce Type: new Abstract: Text-to-image diffusion models are increasingly deployed in open-ended creative contexts, yet their outputs remain impersonal, optimized for aggregate aesthetics rather than individual taste.
By Harini SI, Somesh Singh, Yaman Kumar Singla, David Doermann, Rajiv Ratn Shah