arXiv AI By Potito Aghilar, Vito Walter Anelli, Michelantonio Trizio, Eugenio Di Sciascio, Tommaso Di Noia

Training-Free, Identity-Preserving Image Editing for Fashion Pose Alignment and Normalization

Read the original on arXiv AI →

arXiv:2501. 13692v2 Announce Type: replace-cross Abstract: Diffusion models have recently unlocked new possibilities in editing images of real-world objects.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 3

MMTryon: Multi-Modal Multi-Reference Control for High-Quality Fashion Generation

MMTryon is a multi‑modal, multi‑reference virtual try‑on framework that generates high‑quality compositional try‑on results using text instructions and multiple garment images. It addresses three overlooked problems: supporting multiple try‑on items, allowing dressing style specification via text, and eliminating reliance on segmentation models by using a parsing‑free garment encoder and a scalable data generation pipeline. Experiments on high‑resolution benchmarks and in‑the‑wild test sets show MMTryon outperforms state‑of‑the‑art methods qualitatively and quantitatively.

By Xujie Zhang, Ente Lin, Michael Kampffmeyer, Zhenyu Xie, Jiang Li, Ting Liu, Xiaochao Qu, Luoqi Liu, Xiaodan Liang