MMTryon is a multi‑modal, multi‑reference virtual try‑on framework that generates high‑quality compositional try‑on results using text instructions and multiple garment images. It addresses three overlooked problems: supporting multiple try‑on items, allowing dressing style specification via text, and eliminating reliance on segmentation models by using a parsing‑free garment encoder and a scalable data generation pipeline. Experiments on high‑resolution benchmarks and in‑the‑wild test sets show MMTryon outperforms state‑of‑the‑art methods qualitatively and quantitatively.
By Xujie Zhang, Ente Lin, Michael Kampffmeyer, Zhenyu Xie, Jiang Li, Ting Liu, Xiaochao Qu, Luoqi Liu, Xiaodan Liang
FitControler introduces a fit-aware virtual try‑on system that adds garment fit control to existing VTON models. It uses a fit‑aware layout generator and a multi‑scale fit injector to redraw body‑garment layouts and render garments that match those layouts. The authors also release a new Fit4Men dataset of 13,000 body‑garment pairs and two fit consistency metrics to evaluate fit quality.
By Lu Yang, Yicheng Liu, Letian Zhou, Yanan Li, Xiang Bai, Hao Lu
arXiv:2607.11233v2 Announce Type: replace
Abstract: Virtual try-on (VTON) is a bi-conditional image generation problem that requires not only accurate person preservation but also faithful garment de...
By Lu Yang, Xiaonan Hu, Yanan Li, Daqi Liu, Hao Lu, Xiang Bai
BooM‑VVT is a mask‑free video virtual try‑on framework that builds on a keyframe‑driven paradigm. It introduces a multi‑stage training strategy using image‑level pseudo data to learn mask‑free localization, a garment‑sensitive keyframe sampling method to capture garment appearance, and a Frame‑Shared 3D‑RoPE module to align keyframes with target video frames for accurate garment detail transfer. The authors also release OmniView, a large‑scale multi‑view try‑on dataset, and demonstrate that BooM‑VVT outperforms existing methods in temporal consistency and garment fidelity.
By Wei Zhang, Xin Li, Peishu Shi, Jialin Gao, Xuekang Peng, Zhichao Lian, Yeying Jin
BooM-VVT is a mask‑free video virtual try‑on framework that builds on a keyframe‑driven paradigm. It uses a multi‑stage training strategy with image‑level pseudo data to learn mask‑free localization, introduces Garment‑Sensitive Keyframe Sampling to capture garment appearance, and employs Frame‑Shared 3D‑RoPE for spatiotemporal correspondence. The authors also create the OmniView dataset to support diverse camera viewpoints and tasks, achieving superior temporal consistency and garment fidelity compared to existing methods.
arXiv:2608.29804v1 Announce Type: new
Abstract: Virtual try-on (VTON) requires not only realistic generation but also faithful preservation of garment characteristics. However, existing evaluation me...
By Kaidong Zhang, Yukang Ding, Xiaoyu Liu, Ying Chen
arXiv:2602.24043v2 Announce Type: replace
Abstract: Reconstructing 3D clothed humans from monocular images and videos is a fundamental problem with applications in virtual try-on, avatar creation, an...
By Yingxuan You, Ren Li, Corentin Dumery, Cong Cao, Hao Li, Pascal Fua
arXiv:2608.23302v1 Announce Type: new
Abstract: Fashion complementary image generation (CIG) aims to create garments that stylistically match a seed item based on user intent, making it a natural mul...
By Matteo Attimonelli, Claudio Pomo, Alessandro De Bellis, Danilo Danese, Dietmar Jannach, Tommaso Di Noia
arXiv:2601. 22725v4 Announce Type: replace-cross Abstract: Recent advances in diffusion models have significantly elevated the visual fidelity of Virtual Try-On (VTON) systems, yet reliable evaluation remains a persistent bottleneck.
By Jin Li, Tao Chen, Kai Wen, Siqi Yin, Shuai Jiang, Weijie Wang, Jingwen Luo, Chenhui Wu
arXiv:2608. 05745v1 Announce Type: cross Abstract: Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity, motion, and scene dynamics.
By Yushe Cao, Shikun Feng, Fei Shen, Haikuo Peng, Jianqiang Xia, Yiheng Zhu, Dianxi Shi, Chun Yu
arXiv:2609.39335v1 Announce Type: new
Abstract: Video virtual try-on has attracted increasing attention due to its broad potential in digital fashion and intelligent e-commerce. However, existing met...
By Zijing Qin, Jun Zhou, Ruicheng Zhang, Jiaqi Hou, Zunnan Xu, Ronghui Li, Zhenyu Xie, Xiu Li
Fashion complementary image generation (CIG) aims to create garments that stylistically match a seed item based on user intent, making it a natural multimodal grounding problem where models must inter...