arXiv Computer Vision By Xujie Zhang, Ente Lin, Michael Kampffmeyer, Zhenyu Xie, Jiang Li, Ting Liu, Xiaochao Qu, Luoqi Liu, Xiaodan Liang

MMTryon: Multi-Modal Multi-Reference Control for High-Quality Fashion Generation

Read the original on arXiv Computer Vision →

MMTryon is a multi‑modal, multi‑reference virtual try‑on framework that generates high‑quality compositional try‑on results using text instructions and multiple garment images. It addresses three overlooked problems: supporting multiple try‑on items, allowing dressing style specification via text, and eliminating reliance on segmentation models by using a parsing‑free garment encoder and a scalable data generation pipeline. Experiments on high‑resolution benchmarks and in‑the‑wild test sets show MMTryon outperforms state‑of‑the‑art methods qualitatively and quantitatively.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Sep 1

GarmentWeaver: Schema-Aware Structured Synthesis for Multimodal Sewing Patterns

GarmentWeaver is a new framework for multimodal sewing pattern generation that uses a schema‑aware approach to construct compact hierarchical targets. By activating garment‑relevant structural branches and building on a pretrained vision‑language model, it predicts executable sewing patterns in a structured manner. Experiments show that GarmentWeaver produces more accurate, executable patterns and yields better simulation results than strong baselines.

By Yinwen Lu, Weihao Luo, Yueqi Zhong