arXiv Computer Vision

TexTailor: Texture-Preserving Video Virtual Try-On via Adaptive Garment Conditioning

arXiv Computer Vision
Sep 4

BooM-VVT: Boosting Mask-Free Video Virtual Try-On with Image-Level Pseudo Data

BooM‑VVT is a mask‑free video virtual try‑on framework that builds on a keyframe‑driven paradigm. It introduces a multi‑stage training strategy using image‑level pseudo data to learn mask‑free localization, a garment‑sensitive keyframe sampling method to capture garment appearance, and a Frame‑Shared 3D‑RoPE module to align keyframes with target video frames for accurate garment detail transfer. The authors also release OmniView, a large‑scale multi‑view try‑on dataset, and demonstrate that BooM‑VVT outperforms existing methods in temporal consistency and garment fidelity.

By Wei Zhang, Xin Li, Peishu Shi, Jialin Gao, Xuekang Peng, Zhichao Lian, Yeying Jin
Hugging Face Trending Papers
Sep 3

BooM-VVT: Boosting Mask-Free Video Virtual Try-On with Image-Level Pseudo Data

BooM-VVT is a mask‑free video virtual try‑on framework that builds on a keyframe‑driven paradigm. It uses a multi‑stage training strategy with image‑level pseudo data to learn mask‑free localization, introduces Garment‑Sensitive Keyframe Sampling to capture garment appearance, and employs Frame‑Shared 3D‑RoPE for spatiotemporal correspondence. The authors also create the OmniView dataset to support diverse camera viewpoints and tasks, achieving superior temporal consistency and garment fidelity compared to existing methods.

arXiv AI
Sep 3

ContextAnyone: Context-Aware Diffusion for Character-Consistent Text-to-Video Generation

ContextAnyone is a context‑aware diffusion framework that treats a reference image as an explicitly preserved appearance anchor rather than a simple conditioning signal. By jointly reconstructing the reference image and generating the target video within a shared diffusion transformer, it provides direct supervision for maintaining identity and fine‑grained appearance throughout denoising. The method introduces asymmetric information flow and Gap‑RoPE positional representations to keep the reference stable while allowing selective access by video tokens, and demonstrates improved identity and appearance consistency on an OpenVid‑HD benchmark.

By Ziyang Mai, Yu-Wing Tai
Hugging Face Trending Papers
Sep 24

OmniFabric: Coherent UV Space Texture Synthesis for 3D Garment Reconstruction

OmniFabric is a new method for creating high‑quality, globally coherent texture maps for 3D garment reconstruction from a single image. It first generates a coarse texture initialization on the garment’s sewing pattern using a 3D mesh and Vision‑Language Model priors, then refines this in the UV domain with a diffusion transformer conditioned on 3D positional features. The approach removes distortion and baked‑in artifacts, producing photorealistic 3D garments that outperform existing baselines.

Hugging Face Trending Papers
Aug 27

LiveVVT: High-Fidelity Video Virtual Try-On in Real Time

LiveVVT introduces a rolling streaming diffusion framework for video virtual try‑on that maintains high visual fidelity while enabling real‑time performance. By confining bidirectional spatio‑temporal modeling to a fixed‑size window and using bounded temporal and global appearance memories, it emits clean video chunks with low latency. A progressive distillation pipeline further refines the model, achieving superior quality with 26× lower latency and 11× higher throughput compared to prior methods.

arXiv Computer Vision
Sep 25

OmniFabric: Coherent UV Space Texture Synthesis for 3D Garment Reconstruction

OmniFabric is a new method for creating production‑ready 3D garment assets from a single image. It generates globally coherent texture maps directly in the 2D sewing pattern (UV) space, using a coarse initialization from Vision‑Language Models and refining it with a diffusion transformer conditioned on 3D positional features. The approach removes distortion and baked‑in artifacts, producing photorealistic 3D garments with high‑quality textures that outperform current state‑of‑the‑art baselines.

By Ding-Jiun Huang, Yuanhao Wang, Cheng Zhang, Hugo Bertiche, Alexandru-Eugen Ichim, Thabo Beeler, Fernando De la Torre
arXiv AI
Jun 11

OpenVTON-Bench: A Large-Scale High-Resolution Benchmark for Controllable Virtual Try-On Evaluation

arXiv:2601. 22725v4 Announce Type: replace-cross Abstract: Recent advances in diffusion models have significantly elevated the visual fidelity of Virtual Try-On (VTON) systems, yet reliable evaluation remains a persistent bottleneck.

By Jin Li, Tao Chen, Kai Wen, Siqi Yin, Shuai Jiang, Weijie Wang, Jingwen Luo, Chenhui Wu
arXiv Computer Vision
Sep 3

MMTryon: Multi-Modal Multi-Reference Control for High-Quality Fashion Generation

MMTryon is a multi‑modal, multi‑reference virtual try‑on framework that generates high‑quality compositional try‑on results using text instructions and multiple garment images. It addresses three overlooked problems: supporting multiple try‑on items, allowing dressing style specification via text, and eliminating reliance on segmentation models by using a parsing‑free garment encoder and a scalable data generation pipeline. Experiments on high‑resolution benchmarks and in‑the‑wild test sets show MMTryon outperforms state‑of‑the‑art methods qualitatively and quantitatively.

By Xujie Zhang, Ente Lin, Michael Kampffmeyer, Zhenyu Xie, Jiang Li, Ting Liu, Xiaochao Qu, Luoqi Liu, Xiaodan Liang