Creating photorealistic and temporally coherent animatable human avatars from RGB videos remains challenging. Current methods struggle to capture realistic cloth dynamics, producing over-smoothed appearance or severe artifacts on out-of-distribution poses.
arXiv:2511. 18765v3 Announce Type: replace-cross Abstract: Existing industrial 3D garment meshes already cover most real-world clothing geometries, yet their texture diversity remains limited.
By Hui Shan, Ming Li, Haitao Yang, Kai Zheng, Sizhe Zheng, Yanwei Fu, Xiangru Huang
BooM‑VVT is a mask‑free video virtual try‑on framework that builds on a keyframe‑driven paradigm. It introduces a multi‑stage training strategy using image‑level pseudo data to learn mask‑free localization, a garment‑sensitive keyframe sampling method to capture garment appearance, and a Frame‑Shared 3D‑RoPE module to align keyframes with target video frames for accurate garment detail transfer. The authors also release OmniView, a large‑scale multi‑view try‑on dataset, and demonstrate that BooM‑VVT outperforms existing methods in temporal consistency and garment fidelity.
By Wei Zhang, Xin Li, Peishu Shi, Jialin Gao, Xuekang Peng, Zhichao Lian, Yeying Jin
BooM-VVT is a mask‑free video virtual try‑on framework that builds on a keyframe‑driven paradigm. It uses a multi‑stage training strategy with image‑level pseudo data to learn mask‑free localization, introduces Garment‑Sensitive Keyframe Sampling to capture garment appearance, and employs Frame‑Shared 3D‑RoPE for spatiotemporal correspondence. The authors also create the OmniView dataset to support diverse camera viewpoints and tasks, achieving superior temporal consistency and garment fidelity compared to existing methods.
The paper introduces HyperBones, a real‑time garment simulation framework that combines a reduced‑space neural dynamics simulator with a lightweight neural network correcting Linear Blend Skinning (LBS) at a coarse level, and a convolutional MLP for fine‑scale wrinkle recovery in UV space. By decoupling identity‑specific computation from shape conditioning through a hypernetwork, the method achieves high performance without an offline simulator, delivering physically plausible dynamics across diverse motions and unseen body shapes. Experiments demonstrate a speedup of over 30× compared to state‑of‑the‑art autoregressive neural simulators, reaching interactive inference at roughly 1 ms per frame on a consumer GPU.
By Astitva Srivastava, Hsiao-Yu Chen, Ryan Goldade, Philipp Herholz, Zhongshi Jiang, Gene Wei-Chin Lin, Lingchen Yang, Nikolaos Sarafianos, Tuur Stuyck, Avinash Sharma, Egor Larionov
arXiv:2608. 19900v1 Announce Type: new Abstract: For full-body avatars, modeling surface dynamics is crucial for overcoming the uncanny valley and achieving perceptual realism.
By Guoxing Sun, Heming Zhu, Linjie Lyu, Pascal Fua, Christian Theobalt, Marc Habermann
The paper introduces AvaImg, a multi‑stage optimization pipeline that achieves high‑fidelity SMPL(-X)+D registrations with UV texture for arbitrary clothed scans. By enforcing a body‑inside‑clothing constraint through signed winding numbers and employing a three‑level efficiency cascade, AvaImg significantly reduces runtime and storage while recovering fine surface detail via coarse‑to‑fine displacement optimization. The resulting textured registrations are nearly indistinguishable from scans, and encoding the UV maps with a frozen FLUX VAE demonstrates compatibility with 2D generative models, enabling 3D avatar generation using image‑based priors.
By Margaret Kostyrko, Yuxuan Xue, Garvita Tiwari, Gerard Pons-Moll
arXiv:2607. 23189v1 Announce Type: cross Abstract: AI-generated content (AIGC) has made significant progress, with 2D generative models becoming ready-to-use tools for the digital fashion industry.
By Shenghao Yang, Hongtao Zhang, Yuhan Yi, Zhihao Tang, Zihao Cui, Lian Wen, Han Yan, Yuan Gao, Mingbo Zhao
For full-body avatars, modeling surface dynamics is crucial for overcoming the uncanny valley and achieving perceptual realism. Person-agnostic methods recover static 3D avatars from monocular images, videos, or text prompts, but their skeleton-driven animations lack realistic surface dynamics such as clothing wrinkles.
arXiv:2501. 13692v2 Announce Type: replace-cross Abstract: Diffusion models have recently unlocked new possibilities in editing images of real-world objects.
By Potito Aghilar, Vito Walter Anelli, Michelantonio Trizio, Eugenio Di Sciascio, Tommaso Di Noia
arXiv:2609.01276v1 Announce Type: new
Abstract: Complete 3D perception from egocentric video requires recovering the surrounding scene and the wearer's full-body motion in a shared metric frame. Exis...
By Kai Guan, Minchao Jiang, Ruichen WangLi, Wentao Zhu, Lei Zhang
arXiv:2606. 02000v1 Announce Type: cross Abstract: Diffusion models have shown remarkable success in video generation.
By Jingyun Liang, Min Wei, Shikai Li, Yizeng Han, Hangjie Yuan, Lei Sun, Weihua Chen, Fan Wang