BooM-VVT is a mask‑free video virtual try‑on framework that builds on a keyframe‑driven paradigm. It uses a multi‑stage training strategy with image‑level pseudo data to learn mask‑free localization, introduces Garment‑Sensitive Keyframe Sampling to capture garment appearance, and employs Frame‑Shared 3D‑RoPE for spatiotemporal correspondence. The authors also create the OmniView dataset to support diverse camera viewpoints and tasks, achieving superior temporal consistency and garment fidelity compared to existing methods.
arXiv:2608. 05745v1 Announce Type: cross Abstract: Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity, motion, and scene dynamics.
By Yushe Cao, Shikun Feng, Fei Shen, Haikuo Peng, Jianqiang Xia, Yiheng Zhu, Dianxi Shi, Chun Yu
arXiv:2608.30450v1 Announce Type: new
Abstract: Video virtual try-on aims to transfer a target garment onto a moving person across video frames. Current methods rely on human parsing masks or pose ke...
By Shengyao Chen, Xianbing Sun, Liqing Zhang, Jianfu Zhang
arXiv:2602.24043v2 Announce Type: replace
Abstract: Reconstructing 3D clothed humans from monocular images and videos is a fundamental problem with applications in virtual try-on, avatar creation, an...
By Yingxuan You, Ren Li, Corentin Dumery, Cong Cao, Hao Li, Pascal Fua
arXiv:2604. 06010v2 Announce Type: replace Abstract: Video fundamentally intertwines two crucial axes: the dynamic content of a scene and the camera motion through which it is observed.
By Yukun Wang, Ruihuang Li, Jiale Tao, Shiyuan Yang, Liyi Chen, Zhantao Yang, Handz, Yulan Guo, Shuai Shao, Qinglin Lu
arXiv:2609.01276v1 Announce Type: new
Abstract: Complete 3D perception from egocentric video requires recovering the surrounding scene and the wearer's full-body motion in a shared metric frame. Exis...
By Kai Guan, Minchao Jiang, Ruichen WangLi, Wentao Zhu, Lei Zhang
arXiv:2606. 29531v1 Announce Type: cross Abstract: We propose MotionAtlas, a system for detailed captioning of motion-centric videos, comprising (1) a dedicated human-annotated benchmark, (2) a scalable, high-quality pipeline to construct training samples, and (3) a family of powerful Video-MLLMs.
By Weisong Liu, Haochen Wang, Kuan Gao, Yuhao Wang, Yikang Zhou, Zhongwei Ren, Jacky Mai, Anna Wang, Yanwei Li, Jason Li, Zhaoxiang Zhang
arXiv:2605. 01720v3 Announce Type: replace-cross Abstract: Existing large-scale sign language resources typically provide supervision only at the level of raw video-text alignment and are often produced in laboratory settings.
By Sen Fang, Hongbin Zhong, Yanxin Zhang, Dimitris N. Metaxas
arXiv:2512.12703v2 Announce Type: replace-cross
Abstract: Extracting human motion from large-scale web videos offers a scalable solution to the data scarcity issue in character animation. However, so...
By Boyuan Li, Sipeng Zheng, Bin Cao, Ruihua Song, Zongqing Lu
arXiv:2607. 22830v2 Announce Type: replace Abstract: In visual storytelling, human performances are central to creative intent and narrative meaning.
By Yuancheng Xu, Mingming He, Pablo Salamanca, Li Ma, Yash Kant, Emmett Steven, Paul Debevec, Ning Yu
arXiv:2605. 23045v2 Announce Type: replace-cross Abstract: Video representation learning has seen tremendous progress in recent years.
By Mantas Skackauskas, Xinyue Hao, Laura Sevilla-Lara
arXiv:2607.14935v2 Announce Type: replace
Abstract: Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world application...
By Xinhao Li, Yuhan Zhu, Xiangyu Zeng, Yuhao Dong, Haoning Wu, Zhiqiu Zhang, Yuandong Yang, Changlian Ma, Qingyu Zhang, Yansong Shi, Xinyu Chen, Haoran Chen, Zizheng Huang, Jun Zhang, Kun Ouyang, Lin Sui, Ziang Yan, Yicheng Xu, Chenting Wang, Yinan He, Hongjie Zhang, Yi Wang, Yu Qiao, Yali Wang, Ziwei Liu, Kai Chen, Limin Wang