arXiv:2609.36937v1 Announce Type: cross
Abstract: Human image animation aims to transfer motion from a driving video to subjects in a reference image. Despite remarkable progress in video generation,...
By Sangeyl Lee, Seunghyun Shin, Seungho Park, Wooseok Jeon, Hae-Gon Jeon
BooM‑VVT is a mask‑free video virtual try‑on framework that builds on a keyframe‑driven paradigm. It introduces a multi‑stage training strategy using image‑level pseudo data to learn mask‑free localization, a garment‑sensitive keyframe sampling method to capture garment appearance, and a Frame‑Shared 3D‑RoPE module to align keyframes with target video frames for accurate garment detail transfer. The authors also release OmniView, a large‑scale multi‑view try‑on dataset, and demonstrate that BooM‑VVT outperforms existing methods in temporal consistency and garment fidelity.
By Wei Zhang, Xin Li, Peishu Shi, Jialin Gao, Xuekang Peng, Zhichao Lian, Yeying Jin
BooM-VVT is a mask‑free video virtual try‑on framework that builds on a keyframe‑driven paradigm. It uses a multi‑stage training strategy with image‑level pseudo data to learn mask‑free localization, introduces Garment‑Sensitive Keyframe Sampling to capture garment appearance, and employs Frame‑Shared 3D‑RoPE for spatiotemporal correspondence. The authors also create the OmniView dataset to support diverse camera viewpoints and tasks, achieving superior temporal consistency and garment fidelity compared to existing methods.
arXiv:2606. 29531v1 Announce Type: cross Abstract: We propose MotionAtlas, a system for detailed captioning of motion-centric videos, comprising (1) a dedicated human-annotated benchmark, (2) a scalable, high-quality pipeline to construct training samples, and (3) a family of powerful Video-MLLMs.
By Weisong Liu, Haochen Wang, Kuan Gao, Yuhao Wang, Yikang Zhou, Zhongwei Ren, Jacky Mai, Anna Wang, Yanwei Li, Jason Li, Zhaoxiang Zhang
arXiv:2609.40347v1 Announce Type: new
Abstract: We introduce VideoMSN, a Masked Siamese Network framework for efficient self-supervised spatio-temporal representation learning in videos. Instead of r...
By Owais Iqbal, Sudipta Sarkar, Shyam Marjit, Omprakash Chakraborty, Anirban Chakraborty, Abir Das
arXiv:2608. 05115v1 Announce Type: cross Abstract: Can computer vision help make classrooms safer?
By Paritosh Parmar, Landy Lan, Hong Yang, Chen Yi, Chiat Pin Tay
arXiv:2605. 23045v2 Announce Type: replace-cross Abstract: Video representation learning has seen tremendous progress in recent years.
By Mantas Skackauskas, Xinyue Hao, Laura Sevilla-Lara
arXiv:2410. 19553v2 Announce Type: replace-cross Abstract: This paper explores the impact of occlusions in video action detection.
By Rajat Modi, Vibhav Vineet, Yogesh Singh Rawat
The paper introduces a framework for generating multi‑view images of a person within a natural scene, addressing the scarcity of paired multi‑view datasets for human subjects. It evaluates existing diffusion‑based image‑editing models and finds they often hallucinate head‑turn angles, leading to inconsistent backgrounds. To overcome this, the authors propose the Head Scene Rotation Difference (HSRD) metric, which separates camera movement from head pose changes and enables reliable assessment of 3D spatial parallax for constructing high‑quality synthetic datasets.
By Mahir Majid, Young Kyung Kim, Guillermo Sapiro
Can computer vision help make classrooms safer? In this pilot study, we investigate privacy-aware and computationally efficient classroom incident recognition from CCTV-style observations.
VOR-Bench is a new benchmark for video object removal that addresses shortcomings in current evaluation methods by providing a dataset with paired edited videos and graffiti masks, a realistic motion-capable paired-video acquisition framework (rMPAF), and a perception-driven scoring model (VOR-MDSM). The dataset includes diverse data from model-generated, tool-rendered, and camera-captured sources, ensuring robust real-world assessment. Experiments show that VOR-Bench’s evaluation results correlate strongly (ρ > 0.9) with human subjective judgments, bridging the gap between traditional metrics and human preference.
By Haonan Huang, Tianrui Qiu, Xianghao Zang, Yinan Du, Zhixiang He, Chi Zhang, Hao Sun, Zhongjiang He, Tianwei Cao, Xuchong Zhang, Hongbin Sun, Kongming Liang, Zhanyu Ma
arXiv:2609.38172v1 Announce Type: cross
Abstract: Teaching humanoids loco-manipulation skills, such as carrying diverse objects, via visual imitation is a promising path toward generalist robots. How...
By Zihan Wang, Zhen Wu, Pieter Abbeel, Rocky Duan, Jitendra Malik, Carmelo Sferrazza, C. Karen Liu, Guanya Shi, Angjoo Kanazawa