arXiv Computer Vision

HuC-VideoMAE: Human-Centric Video Masked Autoencoding from synthetic data

arXiv Computer Vision
Sep 4

BooM-VVT: Boosting Mask-Free Video Virtual Try-On with Image-Level Pseudo Data

BooM‑VVT is a mask‑free video virtual try‑on framework that builds on a keyframe‑driven paradigm. It introduces a multi‑stage training strategy using image‑level pseudo data to learn mask‑free localization, a garment‑sensitive keyframe sampling method to capture garment appearance, and a Frame‑Shared 3D‑RoPE module to align keyframes with target video frames for accurate garment detail transfer. The authors also release OmniView, a large‑scale multi‑view try‑on dataset, and demonstrate that BooM‑VVT outperforms existing methods in temporal consistency and garment fidelity.

By Wei Zhang, Xin Li, Peishu Shi, Jialin Gao, Xuekang Peng, Zhichao Lian, Yeying Jin
Hugging Face Trending Papers
Sep 3

BooM-VVT: Boosting Mask-Free Video Virtual Try-On with Image-Level Pseudo Data

BooM-VVT is a mask‑free video virtual try‑on framework that builds on a keyframe‑driven paradigm. It uses a multi‑stage training strategy with image‑level pseudo data to learn mask‑free localization, introduces Garment‑Sensitive Keyframe Sampling to capture garment appearance, and employs Frame‑Shared 3D‑RoPE for spatiotemporal correspondence. The authors also create the OmniView dataset to support diverse camera viewpoints and tasks, achieving superior temporal consistency and garment fidelity compared to existing methods.

arXiv AI
Jun 30

MotionAtlas: Detailed Region Captioning for Motion-Centric Videos

arXiv:2606. 29531v1 Announce Type: cross Abstract: We propose MotionAtlas, a system for detailed captioning of motion-centric videos, comprising (1) a dedicated human-annotated benchmark, (2) a scalable, high-quality pipeline to construct training samples, and (3) a family of powerful Video-MLLMs.

By Weisong Liu, Haochen Wang, Kuan Gao, Yuhao Wang, Yikang Zhou, Zhongwei Ren, Jacky Mai, Anna Wang, Yanwei Li, Jason Li, Zhaoxiang Zhang
arXiv Computer Vision
Sep 7

An Evaluation Framework for Generating Multi-View Images of a Person in a Scene

The paper introduces a framework for generating multi‑view images of a person within a natural scene, addressing the scarcity of paired multi‑view datasets for human subjects. It evaluates existing diffusion‑based image‑editing models and finds they often hallucinate head‑turn angles, leading to inconsistent backgrounds. To overcome this, the authors propose the Head Scene Rotation Difference (HSRD) metric, which separates camera movement from head pose changes and enables reliable assessment of 3D spatial parallax for constructing high‑quality synthetic datasets.

By Mahir Majid, Young Kyung Kim, Guillermo Sapiro
arXiv AI
Sep 16

VOR-Bench: A Human Perception-Driven Benchmark for Video Object Removal

VOR-Bench is a new benchmark for video object removal that addresses shortcomings in current evaluation methods by providing a dataset with paired edited videos and graffiti masks, a realistic motion-capable paired-video acquisition framework (rMPAF), and a perception-driven scoring model (VOR-MDSM). The dataset includes diverse data from model-generated, tool-rendered, and camera-captured sources, ensuring robust real-world assessment. Experiments show that VOR-Bench’s evaluation results correlate strongly (ρ > 0.9) with human subjective judgments, bridging the gap between traditional metrics and human preference.

By Haonan Huang, Tianrui Qiu, Xianghao Zang, Yinan Du, Zhixiang He, Chi Zhang, Hao Sun, Zhongjiang He, Tianwei Cao, Xuchong Zhang, Hongbin Sun, Kongming Liang, Zhanyu Ma