arXiv AI

Robust Motion Generation using Part-level Reliable Data from Videos

arXiv Computer Vision
3d ago

SeMoCo: A Semantic-First Motion Codec for Motion Language Modeling

arXiv:2608.24334v1 Announce Type: new Abstract: Discrete motion representations have substantially advanced autoregressive text-to-motion generation. However, most motion tokenizers are optimized for...

By Tianlv Huang, Hetian Guo, Ziyi Cai, Song Wang, Yanping Zhang, Zipei Fan, Xuan Song, Guangming Wu, Xin Zheng
arXiv AI
Jun 30

MotionAtlas: Detailed Region Captioning for Motion-Centric Videos

arXiv:2606. 29531v1 Announce Type: cross Abstract: We propose MotionAtlas, a system for detailed captioning of motion-centric videos, comprising (1) a dedicated human-annotated benchmark, (2) a scalable, high-quality pipeline to construct training samples, and (3) a family of powerful Video-MLLMs.

By Weisong Liu, Haochen Wang, Kuan Gao, Yuhao Wang, Yikang Zhou, Zhongwei Ren, Jacky Mai, Anna Wang, Yanwei Li, Jason Li, Zhaoxiang Zhang
Hugging Face Trending Papers
Jul 2

PWM-ArtGen: Part World Model for Articulated Object Generation

The key challenge in articulated 3D object generation from a single image is accurately predicting the underlying kinematic structure. Existing methods either infer kinematic parameters directly from a static image that lacks dynamic part-level kinematic relationships, or estimate parameters from visual dynamics generated from a single image, which is prone to accumulated errors of two steps.

arXiv Computer Vision
Aug 21

ID-V2V: Identity-Preserving Video Restylization

arXiv:2607. 22830v2 Announce Type: replace Abstract: In visual storytelling, human performances are central to creative intent and narrative meaning.

By Yuancheng Xu, Mingming He, Pablo Salamanca, Li Ma, Yash Kant, Emmett Steven, Paul Debevec, Ning Yu
arXiv Computer Vision
3d ago

Layer-Aware Video Composition via Split-then-Merge

The paper introduces Split-then-Merge (StM), a new framework for generative video composition that improves control and tackles data scarcity. StM divides a large set of unlabeled videos into dynamic foreground and background layers, then self‑composes them to learn how subjects interact with varied scenes. The method employs a transformation‑aware training pipeline with multi‑layer fusion, augmentation, and an identity‑preservation loss, achieving superior performance over state‑of‑the‑art methods in both quantitative and qualitative evaluations.

By Ozgur Kara, Yujia Chen, Ming-Hsuan Yang, James M. Rehg, Wen-Sheng Chu, Du Tran
arXiv AI
Aug 5

TransVLM: A Vision-Language Framework and Benchmark for Detecting Any Shot Transitions

arXiv:2604. 27975v2 Announce Type: replace-cross Abstract: Traditional Shot Boundary Detection (SBD) inherently struggles with complex transitions by formulating the task around isolated cut points, frequently yielding corrupted video shots.

By Ce Chen, Yi Ren, Yuanming Li, Viktor Goriachko, Zhenhui Ye, Zujin Guo, Zhibin Hong, Mingming Gong