arXiv Computer Vision

Strike a Chord! Modal Kinetic Typography

arXiv Computer Vision
Sep 15

MorphoStyle: Motion Style Transfer with Morphology Control

MorphoStyle is a new framework for shape‑aware motion style transfer that uses a shape‑conditioned FSQ‑VAE. It disentangles style from content through a contrastive style encoder, a text‑guided style‑routing mechanism, and a manifold‑preserving style modulator. Experiments on benchmark datasets show that MorphoStyle outperforms existing baselines in both shape control and motion style transfer.

By Xin Feng, Eleonora D'Arnese, Mohan Sridharan
arXiv Computer Vision
Sep 11

Multi-Modal Controlled Coherent Motion Generation

The paper introduces MOCO, a diffusion-based framework that generates 3D avatar motions from concurrent multimodal inputs such as speech audio, text descriptions, and trajectory data. MOCO decouples motion generation by independently producing modality-specific motions at each denoising step and then assembling them according to spatial rules, iteratively refining the combined motion. This approach yields coherent, lifelike, and synchronized movements, outperforming existing baselines on a multimodal benchmark.

By Yifei Liu, Qiong Cao, Hongwei Yi, Huaiguang Jiang, Changxing Ding
arXiv Computer Vision
Sep 15

SignMimic: Robust High-Quality Sign Language Motion Generation via Human-Shape-Oblivious Pose Transfer Guidance

arXiv:2609.14122v1 Announce Type: new Abstract: We study the challenge of sign language video mimicking: given a driving video and a single reference frame, synthesize a video where the target signer...

By Zhewen He (New York University Abu Dhabi), Junyi Yu (New York University Abu Dhabi), Haomian Huang (New York University Abu Dhabi), Zhenhua Li (ChatSign Technology), Yi Fang (New York University Abu Dhabi, ChatSign Technology)
arXiv Computer Vision
Sep 28

Timo: $\textbf{T}$aming Mult$\textbf{i}$modal Diffusion Transformer for Human $\textbf{Mo}$tion Generation

The paper introduces Timo, a kinematics-aware multimodal diffusion transformer designed for human motion generation. Timo employs fully shared multimodal attention, flow matching, and geometric/rotational-kinematics supervision to better coordinate articulated motion, and uses a two-stage curriculum to align motion with text captions. The authors also present a new benchmark of 40,025 clips from six datasets, showing that Timo outperforms state‑of‑the‑art methods, achieving a 40.8% relative improvement over Kimodo on average.

By Zhao Wang, Jiangtao Hu, Jack Yu, Tao Yu
arXiv Computer Vision
Aug 31

Token-Budget Distillation: Transferring Full-Token Semantics to Compressed Video Vision-Language Models

Token-Budget Distillation (TBD) is a parameter‑efficient fine‑tuning framework that adapts video vision‑language models to a fixed token budget. It freezes the pretrained backbone, updates only LoRA adapters, and incorporates FlashVID visual token compression. TBD uses a dual‑path teacher‑student design with full‑token supervision and compressed student optimization, enabling the student to recover full‑token semantics while remaining efficient under aggressive token reduction.

By Xiaoyang Guo, Guoping Luo, Jusheng Zhang, Keze Wang, Wenhao Wang
arXiv Machine Learning
Sep 7

UniMate: One Unified Model to Animate Diverse Skeletons

UniMate is a unified foundation model that generates articulated motion for any skeleton from a rigged 3D asset and a text prompt, eliminating the need for test‑time optimization or per‑skeleton retraining. It uses a topology‑aware diffusion transformer that incorporates skeletal topology through graph‑aware attention bias, spectral rotary position embedding, and a global topological conditioner. Trained on the newly curated UniML3D dataset of 13,006 diverse motion sequences, UniMate outperforms existing baselines in quality, generalization, and efficiency, and supports zero‑shot cross‑topology transfer, in‑betweening, expansion, and text‑guided editing.

By Linzhan Mou, Jiahui Lei, Zhiyang Dou, Chenyue Cai, Chaoyue Song, Adam Finkelstein, Szymon Rusinkiewicz
Hugging Face Trending Papers
Sep 28

BiMoGen: Bidirectional Motion-Text Generation via Unified Masked Discrete Diffusion

BiMoGen introduces a unified masked discrete diffusion framework for bidirectional motion‑text generation, addressing the limitations of autoregressive models in capturing bidirectional dependencies between language and motion. The approach employs a two‑stage training strategy—decoupled uni‑ and cross‑modal pretraining followed by supervised fine‑tuning—to establish robust cross‑modal correspondence, and incorporates Generation‑Aware Self‑Correction to mitigate error propagation during inference. Experiments on HumanML3D and KIT‑ML show competitive performance on both text‑to‑motion and motion‑to‑text tasks, demonstrating the effectiveness of the proposed training and correction mechanisms.

arXiv Computer Vision
1d ago

The Shape of Speech: A Geometric Measure of Coarticulation for Speech-Driven 3D Facial Animation

The paper introduces a geometric measure of coarticulation for speech‑driven 3D facial animation, comparing lip‑path length to the shortest route through vowel, consonant, and vowel positions. Using only forced alignment, the measure evaluates four state‑of‑the‑art animation methods, revealing that all produce flatter lip trajectories than captured speech and that some methods lose 15–60% of the fast articulatory component. A pre‑registered viewer study confirms that damping real motion lowers perceived quality while exaggeration is not penalized, and viewers prefer real speech in 73.4% of sentence comparisons.

By Danzel Serrano, Przemyslaw Musialski