Kirin is a new framework that reconstructs 3D animal motion from in‑the‑wild videos, learns motion priors at scale, and generates realistic motion conditioned on text and image. It introduces AiM3D, the first large‑scale dataset of aligned video‑text‑motion tuples for quadruped animals, and uses an off‑the‑shelf image‑to‑3D model to automatically rig and animate 3D meshes with the generated motion. The framework and dataset provide a foundation for large‑scale, text and image‑conditioned animal motion generation and animation.
By Brian Nlong Zhao, Zhuoyang Pan, James M. Rehg, Jiajun Wu, Shangzhe Wu
arXiv:2609.09513v1 Announce Type: new
Abstract: Reconstructing a fully animatable 3D animal from a single image remains challenging because animation-ready assets require not only plausible geometry,...
By Chunyi Sun, Ruyi Zha, Weijian Deng, Junlin Han, Dylan Campbell, Stephen Gould
arXiv:2609.36217v1 Announce Type: new
Abstract: A deeper understanding of brain function requires a precise, structured characterization of behavior.Yet, extracting behavioral representations from vi...
By Xinming Dai, Qihang Jin, Tianshu Tan, Baiyuan Chen, Hanrui Lyu, Lenny Aharon, Kyle Daruwalla, Xun Helen Hou, Matthew R. Whiteway, Liam Paninski, Yizi Zhang
Artic-O is an end‑to‑end, feed‑forward framework that reconstructs articulated objects from sparse images by learning latent geometry. It maps multi‑state observations into a pretrained latent geometry space, uses a frozen flow‑matching decoder for complete‑shape priors, and fuses visual tokens with geometry latents in an image‑grounded part‑reasoning module to segment active parts and predict articulation. Trained with a geometry‑to‑articulation curriculum and a decoupled two‑pass strategy, Artic‑O achieves high reconstruction quality and articulation accuracy while drastically reducing inference time from 9 minutes to about 0.3 seconds per object.
By Xuyang Wang, Zhenyu Li, Jian Ding, Habib Slim, Peter Wonka, Hongdong Li, Mohamed Elhoseiny
UniMo introduces a unified point‑cloud based framework for generating 3D motion that works for both humans and animals, overcoming challenges posed by diverse skeletal topologies and limited animal datasets. It converts parametric skeletons into unparametric representations and uses dynamic sampling to focus on active joints. The authors also release UniML3D, a large motion‑language dataset with 145,907 sequences and 433,388 captions, and demonstrate state‑of‑the‑art performance on multiple benchmarks.
By Zeyu Zhang, Zhiyuan Zhang, Siheng Wang, Yiran Wang, Danning Li, Ian Reid, Richard Hartley
arXiv:2509.04276v3 Announce Type: replace
Abstract: We present a method for modeling articulated objects from sparse images with unknown camera poses. Existing approaches require dense multi-view obs...
By Jianning Deng, Kartic Subr, Hakan Bilen