arXiv Computer Vision

ORMA: Optimization-based Monocular 4D Reconstruction of Articulated Animals

ORMA is a training‑free framework that reconstructs articulated 4D representations of animals from monocular videos by decoupling pose and shape. It uses predicted pose as a reference for optimization and generative 3D priors to refine shape, aligning the result with the SMAL+ parametric model. The method combines per‑frame pose estimates with globally consistent camera poses, and further refines the reconstruction using self‑supervised DINO correspondences and temporal consistency, achieving improved accuracy on the new PAW4D benchmark and diverse real‑world videos.

arXiv Computer Vision
Sep 3

Kirin: Animal Motion Generation from In-the-Wild Video

Kirin is a new framework that reconstructs 3D animal motion from in‑the‑wild videos, learns motion priors at scale, and generates realistic motion conditioned on text and image. It introduces AiM3D, the first large‑scale dataset of aligned video‑text‑motion tuples for quadruped animals, and uses an off‑the‑shelf image‑to‑3D model to automatically rig and animate 3D meshes with the generated motion. The framework and dataset provide a foundation for large‑scale, text and image‑conditioned animal motion generation and animation.

By Brian Nlong Zhao, Zhuoyang Pan, James M. Rehg, Jiajun Wu, Shangzhe Wu
arXiv Computer Vision
1d ago

Sparse-View Interpretable 3D Animal Behavior Representations for Neural Encoding and Decoding

arXiv:2609.36217v1 Announce Type: new Abstract: A deeper understanding of brain function requires a precise, structured characterization of behavior.Yet, extracting behavioral representations from vi...

By Xinming Dai, Qihang Jin, Tianshu Tan, Baiyuan Chen, Hanrui Lyu, Lenny Aharon, Kyle Daruwalla, Xun Helen Hou, Matthew R. Whiteway, Liam Paninski, Yizi Zhang
arXiv Computer Vision
Sep 11

Artic-O: End-to-End Articulated Object Reconstruction via Latent Geometry Learning

Artic-O is an end‑to‑end, feed‑forward framework that reconstructs articulated objects from sparse images by learning latent geometry. It maps multi‑state observations into a pretrained latent geometry space, uses a frozen flow‑matching decoder for complete‑shape priors, and fuses visual tokens with geometry latents in an image‑grounded part‑reasoning module to segment active parts and predict articulation. Trained with a geometry‑to‑articulation curriculum and a decoupled two‑pass strategy, Artic‑O achieves high reconstruction quality and articulation accuracy while drastically reducing inference time from 9 minutes to about 0.3 seconds per object.

By Xuyang Wang, Zhenyu Li, Jian Ding, Habib Slim, Peter Wonka, Hongdong Li, Mohamed Elhoseiny
arXiv Computer Vision
Sep 14

UniMo: Unifying Human and Animal Motion Generation

UniMo introduces a unified point‑cloud based framework for generating 3D motion that works for both humans and animals, overcoming challenges posed by diverse skeletal topologies and limited animal datasets. It converts parametric skeletons into unparametric representations and uses dynamic sampling to focus on active joints. The authors also release UniML3D, a large motion‑language dataset with 145,907 sequences and 433,388 captions, and demonstrate state‑of‑the‑art performance on multiple benchmarks.

By Zeyu Zhang, Zhiyuan Zhang, Siheng Wang, Yiran Wang, Danning Li, Ian Reid, Richard Hartley
arXiv Computer Vision
Sep 15

MoCapAnything V2: End-to-End Motion Capture for Arbitrary Skeletons

arXiv:2604.28130v4 Announce Type: replace Abstract: Recent methods for arbitrary-skeleton motion capture from monocular video follow a factorized pipeline, where a Video-to-Pose network predicts join...

By Kehong Gong, Zhengyu Wen, Dao Thien Phong, Mingxi Xu, Weixia He, Qi Wang, Ning Zhang, Zhengyu Li, Guanli Hou, Dongze Lian, Xiaoyu He, Mingyuan Zhang, Hanwang Zhang
arXiv AI
Jun 29

HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration

arXiv:2606. 28215v1 Announce Type: cross Abstract: Extracting dynamic 4D object interactions from massive, in-the-wild monocular videos offers a highly efficient data collection pathway for scaling Embodied AI and training VLAs.

By Jiaxin Li, Yuxiang Wu, Zhenkai Zhang, Xinrui Shi, Haoyuan Wang, Yichen Zhao, Su Linxiang, Chenyang Yu, Mingyu Zhang, Yifan Ding, Boran Wen, Li Zhang, Ruiyang Liu, Yong-Lu Li
arXiv Computer Vision
Aug 26

SceneReGen: Generative Reconstruction of 3D Scenes from a Single Image

SceneReGen is a new framework for reconstructing 3D scenes from a single image by generating and assembling complete object meshes within a shared observation‑aligned scene frame. It uses selective pose factorization to encode each object’s observed orientation directly into the generated mesh, while estimating translation and scale from instance‑level and global scene cues. Evaluated on the 3D‑FUTURE dataset, SceneReGen outperforms existing methods on scene‑level metrics and shows strong performance on object‑level metrics, demonstrating its effectiveness in autonomous‑driving and embodied‑AI scenarios.

By Zefan Tian, Yuteng Ye, Yiheng Zhang, Yuhang Yang, Xueqiang Lv, Shizhou Zhang, Le Liu, Di Xu
Hugging Face Trending Papers
Aug 5

Promptable Animal Pose Tracking Across Species

Animal pose estimation and tracking is important for wildlife monitoring and conservation research, and with limited expert time for labelling automated approaches are imperative. While human pose estimation and tracking has seen rapid progress thanks to large annotated datasets, animal pose remain challenging, due to large morphological and behavioural differences between species and limited annotated data.