arXiv Computer Vision
1d ago

Video-Conditioned Generative Joint 2D-3D Hand Motion Recovery

The paper introduces JoHan, a generative framework that directly recovers 2D and 3D hand motion from video sequences without intermediate per‑frame pose predictions. By jointly learning temporal dynamics and 2D‑3D correspondence, JoHan generates aligned pose sequences that improve temporal consistency and enable accurate estimation of the hand’s global position and orientation relative to the camera. Experiments on challenging benchmarks show that JoHan achieves higher accuracy and faster performance, producing smoother hand‑motion dynamics while maintaining high per‑frame pose accuracy.

By Chen Xu, Yunqi Li, Binbin Huang, Brent Yi, Shenghua Gao, Yi Ma