4D-HOF: Hand-Object Flow Matching for Feed-Forward 4D Interaction Reconstruction
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2607. 11221v1 Announce Type: cross Abstract: Accurate monocular 4D hand reconstruction remains challenging.
arXiv:2605.20992v4 Announce Type: replace Abstract: We ask whether everyday open-world monocular videos can be turned into reusable 4D interaction primitives: articulated hand motion, object shape wi...
Accurate monocular 4D hand reconstruction remains challenging. Per-frame discriminative regressors lack temporal context and often produce jittery predictions.
arXiv:2609.36454v1 Announce Type: new Abstract: We study hand-object interaction (HOI) reconstruction from monocular RGB videos, where partial observations can produce visually plausible yet mechanic...
arXiv:2606.19156v2 Announce Type: replace Abstract: Dynamic 3D hand reconstruction from egocentric videos is essential for next-generation computing platforms such as AR/VR and AI glasses. Despite it...
The paper introduces JoHan, a generative framework that directly recovers 2D and 3D hand motion from video sequences without intermediate per‑frame pose predictions. By jointly learning temporal dynamics and 2D‑3D correspondence, JoHan generates aligned pose sequences that improve temporal consistency and enable accurate estimation of the hand’s global position and orientation relative to the camera. Experiments on challenging benchmarks show that JoHan achieves higher accuracy and faster performance, producing smoother hand‑motion dynamics while maintaining high per‑frame pose accuracy.