HandFlow: Fully Generative 4D Hand Recovery with Flow Matching
Accurate monocular 4D hand reconstruction remains challenging. Per-frame discriminative regressors lack temporal context and often produce jittery predictions.
The paper introduces JoHan, a generative framework that directly recovers 2D and 3D hand motion from video sequences without intermediate per‑frame pose predictions. By jointly learning temporal dynamics and 2D‑3D correspondence, JoHan generates aligned pose sequences that improve temporal consistency and enable accurate estimation of the hand’s global position and orientation relative to the camera. Experiments on challenging benchmarks show that JoHan achieves higher accuracy and faster performance, producing smoother hand‑motion dynamics while maintaining high per‑frame pose accuracy.
Accurate monocular 4D hand reconstruction remains challenging. Per-frame discriminative regressors lack temporal context and often produce jittery predictions.
arXiv:2607. 11221v1 Announce Type: cross Abstract: Accurate monocular 4D hand reconstruction remains challenging.
arXiv:2608.22341v1 Announce Type: cross Abstract: Lifting 3D hand poses from 2D monocular representations remains challenging due to the limited availability of large-scale, diverse 3D-annotated hand...
Grasp in Gaussians (GraG) is a fast, robust method for reconstructing dynamic 3D hand‑object interactions from a single monocular video. It leverages pretrained hand and object priors and represents the scene with a compact Sum‑of‑Gaussians (SoG) model, enabling efficient tracking while preserving geometric fidelity. Experiments show GraG achieves temporally coherent reconstructions on long sequences 4.4×–38.9× faster than prior work.
MINT is a foundation model that directly predicts world-space two-hand trajectories from egocentric RGB video, jointly estimating camera motion, hand states, and hand presence in a single spatiotemporal representation. It uses an open-source labeling pipeline, EGOPIPELINE, to generate large-scale pseudo-labels for pretraining, followed by fine-tuning on a small set of high-quality joint annotations. The model outperforms existing multi-stage approaches in accuracy and speed, and generalizes zero‑shot to unseen egocentric datasets.
arXiv:2606.19156v2 Announce Type: replace Abstract: Dynamic 3D hand reconstruction from egocentric videos is essential for next-generation computing platforms such as AR/VR and AI glasses. Despite it...
arXiv:2610.08782v1 Announce Type: cross Abstract: Existing methods for 4D hand-object reconstruction often rely on costly per-sequence optimization, while generative approaches typically synthesize i...
arXiv:2609.24424v1 Announce Type: new Abstract: Monocular RGB-based hand pose estimation has emerged as a critical research frontier in computer vision. The local hand pose estimation methods predict...
arXiv:2608. 20093v1 Announce Type: new Abstract: In this work, we present HandMvNet, one of the first real-time method designed to estimate 3D hand motion and shape from multi-view camera images.
arXiv:2608.20308v2 Announce Type: replace Abstract: Egocentric video offers scalable manipulation data for embodied AI, yet recovering metric 3D hand trajectories remains challenging due to severe ob...
PACT is an end‑to‑end model that jointly learns human pose, contacts, and contact forces from monocular video. It augments a human reconstruction foundation model with learnable contact‑force tokens and a temporal transformer, and uses physics‑based supervision to enforce consistency between motion and forces. The authors also create a data annotation pipeline and a real‑world climbing benchmark, ForceWall, to train and evaluate the system, achieving state‑of‑the‑art performance and better generalization than staged approaches.
arXiv:2608. 20308v1 Announce Type: new Abstract: Egocentric video offers scalable manipulation data for embodied AI, yet recovering metric 3D hand trajectories remains challenging due to severe object occlusion and frequent out-of-sight gaps.