HandFlow: Fully Generative 4D Hand Recovery with Flow Matching
Accurate monocular 4D hand reconstruction remains challenging. Per-frame discriminative regressors lack temporal context and often produce jittery predictions.
Accurate monocular 4D hand reconstruction remains challenging. Per-frame discriminative regressors lack temporal context and often produce jittery predictions.
arXiv:2607. 11221v1 Announce Type: cross Abstract: Accurate monocular 4D hand reconstruction remains challenging.
arXiv:2609.24424v1 Announce Type: new Abstract: Monocular RGB-based hand pose estimation has emerged as a critical research frontier in computer vision. The local hand pose estimation methods predict...
arXiv:2603.25175v2 Announce Type: replace Abstract: Monocular egocentric 3D pose estimation is difficult because severe foreshortening, self-occlusion, and a restricted field of view often remove the...
arXiv:2603.12789v3 Announce Type: replace Abstract: Recent advances in 3D foundation models have enabled joint reconstruction of humans and their surrounding environments. However, combining independ...
arXiv:2601.13913v3 Announce Type: replace Abstract: We consider monocular 3D human pose estimation (HPE), where the goal is to predict 3D human skeletal joints from a single 2D image, typically via 2...
arXiv:2608.30521v1 Announce Type: new Abstract: Conventional algebraic triangulation solves 3D human pose estimation (HPE) from multi-view 2D keypoints. The typical approach, decoding 2D keypoints fr...
arXiv:2608. 20093v1 Announce Type: new Abstract: In this work, we present HandMvNet, one of the first real-time method designed to estimate 3D hand motion and shape from multi-view camera images.
arXiv:2609.27227v1 Announce Type: new Abstract: Objective assessment of robotic surgery uses instrument kinematics, which must be reconstructed when only video is available. We introduce a kinematic...
The paper introduces CalfVO, a monocular visual odometry system that operates without camera intrinsics, test‑time optimization, bundle adjustment, or loop closure. Using a transformer, it predicts relative poses with separate rotation and translation confidences over overlapping image windows, then aggregates these predictions via a confidence‑weighted module to produce a single trajectory. CalfVO achieves the highest accuracy among calibration‑free methods across five benchmarks and runs at 53 FPS, outperforming all baselines.
arXiv:2604.28130v4 Announce Type: replace Abstract: Recent methods for arbitrary-skeleton motion capture from monocular video follow a factorized pipeline, where a Video-to-Pose network predicts join...
arXiv:2512. 16919v2 Announce Type: replace-cross Abstract: Perceiving and reconstructing 3D scene geometry from visual inputs is crucial for autonomous driving.