HandFlow: Fully Generative 4D Hand Recovery with Flow Matching
arXiv:2607. 11221v1 Announce Type: cross Abstract: Accurate monocular 4D hand reconstruction remains challenging.
arXiv:2607. 11221v1 Announce Type: cross Abstract: Accurate monocular 4D hand reconstruction remains challenging.
arXiv:2605.20992v4 Announce Type: replace Abstract: We ask whether everyday open-world monocular videos can be turned into reusable 4D interaction primitives: articulated hand motion, object shape wi...
Accurate monocular 4D hand reconstruction remains challenging. Per-frame discriminative regressors lack temporal context and often produce jittery predictions.
arXiv:2609.36454v1 Announce Type: new Abstract: We study hand-object interaction (HOI) reconstruction from monocular RGB videos, where partial observations can produce visually plausible yet mechanic...
arXiv:2606.19156v2 Announce Type: replace Abstract: Dynamic 3D hand reconstruction from egocentric videos is essential for next-generation computing platforms such as AR/VR and AI glasses. Despite it...
The paper introduces JoHan, a generative framework that directly recovers 2D and 3D hand motion from video sequences without intermediate per‑frame pose predictions. By jointly learning temporal dynamics and 2D‑3D correspondence, JoHan generates aligned pose sequences that improve temporal consistency and enable accurate estimation of the hand’s global position and orientation relative to the camera. Experiments on challenging benchmarks show that JoHan achieves higher accuracy and faster performance, producing smoother hand‑motion dynamics while maintaining high per‑frame pose accuracy.
arXiv:2606. 28215v1 Announce Type: cross Abstract: Extracting dynamic 4D object interactions from massive, in-the-wild monocular videos offers a highly efficient data collection pathway for scaling Embodied AI and training VLAs.
Grasp in Gaussians (GraG) is a fast, robust method for reconstructing dynamic 3D hand‑object interactions from a single monocular video. It leverages pretrained hand and object priors and represents the scene with a compact Sum‑of‑Gaussians (SoG) model, enabling efficient tracking while preserving geometric fidelity. Experiments show GraG achieves temporally coherent reconstructions on long sequences 4.4×–38.9× faster than prior work.
arXiv:2609.38615v1 Announce Type: cross Abstract: Egocentric videos of human manipulation provide valuable visual experience for embodied intelligence, yet collecting such data at scale is costly. Ex...
arXiv:2609.08636v1 Announce Type: cross Abstract: Egocentric 4D interaction forecasting aims to anticipate both where future interactions will occur in 3D and how the human body will move to realize...
arXiv:2609.19119v1 Announce Type: new Abstract: Human videos contain rich causal evidence for robot manipulation: they reveal how hand motion induces object motion and produces task-relevant changes...
FAMOS is a feed‑forward model that predicts movable‑part segmentation and joint parameters from a sparse, unordered set of partial point clouds. It jointly reasons over multiple observations using a Multi‑state Articulation Transformer that alternates state‑wise and global attention, and introduces an observed articulation span objective to supervise motion ranges across inputs. A procedural data generator supplies self‑annotated assets for training, and experiments on PartNet‑Mobility, ACD, and ArtiCraft‑10K show consistent improvements over existing feed‑forward and optimization‑based baselines.