Foundational feature fusion for conditional flow matching in 6D pose estimation
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2609.39748v1 Announce Type: new Abstract: Scaling has become a primary driver of progress in language and vision foundation models, yet its role in precise correspondence matching remains under...
Moving6DPoSe is a multimodal database for monocular 6D pose estimation and segmentation of moving objects, comprising two subsets: real-world recordings (Moving6DPoSe‑R) and synthetic sequences (Moving6DPoSe‑S). It includes 16 scanned objects, 1,702 real and synthetic rosbags, and annotations for semantic segmentation, object detection, and monocular 6D pose estimation. Baseline results show that event-based representations outperform conventional RGB images for moving‑object segmentation, while monocular orientation estimation remains challenging.
arXiv:2607. 11221v1 Announce Type: cross Abstract: Accurate monocular 4D hand reconstruction remains challenging.
arXiv:2608.30450v1 Announce Type: new Abstract: Video virtual try-on aims to transfer a target garment onto a moving person across video frames. Current methods rely on human parsing masks or pose ke...
Accurate monocular 4D hand reconstruction remains challenging. Per-frame discriminative regressors lack temporal context and often produce jittery predictions.
arXiv:2610.01758v1 Announce Type: new Abstract: Category-level object pose estimation (COPE), capable of generalizing to intra-class unknown objects, has become a core technique for robotic 3D scene...