FAST: Flow Any Scene Transformer
arXiv:2609.39748v1 Announce Type: new Abstract: Scaling has become a primary driver of progress in language and vision foundation models, yet its role in precise correspondence matching remains under...
arXiv:2609.39748v1 Announce Type: new Abstract: Scaling has become a primary driver of progress in language and vision foundation models, yet its role in precise correspondence matching remains under...
Moving6DPoSe is a multimodal database for monocular 6D pose estimation and segmentation of moving objects, comprising two subsets: real-world recordings (Moving6DPoSe‑R) and synthetic sequences (Moving6DPoSe‑S). It includes 16 scanned objects, 1,702 real and synthetic rosbags, and annotations for semantic segmentation, object detection, and monocular 6D pose estimation. Baseline results show that event-based representations outperform conventional RGB images for moving‑object segmentation, while monocular orientation estimation remains challenging.
arXiv:2607. 11221v1 Announce Type: cross Abstract: Accurate monocular 4D hand reconstruction remains challenging.
arXiv:2608.30450v1 Announce Type: new Abstract: Video virtual try-on aims to transfer a target garment onto a moving person across video frames. Current methods rely on human parsing masks or pose ke...
Accurate monocular 4D hand reconstruction remains challenging. Per-frame discriminative regressors lack temporal context and often produce jittery predictions.
arXiv:2610.01758v1 Announce Type: new Abstract: Category-level object pose estimation (COPE), capable of generalizing to intra-class unknown objects, has become a core technique for robotic 3D scene...
arXiv:2609.39116v1 Announce Type: new Abstract: Prior-free 6D object pose tracking seeks to recover the trajectory of an unseen object from a single RGB video without object-specific CAD models, pose...
arXiv:2610.01286v1 Announce Type: new Abstract: Recent depth foundation models like Depth Anything 3 (DA3) achieve remarkable multi-view depth estimation but assume static 3D scenes, limiting their a...
arXiv:2603.12789v3 Announce Type: replace Abstract: Recent advances in 3D foundation models have enabled joint reconstruction of humans and their surrounding environments. However, combining independ...
arXiv:2606. 06903v1 Announce Type: cross Abstract: Human image animation aims to generate a video from a static reference image, guided by pose information extracted from a driving video.
arXiv:2607. 04930v1 Announce Type: cross Abstract: In the pursuit of robust and generalizable category-level object pose estimation, most existing methods adopt parametric formulations that learn effective representations from data, yet they primarily encode category-level patterns into fixed shape priors or static parameter weights, which limits their scalability to highly diverse instances.
arXiv:2512. 16919v2 Announce Type: replace-cross Abstract: Perceiving and reconstructing 3D scene geometry from visual inputs is crucial for autonomous driving.