FAST: Flow Any Scene Transformer
arXiv:2609.39748v1 Announce Type: new Abstract: Scaling has become a primary driver of progress in language and vision foundation models, yet its role in precise correspondence matching remains under...
arXiv:2609.39748v1 Announce Type: new Abstract: Scaling has become a primary driver of progress in language and vision foundation models, yet its role in precise correspondence matching remains under...
arXiv:2509.22650v3 Announce Type: replace Abstract: Most existing approaches to referring segmentation achieve strong performance only through fine-tuning or by composing multiple pre-trained models,...
arXiv:2607. 03612v1 Announce Type: cross Abstract: Feed-forward 3D reconstruction (F3R) transformers have recently achieved remarkable success.
Classical image correspondence is solved at the level of sparse keypoints or dense pixels, but the systems that consume these matches - object-level mapping, topological navigation, scene-graph maintenance - reason about whole objects. Recent work narrows this gap by matchng directly at the level of instance segments: a class-agnostic segmenter partitions each image, and per-segment descriptors are obtained by pooling features from large 3D foundation models over the masks.
Dyna3 is a training‑free framework that extends the depth foundation model DA3 to perform 4D dynamic scene reconstruction without fine‑tuning. By leveraging DA3’s cross‑view features and a best‑match search, it distinguishes static surfaces from moving objects, and uses vision‑language models to generate semantic prompts for SAM 3 to achieve precise instance‑level segmentation. Experiments on four datasets show Dyna3 outperforms correspondence‑trained methods, improving dynamic object segmentation by +5.5 pp, speeding pose estimation 13×, and reducing memory usage 4–8×.
Local feature matching is a fundamental component of photogrammetry, enabling accurate image correspondence critical for tasks such as 3D reconstruction, stereo mapping, and visual localization. While recent detector-free matching methods, like LoFTR, have advanced the field, the global features obtained by leveraging the global-range modeling capacity of the unconstrained attention mechanism compromise the model's attention to the salient structures in certain scenarios.
arXiv:2604. 00086v2 Announce Type: replace-cross Abstract: The field of computer vision has experienced significant advancements through scalable vision encoders and multimodal pre-training frameworks.
arXiv:2608. 11093v1 Announce Type: new Abstract: Cross-view feature matching aims to establish reliable correspondences across images with large viewpoint variations.
arXiv:2608.29183v1 Announce Type: new Abstract: Conditional flow matching has enabled a step forward in object 6D pose estimation, achieving state-of-the-art performance by progressively denoising an...
arXiv:2608. 05745v1 Announce Type: cross Abstract: Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity, motion, and scene dynamics.
arXiv:2605.12491v2 Announce Type: replace Abstract: Vision Transformers (ViTs) learn rich visual-semantic representations through all-to-all self-attention among patch tokens. However, this design im...
arXiv:2512.15708v2 Announce Type: replace Abstract: Foundation models are vital tools in various Computer Vision applications. They take as input a single RGB image and output a deep feature represen...