VASC: Value-Aware Sparse Attention with Cross-Layer Memory for Efficient 3D Reconstruction
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2607. 03612v1 Announce Type: cross Abstract: Feed-forward 3D reconstruction (F3R) transformers have recently achieved remarkable success.
arXiv:2608. 10519v2 Announce Type: replace Abstract: InfinityStar extends visual autoregressive generation to video through a sequence of image and clip pyramids.
arXiv:2609.23733v1 Announce Type: new Abstract: Feed-forward visual geometry models such as the Visual Geometry Grounded Transformer (VGGT) have recently enabled direct 3D reconstruction from multi-v...
Visual Geometry Grounded Transformer (VGGT) recovers dense 3D scene structure from multi-view images in one forward pass, but quadratic cross-frame attention limits its scalability. Existing training-free accelerators reduce computation uniformly along one axis, missing layer heterogeneity.
InfinityStar extends visual autoregressive generation to video through a sequence of image and clip pyramids. Its changing scale and cross-clip context, however, leave late-scale attention costly and make sparse patterns reused from diffusion or image VAR models unreliable.
LoG-VGGT is a memory‑efficient framework for long‑sequence 3D reconstruction that balances local temporal modeling with global camera consistency. It uses cross‑window attention in a small subset of transformer blocks to propagate information across adjacent temporal windows while keeping memory usage bounded. A global camera consistency refinement module further improves long‑horizon pose stability by enforcing scene‑level constraints through cross‑attention between camera and compact register tokens, leading to better depth accuracy and robust camera pose estimation on multiple benchmarks.