PrePARE: Pre-AA Token Pruning for Frozen Multi-View Geometry Transformers
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2609.23733v1 Announce Type: new Abstract: Feed-forward visual geometry models such as the Visual Geometry Grounded Transformer (VGGT) have recently enabled direct 3D reconstruction from multi-v...
arXiv:2606. 12412v1 Announce Type: cross Abstract: Vision-language models (VLMs) project images into hundreds to thousands of visual tokens, making decoder inference expensive in both attention computation and KV-cache memory.
arXiv:2609.36374v1 Announce Type: new Abstract: Feed-forward 3D reconstruction models have achieved impressive performance by scaling model and dataset size, but their cost excludes most research gro...
arXiv:2608.22526v1 Announce Type: new Abstract: We introduce RS$^3$-Prune, a training-free token-pruning recipe that instantiates as a small set of inference time hooks atop existing video object seg...
arXiv:2609.05916v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) achieve strong multimodal understanding, but the hundreds to thousands of visual tokens they process impose subs...
arXiv:2608.28706v1 Announce Type: new Abstract: ViT detectors fix a uniform token grid before any learned stage. A native-resolution aerial detector must then choose between resolving few-pixel objec...