arXiv Computer Vision By Fengyi Zhang, Holger Caesar, Xiangyu Sun, Zheng Zhang, Zi Huang, Yadan Luo

Does the VGGT Family Need All Its Layers?

Read the original on arXiv Computer Vision →

The study investigates which layers of feed‑forward geometry models—specifically VGGT, π³, and VGGT‑Ω—are essential for preserving camera poses and dense 3D structure. By pruning 3,018 configurations and evaluating seven metrics across indoor and outdoor datasets, the authors identify two redundancy regions (early and late) and show that combined deletions degrade performance additively, enabling more efficient pruning. They also demonstrate that CKA can serve as a cheaper proxy for interval degradation, and that closed‑form linear calibration can recover accuracy without retraining, reducing aggregator parameters by up to 44% while maintaining comparable performance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Jun 4

Revisiting Model Stitching In the Foundation Model Era

arXiv:2603. 12433v3 Announce Type: replace-cross Abstract: Model stitching, connecting early layers of one model (source) to later layers of another (target) via a light stitch layer, has served as a probe of representational compatibility.

By Zheda Mai, Ke Zhang, Fu-En Wang, Zixiao Ken Wang, Albert Y. C. Chen, Lu Xia, Min Sun, Wei-Lun Chao, Cheng-Hao Kuo
arXiv Computer Vision
Sep 17

GeoCond: A Conditioning-Aware Reliability Adapter for Feed-Forward 3D Reconstruction

GeoCond is a lightweight reliability adapter that enhances frozen feed‑forward 3D reconstruction backbones by reading their predicted geometry to produce pose‑level uncertainty and a refinement gate. It can be trained using permutation‑orbit variance, ground‑truth pose error, or cycle residuals from unlabeled pose graphs, and at inference requires only a single backbone pass plus a small MLP. On the VGGT backbone, GeoCond reduces out‑of‑distribution AUSE from 0.32 to 0.20, transfers zero‑shot to outdoor extreme‑view scenes, and prevents collapse from uniform bundle adjustment, while also enabling gated refinement, pose‑graph weighting, calibration, curation, and capture decisions.

By David Ahmedt-Aristizabal, Mohammad Ali Armin, Russell Tsuchida, Lars Petersson