arXiv:2609.23733v1 Announce Type: new
Abstract: Feed-forward visual geometry models such as the Visual Geometry Grounded Transformer (VGGT) have recently enabled direct 3D reconstruction from multi-v...
By Abteen Arab, Guile Wu, Chengjie Huang, Dongfeng Bai
arXiv:2603. 12433v3 Announce Type: replace-cross Abstract: Model stitching, connecting early layers of one model (source) to later layers of another (target) via a light stitch layer, has served as a probe of representational compatibility.
By Zheda Mai, Ke Zhang, Fu-En Wang, Zixiao Ken Wang, Albert Y. C. Chen, Lu Xia, Min Sun, Wei-Lun Chao, Cheng-Hao Kuo
GeoCond is a lightweight reliability adapter that enhances frozen feed‑forward 3D reconstruction backbones by reading their predicted geometry to produce pose‑level uncertainty and a refinement gate. It can be trained using permutation‑orbit variance, ground‑truth pose error, or cycle residuals from unlabeled pose graphs, and at inference requires only a single backbone pass plus a small MLP. On the VGGT backbone, GeoCond reduces out‑of‑distribution AUSE from 0.32 to 0.20, transfers zero‑shot to outdoor extreme‑view scenes, and prevents collapse from uniform bundle adjustment, while also enabling gated refinement, pose‑graph weighting, calibration, curation, and capture decisions.
By David Ahmedt-Aristizabal, Mohammad Ali Armin, Russell Tsuchida, Lars Petersson
arXiv:2607. 03784v1 Announce Type: cross Abstract: While prior studies have successfully compressed vision Transformers (ViTs) through various pruning techniques, most have concentrated on width pruning to achieve significant reductions in model size.
By Zhenfeng Su, Kang Zhao, Han Bao, Tao Yuan, Zhongzhe Hu, Xianzhi Yu, Wenxuan Wang
Feed-forward 3D foundation models such as VGGT predict cameras, depth, and point maps in a single pass, but can fail silently under low overlap, low parallax, and extreme relative rotation. Stratified...
arXiv:2608.21008v1 Announce Type: new
Abstract: Mobile AR frameworks attach a metric pose prior to every casual phone capture, and turning it into reconstruction-grade poses cheaply on CPU is the ste...
By Nikolaos Kyriazis
arXiv:2605.08371v2 Announce Type: replace
Abstract: Multi-view geometry transformers are feed-forward 3D foundation models that jointly predict depth maps, point maps, and camera poses for N images i...
By Haotang Li, Zhenyu Qi, Shaohan Henry Wang, Kebin Peng, Zi Wang, Qing Guo, Sen He, Huanrui Yang
arXiv:2608.29680v1 Announce Type: new
Abstract: Feed-forward 3D foundation models reconstruct perspective scenes in one pass. Satellite photogrammetry needs a different product, one that domain adapt...
By Zhe Dong, Wanqing Wu, Yuzhe Sun, Haochen Jiang, Yuchen Ma, Lecheng Ren, Tianzhu Liu, Yanfeng Gu
arXiv:2608. 08904v1 Announce Type: cross Abstract: How much of a vision-language model's (VLM) spatial understanding remains after the action post-training process of building a vision-language-action model (VLA)?
By Alexander Hackett, Arnaud Denis-Remillard, Axel Cassou
arXiv:2609.37013v1 Announce Type: cross
Abstract: Rapid assessment of building damage after natural disasters is essential to support emergency response. Earth Observation satellites can acquire rele...
By Thomas Goudemant, Benjamin Francesconi, Marjorie Bellizzi, Adrien Dorise
arXiv:2606.21596v2 Announce Type: replace
Abstract: Recent image-to-3D scene methods recover high-fidelity 3D objects with plausible arrangements, but often leave floatings and interpenetrations that...
By Haodong Li, Lulu Shao, Haolin Lu, Yu Fu, Yen-Ru Chen, Seemandhar Jain, Manmohan Chandraker
Syn2RealTrack addresses the synthetic‑to‑real gap in multi‑camera 3D perception for warehouses by decomposing it into three distinct issues: camera calibration, object shape prior, and known object census. The pipeline corrects lens distortion from images, fuses detections with a visibility‑weighted part‑based descriptor, measures person height directly from calibration, and uses a closed‑world cardinality prior with a causal filter to eliminate phantom boxes. These local remedies allow the system to adapt without retraining a feature extractor, achieving a 3D HOTA of 52.0118% on the AI City Challenge 2026 Track 1.
By Duong Nguyen-Ngoc Tran, Ngoc Doan-Minh Huynh, Cu Quoc Le, Hoang-Khang Nguyen, Long Hoang Pham, Huy-Hung Nguyen, Quoc Pham-Nam Ho, Trinh Le Ba Khanh, Chi Dai Tran, Duong Khac Vu, Son Hong Phan, Hyung-Min Jeon, Jae Wook Jeon