LoG-VGGT is a memory‑efficient framework for long‑sequence 3D reconstruction that balances local temporal modeling with global camera consistency. It uses cross‑window attention in a small subset of transformer blocks to propagate information across adjacent temporal windows while keeping memory usage bounded. A global camera consistency refinement module further improves long‑horizon pose stability by enforcing scene‑level constraints through cross‑attention between camera and compact register tokens, leading to better depth accuracy and robust camera pose estimation on multiple benchmarks.
By Jingke Zhou, Chenhang Ma, Zhizhou Zhong, Mingkai Liu, Zhuang Zhou, Yicheng ji, Binghua Su, Bo Cai, Xianliang Huang
arXiv:2604. 05182v2 Announce Type: replace-cross Abstract: We introduce the Large Sparse Reconstruction Model to study how scaling transformer context windows affects feed-forward 3D reconstruction.
By Zhengqin Li, Cheng Zhang, Jakob Engel, Zhao Dong
arXiv:2609.23796v2 Announce Type: replace
Abstract: Single-image 3D object generation can now produce high-fidelity assets, yet accurately placing them into a coherent scene layout remains an open ch...
By Yang-Tian Sun, Tianjia Liu, Zehuan Huang, Yi-Hua Huang, Xiaoyang Lyu, Ziyi Yang, Zi-Xin Zou, Yuan-Chen Guo, Yan-Pei Cao, Xiaojuan Qi
arXiv:2609.23796v1 Announce Type: new
Abstract: Single-image 3D object generation can now produce high-fidelity assets, yet accurately placing them into a coherent scene layout remains an open challe...
By Yang-Tian Sun, Tianjia Liu, Zehuan Huang, Yi-Hua Huang, Xiaoyang Lyu, Ziyi Yang, Zi-Xin Zou, Yuan-Chen Guo, Yan-Pei Cao, Xiaojuan Qi
arXiv:2607. 03612v1 Announce Type: cross Abstract: Feed-forward 3D reconstruction (F3R) transformers have recently achieved remarkable success.
By Jianing Deng, Yuanzhe Li, Jialu Wang, Song Wang, Tianlong Chen, Huanrui Yang, Jingtong Hu
The paper proposes a new sequence-to-sequence formulation for multi-view stereo (MVS) that jointly predicts 3D geometry for all input views using a global transformer architecture. It introduces ray‑map embeddings to inject camera parameters into image tokens and a unified global cost volume to capture 3D structure across all views. Experiments on public benchmarks demonstrate state‑of‑the‑art performance, outperforming both traditional MVS and feed‑forward reconstruction baselines.
By Aoxiang Fan, Corentin Dumery, Nicolas Talabot, Pascal Fua