arXiv Computer Vision

Rethinking the State Update Gate for Long-Sequence Recurrent 3D Reconstruction

arXiv Computer Vision
Sep 15

Anchor3R: Streaming 3D Reconstruction with Transient Anchors for Long-Horizon Visual Mapping

Anchor3R is a streaming 3D reconstruction framework that predicts window-relative poses and local geometry in the current‑frame coordinate system, forming a dense relative‑pose graph for online pose updates and loop‑aware motion averaging. It improves long‑horizon pose accuracy and dense reconstruction quality on indoor, outdoor, driving, and RGB‑D benchmarks, and generalizes from 48‑frame training sequences to streams exceeding 10,000 frames while keeping GPU memory bounded. The method addresses issues of train‑test mismatch, early‑anchor bias, and accumulated drift found in previous streaming models.

By Peilin Tao, Chong Cheng, Yuansen Du, Caiwei Song, Zhengqing Chen, Xiaoyang Guo, Wei Yin, Weiqiang Ren, Qian Zhang, Hainan Cui, Shuhan Shen
arXiv AI
Sep 21

Think Locally, Refine Globally for Memory-Efficient 3D Reconstruction

LoG-VGGT is a memory‑efficient framework for long‑sequence 3D reconstruction that balances local temporal modeling with global camera consistency. It uses cross‑window attention in a small subset of transformer blocks to propagate information across adjacent temporal windows while keeping memory usage bounded. A global camera consistency refinement module further improves long‑horizon pose stability by enforcing scene‑level constraints through cross‑attention between camera and compact register tokens, leading to better depth accuracy and robust camera pose estimation on multiple benchmarks.

By Jingke Zhou, Chenhang Ma, Zhizhou Zhong, Mingkai Liu, Zhuang Zhou, Yicheng ji, Binghua Su, Bo Cai, Xianliang Huang
Hugging Face Trending Papers
Aug 3

StreamSplat: Streaming Feed-Forward 3D Gaussian Splatting

Feed-forward 3D Gaussian Splatting enables efficient novel-view synthesis without per-scene optimization, but most existing methods assume a fixed set of context views and process them jointly. This limits their applicability to online scenarios where calibrated views arrive sequentially and the scene must be updated causally.

arXiv Computer Vision
Aug 27

LongVU-TTT: Causal Test-Time Training for Visual Resampling in Long Video Understanding

LongVU‑TTT is a causal test‑time training method for long‑video multimodal large language models that inserts a convolutional resampler with fast‑weight updates between the vision encoder and the LLM. The fast weights adapt per video and contextualize frame features before compression, while a hybrid selector keeps explicit visual evidence for downstream reasoning. Experiments show that TTT‑Conv outperforms TTT‑MLP and bidirectional Mamba2 on MLVU, and beats attention‑ and fixed‑state recurrent resamplers on three benchmarks, achieving competitive results on five video‑understanding tasks after reducing 512 frames to 128 LLM frames.

By Mahmoud Ahmed, Sameh Abdulah, Olatunji Ruwase, Sam Ade Jacobs, Mathis Bode, Mohamed Elhoseiny
arXiv Computer Vision
Sep 22

RegVGGT: Sustainable Visual Geometry Grounding for Streaming via Regulated Memory

arXiv:2609.23286v1 Announce Type: new Abstract: 3D reconstruction from a lengthy video stream input poses a dilemma for feed-forward reconstruction models (FFRMs), that a whole-stream inference conte...

By Hongbo Mao (Harbin Institute of Technology), Junjun Jiang (Harbin Institute of Technology), Youyu Chen (Harbin Institute of Technology), Jiaxin Zhang (Harbin Institute of Technology), Zhemeng Dong (Harbin Institute of Technology), Xianming Liu (Harbin Institute of Technology)
arXiv Computer Vision
1d ago

TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining

TT-VidT is a video pretraining method that decouples the temporal axis by combining a per‑frame ViT-B/16 spatial encoder with a compact Temporal Transfer Layer trained via Diff Compression. The authors conduct a systematic 24‑configuration study to isolate architecture, objective, and decoder effects, showing that the full TT-VidT design yields the strongest motion‑sensitive representations. In downstream fine‑tuning, TT‑VidT outperforms state‑of‑the‑art baselines on Jester, Something‑Something V2, ARID, and Diving48 while using significantly fewer encoder FLOPs.

By Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai