Feed-forward 3D Gaussian Splatting enables efficient novel-view synthesis without per-scene optimization, but most existing methods assume a fixed set of context views and process them jointly. This limits their applicability to online scenarios where calibrated views arrive sequentially and the scene must be updated causally.
arXiv:2605.16981v3 Announce Type: replace
Abstract: Streaming 3D reconstruction under a strict constant-memory budget hinges on how the recurrent state is updated as the stream evolves. We profile TT...
By Kejun Ren, Lei Jin, Tianxin Huang, Lianming Xu, Li Wang
arXiv:2609.23601v1 Announce Type: new
Abstract: Long-video understanding must capture transient visual evidence under strict token budgets, yet existing methods compress frames, append memory tokens,...
By Siru Zhong, Qiongyan Wang, Xiaohui Lv, Yuzheng Zhuang, Shuai Tao, Wulong Liu, Haohuan Fu, Yuxuan Liang
arXiv:2510. 09608v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) could power real-time assistants and autonomous agents, but they face a critical challenge: understanding near-infinite video streams without escalating latency and memory usage.
By Ruyi Xu, Guangxuan Xiao, Yukang Chen, Liuning He, Yao Lu, Song Han
arXiv:2609.36929v1 Announce Type: new
Abstract: Recent event-based depth estimation methods successfully transfer geometric priors from vision foundation models via cross-modal distillation. However,...
By Thai Duy Nguyen, Addison Lin Wang
arXiv:2609.37042v1 Announce Type: cross
Abstract: Video Large Language Models (VideoLLMs) have achieved strong video understanding capabilities but incur substantial inference overhead due to the lar...
By Shuo Yang, Changbai Li, Rui Tang, Xinyu Zhao, Linlin Yang, Baochang Zhang
arXiv:2609.17230v1 Announce Type: new
Abstract: Streaming 3D reconstruction demands both speed and temporal fidelity, goals that existing methods undermine by updating every Gaussian every frame, eve...
By Idil Sulo, Alexey Supikov, Ilke Demir, Sainan Liu
arXiv:2605. 16366v2 Announce Type: replace-cross Abstract: Video MLLMs face a persistent tension between spatial fidelity and temporal coverage: preserving fine-grained visual details requires many spatial tokens, while capturing short-lived events requires dense temporal sampling.
By Yigui Feng (The College of Computer Science, National University of Defense Technology, Changsha, Hunan, China), Qinglin Wang (The College of Computer Science, National University of Defense Technology, Changsha, Hunan, China), Yang Liu (The Shien-Ming Wu School of Intelligent Engineering, South China University of Technology, Guangzhou, Guangdong, China), Jie Liu (The College of Computer Science, National University of Defense Technology, Changsha, Hunan, China)
arXiv:2608.27529v1 Announce Type: new
Abstract: Streaming 3D reconstruction from extremely long videos requires estimating camera motion and scene geometry online under bounded memory and computation...
By Jiarong Han, Jincheng Xiong, Yuzhou Liu, Linzhe Shi, Changjie Wu, Ning Guo, Mu Xu, Hang Zhang, Ming Qian
TempoGround is a vision‑language model–native framework for streaming visual grounding that detects cross‑frame object correspondence and explicitly models object presence states. It uses a curriculum prediction mechanism to resolve 2D instance association, predict object entry, continuation, or exit, decode 2D boxes, and lift them to 3D camera‑frame boxes. The approach is further refined with Streaming Grounding Reinforcement, which optimizes grounding, identity, and consistency rewards, and achieves significant improvements on multiple streaming visual grounding benchmarks.
By Leqian Ding, Junning Qiu, Manwen Yang, Yu Guo, Fei Wang
arXiv:2606. 26762v1 Announce Type: cross Abstract: Streaming video understanding (SVU) must answer queries that arrive asynchronously while visual tokens stream continuously under strict GPU-memory and query-time latency budgets.
By Le Tu Ngoc Minh (KAIST), Jinyeong Lim (KAIST), Dongsu Han (KAIST)
Camera intrinsics are vital for recovering 3D structure from 2D video. However, most 3D algorithms assume fixed intrinsics throughout a video, an assumption that often fails for real-world in-the-wild videos.