Streaming 3D reconstruction relies on a compact recurrent scene state to process long image streams in linear time and bounded memory. However, repeated updates can gradually corrupt this state, causing reliable historical information to be overwritten by noisy or ambiguous observations.
arXiv:2609.21938v1 Announce Type: new
Abstract: Transformer-based models have recently achieved strong performance on 3D reconstruction from images, and recent works extend them to process video stre...
By Sunghyun Baek, Hanna Bae, Minchan Kwon, Junmo Kim
arXiv:2608.27529v1 Announce Type: new
Abstract: Streaming 3D reconstruction from extremely long videos requires estimating camera motion and scene geometry online under bounded memory and computation...
By Jiarong Han, Jincheng Xiong, Yuzhou Liu, Linzhe Shi, Changjie Wu, Ning Guo, Mu Xu, Hang Zhang, Ming Qian
arXiv:2605.16981v3 Announce Type: replace
Abstract: Streaming 3D reconstruction under a strict constant-memory budget hinges on how the recurrent state is updated as the stream evolves. We profile TT...
By Kejun Ren, Lei Jin, Tianxin Huang, Lianming Xu, Li Wang
arXiv:2606. 12691v1 Announce Type: cross Abstract: Auto-regressive models have emerged as powerful tools for sequential data, from language to video.
By Yahya Sattar, Sunmook Choi, Leo Maynard-Zhang, Yassir Jedra, Maryam Fazel, Sarah Dean
The paper introduces AURA, a meta‑learning framework that learns a low‑dimensional latent state‑space model for the evolution of optimal model parameters under distribution shift. Online adaptation is performed via extended Kalman filtering in this latent space, followed by reconstruction of full model parameters through a learned lifting map, enabling efficient single‑step updates. Experiments on neural wireless receivers and non‑stationary image classification show that AURA improves adaptation speed, accuracy, and computational efficiency compared to existing online learning and Bayesian filtering baselines.
By Guy Gerson, Tomer Raviv, Nir Shlezinger, Tirza Routtenberg, Osvaldo Simeone
Feed-forward 3D Gaussian Splatting enables efficient novel-view synthesis without per-scene optimization, but most existing methods assume a fixed set of context views and process them jointly. This limits their applicability to online scenarios where calibrated views arrive sequentially and the scene must be updated causally.
arXiv:2608.29029v3 Announce Type: replace-cross
Abstract: Joint-Embedding Predictive Architectures (JEPAs) provide a powerful framework for latent world modeling and planning in a reconstruction-free...
By Yanchen Huo, Ziying Song, Yadan Luo
arXiv:2609.17230v1 Announce Type: new
Abstract: Streaming 3D reconstruction demands both speed and temporal fidelity, goals that existing methods undermine by updating every Gaussian every frame, eve...
By Idil Sulo, Alexey Supikov, Ilke Demir, Sainan Liu
arXiv:2608. 19556v1 Announce Type: cross Abstract: Streaming autoregressive diffusion models enable real-time, long-horizon video generation, but their training objectives optimize local frame prediction rather than the geometry and dynamics of a coherent world: long rollouts accumulate geometric drift and degrade into static or unnatural motion.
By Yuanhao Ban, Jiaqi Feng, Hengguang Zhou, Xiaohuan Pei, Justin Cui, Cho-Jui Hsieh
LayerRecall is a memory router for autoregressive video diffusion that selectively retrieves and injects historical key/value states into specific layers of the model, based on the current context. It addresses the problem that existing memory mechanisms expose nonlocal history but do not guarantee effective use, by recognizing that different layers prefer current, recent, or distant context. The method, combined with Cross‑Horizon Prediction Matching, achieves state‑of‑the‑art long‑range consistency on MemoBench and MovieBench while maintaining local continuity and incurring negligible inference overhead.
By Yixuan Ding, Jiahao Kong, Wei Huang, Ruijie Quan, Yi Yang
arXiv:2609.23286v1 Announce Type: new
Abstract: 3D reconstruction from a lengthy video stream input poses a dilemma for feed-forward reconstruction models (FFRMs), that a whole-stream inference conte...
By Hongbo Mao (Harbin Institute of Technology), Junjun Jiang (Harbin Institute of Technology), Youyu Chen (Harbin Institute of Technology), Jiaxin Zhang (Harbin Institute of Technology), Zhemeng Dong (Harbin Institute of Technology), Xianming Liu (Harbin Institute of Technology)