Hugging Face Trending Papers

ReCal3R: Reliability-Calibrated Learning Rates for Streaming 3D Reconstruction

Read the original on Hugging Face Trending Papers →

Streaming 3D reconstruction relies on a compact recurrent scene state to process long image streams in linear time and bounded memory. However, repeated updates can gradually corrupt this state, causing reliable historical information to be overwritten by noisy or ambiguous observations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Sep 10

FILT3R: Latent State Adaptive Kalman Filter for Streaming 3D Reconstruction

FILT3R is a training‑free latent filtering layer for streaming 3D reconstruction that treats recurrent state updates as stochastic state estimation in token space. It maintains per‑token variance and computes a Kalman‑style gain to balance memory retention with new observations, estimating process noise online from temporal drift of candidate tokens. Experiments show that FILT3R generalizes overwrite and gating policies, shrinking gains in stable regimes and increasing them during genuine scene changes, thereby improving long‑horizon stability for depth, pose, and 3D reconstruction.

By Seonghyun Jin, Jong Chul Ye
arXiv AI
Sep 21

Think Locally, Refine Globally for Memory-Efficient 3D Reconstruction

LoG-VGGT is a memory‑efficient framework for long‑sequence 3D reconstruction that balances local temporal modeling with global camera consistency. It uses cross‑window attention in a small subset of transformer blocks to propagate information across adjacent temporal windows while keeping memory usage bounded. A global camera consistency refinement module further improves long‑horizon pose stability by enforcing scene‑level constraints through cross‑attention between camera and compact register tokens, leading to better depth accuracy and robust camera pose estimation on multiple benchmarks.

By Jingke Zhou, Chenhang Ma, Zhizhou Zhong, Mingkai Liu, Zhuang Zhou, Yicheng ji, Binghua Su, Bo Cai, Xianliang Huang
arXiv Computer Vision
Sep 15

Anchor3R: Streaming 3D Reconstruction with Transient Anchors for Long-Horizon Visual Mapping

Anchor3R is a streaming 3D reconstruction framework that predicts window-relative poses and local geometry in the current‑frame coordinate system, forming a dense relative‑pose graph for online pose updates and loop‑aware motion averaging. It improves long‑horizon pose accuracy and dense reconstruction quality on indoor, outdoor, driving, and RGB‑D benchmarks, and generalizes from 48‑frame training sequences to streams exceeding 10,000 frames while keeping GPU memory bounded. The method addresses issues of train‑test mismatch, early‑anchor bias, and accumulated drift found in previous streaming models.

By Peilin Tao, Chong Cheng, Yuansen Du, Caiwei Song, Zhengqing Chen, Xiaoyang Guo, Wei Yin, Weiqiang Ren, Qian Zhang, Hainan Cui, Shuhan Shen