arXiv AI

MoPe: Motion Permanence for Robust Monocular Gaussian Mapping in Dynamic Environments

arXiv:2606. 29237v1 Announce Type: cross Abstract: Robust robot autonomy depends on scene representations that remain stable enough to support localization, navigation, and downstream decision making in dynamic environments.

arXiv Computer Vision
6d ago

Reliability-Regulated Trajectory Optimization for Progressive COLMAP-Free 3D Gaussian Splatting

The paper introduces a reliability-regulated trajectory optimization framework for progressive COLMAP‑free 3D Gaussian Splatting (3DGS). It uses a self‑supervised bidirectional cycle‑consistency mechanism to control camera trajectory estimation through forward motion propagation and retrospective trajectory correction, thereby reducing error compounding without external priors. Experiments on Tanks and Temples and CO3D‑V2 demonstrate improved camera trajectory accuracy and novel‑view rendering quality compared to existing unposed baselines.

By Zijian Wu, Jinliang Wang, Zidian Lin, Ying Song, Ziqian Lu, Hanjie Ma, Zhen Ye, Mingfeng Jiang
Hugging Face Trending Papers
Jul 23

GLAM-SLAM: Real-time Gaussian Large-scale Mapping via Flow Densification and Spatial Decomposition

Existing Gaussian-splatting-based monocular Simultaneous Localization and Mapping (SLAM) systems are either tailored to short sequences, are not real-time, or suffer from prohibitive GPU memory requirements, limiting their applicability in realistic, long-horizon scenarios. To address this, we present GLAM-SLAM, a real-time, decoupled Gaussian-splatting SLAM system designed for large-scale outdoor scenes.

Hugging Face Trending Papers
Sep 17

EliGSiR: Continual RGB-D Mapping with Gaussian Splatting under Bounded Compute

EliGSiR is a continual RGB‑D mapping system that extends Gaussian splatting to handle online, bounded‑compute scenarios. It introduces Map‑Guided View Scheduling to filter redundant views, Load‑Adaptive Fidelity to adjust supervision resolution, and Targeted Geometry Growth to add structure only where needed. Experiments on Replica, TUM RGB‑D, ScanNet++ and real sensor data show that EliGSiR outperforms baselines in reconstruction quality while efficiently using the available mapping budget.

arXiv Computer Vision
Sep 18

EliGSiR: Continual RGB-D Mapping with Gaussian Splatting under Bounded Compute

EliGSiR is a continual RGB‑D mapping method that extends 3D Gaussian Splatting to online settings by adaptively allocating optimization resources. It introduces Map‑Guided View Scheduling to filter redundant views, Load‑Adaptive Fidelity to adjust supervision resolution, and Targeted Geometry Growth to add geometric capacity only where needed. The approach is evaluated on Replica, TUM RGB‑D, ScanNet++, and real sensor sequences, achieving higher reconstruction quality and faster performance than several baselines.

By Bj\"orn Ellensohn, Elmar Rueckert, Christian Rauch
arXiv AI
Aug 20

GS-VLA: Plug-and-Play Viewpoint Canonicalization for Frozen VLA Policies via Gaussian Splatting

GS‑VLA introduces a lightweight, plug‑and‑play framework that uses a 4 M‑parameter 3D‑Gaussian canonicalizer to adapt frozen Vision‑Language‑Action (VLA) policies to viewpoint shifts without retraining the policy. By treating viewpoint changes as a localized novel‑view synthesis problem under a locality assumption, the method normalizes observations through a scene‑ and policy‑independent disocclusion task. Experiments on the LIBERO benchmark demonstrate that GS‑VLA recovers a large portion of performance lost due to camera displacement, improving results across different policy architectures, unseen task suites, and perturbation scales. whyItMatters":"The approach offers a computationally efficient alternative to costly fine‑tuning or generative augmentation, enabling robust VLA deployment in real‑world settings where camera configurations may vary."

By Yechan Park, HyunJin Kim
arXiv Computer Vision
Sep 21

Adaptive World Memory 3D Foundation Model for Scalable 3D Mapping, Localization, and Rendering

The paper introduces Adaptive World Memory 3D Foundation Model (AWM-3DFM), a memory‑centric 3D foundation model that scales to large‑scale robotic localization, reconstruction, and Gaussian rendering. It employs transformer‑based gated updates, test‑time temporal‑spatial regulation, and local submap organization to maintain persistent memory, accuracy, and consistency across long image sequences. A Gaussian reconstruction head unifies pose estimation, dense point‑cloud reconstruction, and photorealistic rendering, achieving superior trajectory accuracy, reconstruction completeness, and rendering quality on public benchmarks and diverse robotic datasets.

By Tianchen Deng, Guole Shen, Yilin Shen, Wenhua Wu, Yilin Fang, Ziqi Ma, Tianjun Zhang, Shenghai Yuan, Wolfram Burgard, Hesheng Wang
arXiv Computer Vision
Sep 24

Know-Your-Scene (KYS)-SLAM: Hierarchical Semantic-Motion Priors for Feature Matching in Stereo Visual SLAM

Know-Your-Scene (KYS)-SLAM extends ORB‑SLAM3 by replacing binary feature rejection with continuous correspondence modulation based on semantic, panoptic, and motion priors. Each keypoint is augmented with hierarchical compatibility scores that down‑weight features on independently moving objects while preserving static structure, using a training‑free depth‑aware ego‑motion model and self‑calibrating thresholds. Across 21 stereo sequences, KYS‑SLAM achieves a 17.4% ATE RMSE reduction on outdoor KITTI, 27.7% on indoor EuRoC, and significant improvements on dynamic and synthetic datasets without per‑sequence tuning.

By Preeti Chatterjee, Jin Lu, Jin Sun, Suchendra M. Bhandarkar