GeoWAM introduces a visual geometry world action model that predicts future scene geometry instead of future images, using point clouds to capture spatial structure and transformations. The model is pretrained to forecast geometry, then a geometry-conditioned action head predicts ego trajectories. Experiments show that this geometry-based approach yields stronger driving policies than image-based alternatives.
By Yiren Lu, Xin Ye, Jiaming Liu, Jin Yao, Yi-chung Chen, Liam Merino, Dhruva Dixith Kurra, Min Cai, Tom Lampo, Yu Yin, Danhua Guo, Burhan Yaman
The paper proposes a new sequence-to-sequence formulation for multi-view stereo (MVS) that jointly predicts 3D geometry for all input views using a global transformer architecture. It introduces ray‑map embeddings to inject camera parameters into image tokens and a unified global cost volume to capture 3D structure across all views. Experiments on public benchmarks demonstrate state‑of‑the‑art performance, outperforming both traditional MVS and feed‑forward reconstruction baselines.
By Aoxiang Fan, Corentin Dumery, Nicolas Talabot, Pascal Fua
arXiv:2512. 02473v2 Announce Type: replace-cross Abstract: Video world models have attracted significant attention for their ability to produce high-fidelity future visual observations conditioned on past observations and navigation actions.
By Yuta Oshima, Yusuke Iwasawa, Masahiro Suzuki, Yutaka Matsuo, Hiroki Furuta
DiffWAM is a geometry‑conditioned navigation world‑action model that transforms predictive features from a frozen video foundation model into continuous camera trajectories, eliminating the need for future‑video synthesis and multi‑frame reconstruction during deployment. Its Grid‑Motion module preserves spatial‑temporal motion associations, while Latent2Pose grounds them with first‑frame geometry to recover metrically meaningful 3D motion. The system, complemented by FastDreamer for asynchronous trajectory handoff, achieves a trajectory RMSE of 0.3492 m and a 74.40 % endpoint success rate on the DiffWAM‑1000 benchmark, with real‑world tests showing complex UAV behaviors and an onboard implementation reaching 1.08 s latency on NVIDIA Jetson AGX Thor.
By Mo Zhu, Yuze Wu, Xijie Huang, Xiao Cui, Fei Gao, Xin Zhou
arXiv:2512. 16919v2 Announce Type: replace-cross Abstract: Perceiving and reconstructing 3D scene geometry from visual inputs is crucial for autonomous driving.
By Sicheng Zuo, Zixun Xie, Wenzhao Zheng, Shaoqing Xu, Fang Li, Shengyin Jiang, Long Chen, Zhi-Xin Yang, Jiwen Lu
arXiv:2608.27529v1 Announce Type: new
Abstract: Streaming 3D reconstruction from extremely long videos requires estimating camera motion and scene geometry online under bounded memory and computation...
By Jiarong Han, Jincheng Xiong, Yuzhou Liu, Linzhe Shi, Changjie Wu, Ning Guo, Mu Xu, Hang Zhang, Ming Qian
Long-horizon future-frame prediction is important for autonomous driving, traffic surveillance, and intelligent transportation systems, yet remains challenging due to temporal ghosting, geometry drift, and inconsistent object motion. Recent latent video diffusion models have achieved impressive visual quality, but directly applying them to structured traffic scenes often leads to unstable geometry and degraded temporal coherence over extended horizons.
arXiv:2606.18250v2 Announce Type: replace
Abstract: Forecasting the evolution of dynamic environments is crucial for autonomous agents. While generative world models have achieved high photorealism i...
By Nils Morbitzer, Jonathan Evers, Artem Savkin, Thomas Stauner, Nassir Navab, Federico Tombari, Stefano Gasperini
Anchor3R is a streaming 3D reconstruction framework that predicts window-relative poses and local geometry in the current‑frame coordinate system, forming a dense relative‑pose graph for online pose updates and loop‑aware motion averaging. It improves long‑horizon pose accuracy and dense reconstruction quality on indoor, outdoor, driving, and RGB‑D benchmarks, and generalizes from 48‑frame training sequences to streams exceeding 10,000 frames while keeping GPU memory bounded. The method addresses issues of train‑test mismatch, early‑anchor bias, and accumulated drift found in previous streaming models.
By Peilin Tao, Chong Cheng, Yuansen Du, Caiwei Song, Zhengqing Chen, Xiaoyang Guo, Wei Yin, Weiqiang Ren, Qian Zhang, Hainan Cui, Shuhan Shen
PhysWAM is a unified world-action model for autonomous driving that jointly denoises multiview video, metric depth, and ego motion using a flow‑matching transformer. It introduces Coupled Point Projection (CPP), a geometric constraint that aligns generated depth points with LiDAR data after applying the predicted SE(3) ego motion, thereby enforcing physical consistency. At inference, trajectory selection uses a simple label‑free consensus rule, and the model demonstrates strong planning performance, zero‑shot transfer to unseen environments, and accurate, temporally coherent depth and video predictions.
By Dhruv Parikh, Fengcheng Yu, Quankai Gao, Jiawei Yang, Junjie Ye, Maulik Bhatt, Thang Vu, Charles Ochoa, Rowan McAllister, Igor Vasiljevic, Rajgopal Kannan, Viktor Prasanna, Vitor Guizilini, Yue Wang
The paper introduces GeoNeXt, a framework that repurposes pretrained video generative models for geometry estimation by framing it as a next‑frame prediction task. Unlike prior methods that either train separate depth/normal models or fine‑tune image diffusion backbones, GeoNeXt jointly models images and geometric targets, leveraging the structured knowledge of video models for more data‑efficient learning. Experiments show zero‑shot monocular depth and surface normal estimation that outperforms existing generative approaches and rivals discriminative state‑of‑the‑art methods while using far less training data.
By Haosen Yang, Jifei Song, Zhensong Zhang, Xiatian Zhu, Jiankang Deng
Image-goal visual navigation is a fundamental capability for embodied agents. Existing navigation policies efficiently predict waypoint trajectories but lack visual foresight, while navigation world models can anticipate future observations but often require costly planning rollouts.