arXiv Computer Vision

Matisse: Evidence-Space Reasoning for Active 3D Reconstruction

Matisse is a training‑free framework that combines active 3D reconstruction with keyframe selection by using evidence from a pretrained generative 3D model. It estimates evidential uncertainty via cross‑attention on 3D latent tokens and derives an evidential information gain to guide view acquisition and keyframe selection, reducing redundant observations and supporting multi‑object scenes with occlusion‑aware aggregation. On GSO30, YCB‑V, and Replica, Matisse improves Chamfer distance by 12.7%, 3.8%, and 9.2% respectively, and speeds up end‑to‑end reconstruction by 1.5× compared to the best baseline.

arXiv AI
Aug 3

AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance

arXiv:2512. 05131v2 Announce Type: replace-cross Abstract: Active 3D reconstruction enables an agent to autonomously select viewpoints to efficiently obtain accurate and complete scene geometry, rather than passively reconstructing scenes from pre-collected images.

By Tianling Xu, Shengzhe Gan, Leslie Gu, Yuelei Li, Fangneng Zhan, Hanspeter Pfister
arXiv Computer Vision
Sep 25

One View Is Enough: In-the-Wild Monocular Pretraining for Novel View Generation

The paper introduces OVIE, a monocular novel-view synthesis method that eliminates the need for multi‑view training data. By using a frozen depth estimator to generate pseudo‑target views from single images and applying masked and adversarial losses, OVIE is trained on 30 million uncurated images. It achieves state‑of‑the‑art performance on RealEstate10K and DL3DV, produces highly consistent multi‑view trajectories, and runs at 116 FPS—over 600× faster than the fastest baseline.

By Adrien Ramanana Rahary, Nicolas Dufour, Patrick Perez, David Picard