arXiv:2609.18034v1 Announce Type: new
Abstract: Novel view synthesis from unposed multi-view images remains challenging, as the model must jointly learn scene representations and camera parameters wi...
By Wenyu Li, Sidun Liu, Peng Qiao, Yong Dou, Tongrui Hu
The paper introduces a biologically inspired framework that learns object‑centric visual representations from raw videos without human annotations or camera calibration. By using motion boundaries detected via optical flow and clustering to create pseudo‑instance masks, the method supervises a single‑image encoder with pixel‑level pairwise metric learning. Training on 195 million pseudo‑labeled frames and expanding to 421 million frames through Motion‑Verified Self‑Training, the approach yields Swin‑based encoders that outperform or match supervised and self‑supervised baselines on tasks such as monocular depth estimation, 3D object detection, 3D occupancy prediction, and end‑to‑end planning.
By Boshi Li, Xiaohui Wang, Xiaoyang Wu, Zhichao Li, Ya Yang, Naiyan Wang
DINOcular is a self‑supervised framework that learns joint visuospatial representations from RGB‑D observations. It fuses depth‑derived geometric priors with a visual backbone using inter‑patch and intra‑patch fusion, allowing the model to encode both appearance and spatial structure efficiently. The resulting representation improves 3D awareness on multiple geometry benchmarks while staying competitive on standard RGB‑D semantic segmentation tasks.
By Farkhat Almukhamedov, Sami Azirar, Hermann Blum
DINOcular is a self‑supervised framework that learns joint visuospatial representations from RGB‑D data. It fuses depth‑derived geometric priors with a visual backbone using inter‑patch and intra‑patch techniques, allowing the model to encode both appearance and spatial structure efficiently. The resulting representation improves 3D awareness on multiple geometry benchmarks while staying competitive on standard RGB‑D semantic segmentation tasks.
arXiv:2609.39590v1 Announce Type: new
Abstract: Compositional 3D scene generation aims to recover complete 3D object shapes and their spatial arrangement from visual observations. Recent image-condit...
By Guibiao Liao, Mochu Xiang, Heng Li, Ken Deng, Zijie Wang, Guanbin Li, Ping Tan, Shenghua Gao, Yizhou Yu
Real-world spatial intelligence requires agents to understand scenes from continuous video streams, where objects move, persist, disappear, and reappear over time. While recent spatial foundation models have enabled generalizable feed-forward 3D reconstruction, most streaming methods remain geometry-centric and lack temporally consistent object-level understanding.