arXiv:2608.23238v1 Announce Type: new
Abstract: We present Mover360, a controllable object manipulation framework for 360{\deg} images. Unlike perspective images, 360{\deg} images in equirectangular...
By Haoyi Zhong, Fang-Lue Zhang, Andrew Chalmers, Taehyun Rhee
arXiv:2607. 00832v1 Announce Type: cross Abstract: A single panorama captures the full visual sphere from one camera center, yet confines users to looking around in place without enabling true scene exploration.
By Zhenjia Li, Jinrang Jia, Yifeng Shi
TDFNet introduces a Tri-projection Deformable Fusion Network that uses equirectangular, cube map, and tangent projections to mitigate geometric distortions in panoramic salient object detection. It incorporates a cross-projection deformable attention module for geometry-aware sampling and a latitude-guided fusion module that balances ERP and CMP features using spherical latitude priors. The network’s three-branch encoding preserves global continuity, local detail, and boundary precision, improving detection performance over existing projection-based methods.
By Qiangqiang Zhou, Jiacong Yu, Jiawei Xu, Yong Chen, Xin Huang, Ping Li
arXiv:2606. 19253v1 Announce Type: cross Abstract: Existing approaches to 3D scene understanding in Vision-Language Models (VLMs) either rely on complex, model-specific geometry encoders or large training budgets in pursuit of spatial reasoning.
By Bart{\l}omiej Baranowski, Dave Zhenyu Chen, Matthias Nie{\ss}ner
arXiv:2608.29081v1 Announce Type: new
Abstract: Spherical Transformers have emerged as a promising framework for panoramic semantic segmentation (PASS) by operating directly on spherical geometry and...
By Soumyaratna Debnath, Weiming Zhang, Shriram Damodaran, Dingwen Xiao, Addison Lin Wang
Object-Uni is a unified model that integrates pose perception, spatial reasoning, pose-conditioned generation, and novel view synthesis for object-centric spatial understanding and controllable image generation. It treats object pose as an explicit geometric variable shared across tasks and introduces a viewpoint-based orientation abstraction to make pose interpretable by multimodal large language models. The authors also create a new benchmark, UniSpatial-80K, and demonstrate that Object‑Uni improves both pose understanding and pose‑controllable generation compared to existing models.
By Mining Tan, Yinuo Wang, Ziqi Zhou, Weize Quan, Sifei Li, Jingdong Chen, DanDan Zheng, Libin Wang, Weiming Dong
arXiv:2606. 31148v1 Announce Type: cross Abstract: 3D Visual Grounding (3DVG) aims to localize target objects in 3D scenes given natural language descriptions.
By Duc Cao Dinh, Khai Le-Duc, Florent Draye, Chris Ngo, Terry Jingchen Zhang, Bernhard Sch\"olkopf, Zhijing Jin
arXiv:2601.14056v2 Announce Type: replace-cross
Abstract: Training robust visual surveillance models requires large-scale datasets with precise spatial annotations, yet collecting real surveillance d...
By Andrea Rigo, Luca Stornaiuolo, Weijie Wang, Mauro Martino, Bruno Lepri, Nicu Sebe
Driving World Models (DWMs) have recently advanced rapidly with generative models, yet most existing methods mainly focus on conditional scene generation and lack explicit 3D scene understanding, lang...
arXiv:2607. 06097v1 Announce Type: cross Abstract: 3D dense captioning, an emerging vision-language task, aims to generate descriptive sentences for each object in the 3D scene.
By Xiaopei Wu, Chenshu Hou, Liang Peng, Dan Xu, Binbin Lin, Xiaoshui Huang, Yuenan Hou, Yu Li, Wenxiao Wang, Haifeng Liu, Deng Cai, Wanli Ouyang
arXiv:2606. 03100v1 Announce Type: cross Abstract: Recently, zero-shot 3D scene understanding via 2D Vision-Language Models (VLMs) has gained increasing research interest due to their promising spatial reasoning capabilities.
By Dongsheng Wang, Dawei Su, Hui Huang
arXiv:2608.30342v1 Announce Type: new
Abstract: 3D Gaussian Splatting (3DGS) enables photorealistic real-time novel view synthesis, yet placing a virtual camera to capture a desired frame remains lar...
By Jirong Li, Satoshi Ikehata, Shuhei Kurita, Ikuro Sato