The paper introduces PLANET, a multi‑object tracker that transcends traditional image‑plane limitations by incorporating 3D scene geometry into its query formation. By lifting 2D tracking datasets into 3D and embedding reconstructed geometry into features and positional encodings, PLANET encourages queries to encode object positions. An auxiliary 3D location prediction task and a dual‑resolution temporal memory further enhance performance, enabling state‑of‑the‑art results on three diverse benchmarks.
By Orcun Cetintas, Guillem Bras\'o, Tim Meinhardt, Laura Leal-Taix\'e
Synthesizing a novel-view video from a monocular reference video along a target camera trajectory requires both geometric consistency and motion fidelity with respect to the reference video. Existing methods based on explicit 3D representations are limited by the accuracy of off-the-shelf reconstruction modules, which often produce inaccurate geometry for dynamic objects in monocular videos.
arXiv:2609.18363v1 Announce Type: new
Abstract: Online multi camera 3D tracking must maintain scene global identities across synchronized views, yet query-based trackers carry these identities only i...
By Pragyan Shrestha, Haruto Nakayama, Atom Scott
arXiv:2606. 30576v1 Announce Type: cross Abstract: Cross-view object geo-localization (CVOGL) aims to locate a target object from a query view (e.
By Liyao Wang, Ruipu Wu, Haojun Xu, Lei Shi, Linjiang Huang, Si Liu
arXiv:2607. 17778v1 Announce Type: cross Abstract: Class-agnostic 3D instance segmentation is critical for robotic systems operating in unknown environments, enabling perception of previously unseen objects for reliable manipulation and navigation.
By Juno Kim, Hye-Jung Yoon, Yesol Park, Byoung-Tak Zhang
Class-agnostic 3D instance segmentation is critical for robotic systems operating in unknown environments, enabling perception of previously unseen objects for reliable manipulation and navigation. Existing approaches typically project per-frame 2D instance masks into 3D and merge them, which often breaks object identities across time and yields fragmented 3D instances.
MAETrack introduces a lightweight framework to adapt pretrained masked autoencoder (MAE) representations for 3D single object tracking (SOT). It uses Layer‑Selective Initialization (LSI) to keep shallow geometric layers from the pre‑training while re‑initializing deeper layers, and Geometric Residual Gating (GRG) to emphasize salient regions in BEV features before template‑search fusion. Experiments on standard 3D SOT benchmarks show consistent improvements over vanilla fine‑tuning with minimal computational cost.
By Sifan Zhou, Qiwei Wang, Linyue Tan, Ziyu Liu, Ziyu Zhao, Xiaobo Lu
Open vocabulary 3D scene understanding is essential for next-generation interactive systems, empowering users to intuitively query and navigate reconstructed environments using natural language. However, current 3D Gaussian frameworks are often bottlenecked by restrictive multiview capture requirements, costly scene-specific optimization, and the massive memory overhead of storing dense language features.
TQD-Track introduces Temporal Query Denoising (TQD) for 3D Multi‑Object Tracking, extending query denoising beyond single frames by initializing denoising queries from previous‑frame ground truths and propagating them as independent association candidates. The method enriches track queries with temporal context and instance‑specific features, and incorporates diverse noise types to emulate real‑world tracking challenges. Experiments on nuScenes and Argoverse 2 show consistent improvements across multiple MOT baselines with only training‑process modifications.
By Yutong Yang, Shuxiao Ding, Mohammed Amine Bencheikh Lehocine, Julian Wiederer, Markus Braun, Peizheng Li, Juergen Gall, Bin Yang
Real-world spatial intelligence requires agents to understand scenes from continuous video streams, where objects move, persist, disappear, and reappear over time. While recent spatial foundation models have enabled generalizable feed-forward 3D reconstruction, most streaming methods remain geometry-centric and lack temporally consistent object-level understanding.
Image-goal visual navigation is a fundamental capability for embodied agents. Existing navigation policies efficiently predict waypoint trajectories but lack visual foresight, while navigation world models can anticipate future observations but often require costly planning rollouts.
arXiv:2606. 19733v1 Announce Type: cross Abstract: Efficiently retrieving specific 3D instances from large-scale scenes via natural language prompts remains a formidable challenge in multimedia analysis.
By Xiuyuan Zhu, Ke Lu, Zijie Yang, Chao Yue, Jian Xue, Dongming Zhang