arXiv:2512. 09065v2 Announce Type: replace-cross Abstract: Many indoor workspaces are quasi-static: their global geometric layout is stable, but local semantics change continually, producing repetitive geometry, dynamic clutter, and perceptual noise that defeat standard vision-based localization.
By Shivendra Agrawal, Jake Brawer, Ashutosh Naik, Alessandro Roncone, Bradley Hayes
AnyBox is a zero‑shot framework that estimates the full 9DoF pose (6D pose plus 3D dimensions) of boxes from a single RGB‑D image, leveraging the geometric regularity of boxes. It alternates between pose and scale estimation, using a binary search guided by the discrepancy between a reprojected template and the observed mask, and employs a depth‑consistency filter and an early‑stopping rule to prune implausible hypotheses. On public benchmarks and a warehouse dataset, AnyBox improves detection AP by up to 36 points and boosts robotic box‑shelving success by 28%.
By Yintao Ma, Sajjad Pakdamansavoji, Charles Eret, Rui Heng Yang, Xuan Zhao, Yingxue Zhang, Tongtong Cao, Amir Rasouli
arXiv:2512.01352v2 Announce Type: replace
Abstract: Unsupervised and open-vocabulary 3D object detection have recently gained attention, particularly in autonomous driving, where reducing annotation...
By In-Jae Lee, Mungyeom Kim, Kwonyoung Ryu, Pierre Musacchio, Jaesik Park
GenCOPE introduces a synthetic-to-real (Syn2Real) approach for category-level object pose estimation (COPE) that eliminates the need for labor-intensive real-world data collection. By learning domain-invariant representations through 2D and 3D semantic consistency constraints and employing an end-to-end pose regression framework with 2D-3D cross consistency, the model achieves robust generalization across synthetic and real domains. The architecture relies solely on global features, resulting in a lightweight and efficient design validated on REAL275, Wild6D, and real-world robotic manipulation scenes.
By Jian Liu, Wei Sun, Zhenqi Dai, Hui Yang, Jian Xiao, Nicu Sebe, Na Zhao
arXiv:2609.25654v1 Announce Type: cross
Abstract: Robots operating safely in cluttered everyday environments often need to infer scene geometry from partial observations. Methods that detect objects...
By Dongwon Son, Junhyek Han, Yoontae Cho, Minseok Lee, Hong-seok Choi, Jiwook Choi, Hyungjin Kim, Beomjoon Kim
arXiv:2607. 02921v1 Announce Type: cross Abstract: Quantitative 3D spatial reasoning from egocentric RGB-D video is a critical capability for next-generation wearable assistants.
By Maxwell Horton, Wei Lu, Quan Tran, Yury Astashonok, Kirmani Ahmed, Babak Damavandi, Anuj Kumar, Xiao Zhang, Seungwhan Moon
Robots operating safely in cluttered everyday environments often need to infer scene geometry from partial observations. Methods that detect objects in 2D and reconstruct them independently struggle i...
arXiv:2610.03013v1 Announce Type: new
Abstract: Category-level object pose estimation predicts the rotation, translation, and metric size of unseen instances within known categories. Many accurate RG...
By Hakjin Lee, Junghoon Seo, Jaehoon Sim
arXiv:2608. 19973v1 Announce Type: cross Abstract: Recently, open-vocabulary 3D object detection (3D-OVD) has gained increasing attention for its ability to detect unseen objects in 3D scenes.
By Shangbo Yuan, Jie Xu, Xiaofeng Zhu, Na Zhao
arXiv:2607. 19036v1 Announce Type: cross Abstract: V2X collaborative object detection features overcoming the limitations of single-vehicle systems by aggregating environmental features from multiple collaborative agents.
By Zhihao Yang, Zhiyu Xiang, Peng Xu, Tianyu Pu, Kai Wang, Eryun Liu, Dongping Zhang, Yong Ding
arXiv:2607. 17778v1 Announce Type: cross Abstract: Class-agnostic 3D instance segmentation is critical for robotic systems operating in unknown environments, enabling perception of previously unseen objects for reliable manipulation and navigation.
By Juno Kim, Hye-Jung Yoon, Yesol Park, Byoung-Tak Zhang
Physical AI Smart Spaces is the first benchmark that offers large‑scale, multi‑class, multi‑camera 3D perception data for indoor smart spaces. It includes over 280 hours of synchronized 1080p footage from nearly 1,800 cameras in warehouses, hospitals, and retail venues, with automatic annotations for identities, 2D and 3D bounding boxes, camera calibration, and depth. The benchmark spans synthetic generation, appearance augmentation, and real‑world Sim2Real evaluation, and introduces a 3D version of Higher Order Tracking Accuracy (HOTA) for evaluating multi‑class 3D box tracking.
By Yuxing Wang, Yizhou Wang, Anqi Li, Shuo Wang, Sameer Satish Pusegaonkar, Haoquan Liang, Jiajun Li, Shenxin Jiang, Jianhe Yuan, Shangru Li, Tongwei Dai, Zihao Chen, David C. Anastasiu, Sujit Biswas, Xunlei Wu, Zheng Tang