arXiv Computer Vision

ShelfChange3D: Object-Level 3D Change Detection for Retail Shelf Monitoring

arXiv AI
Jun 2

ShelfAware: Real-Time Semantic Localization in Quasi-Static Environments with Low-Cost Sensors

arXiv:2512. 09065v2 Announce Type: replace-cross Abstract: Many indoor workspaces are quasi-static: their global geometric layout is stable, but local semantics change continually, producing repetitive geometry, dynamic clutter, and perceptual noise that defeat standard vision-based localization.

By Shivendra Agrawal, Jake Brawer, Ashutosh Naik, Alessandro Roncone, Bradley Hayes
arXiv AI
Sep 4

AnyBox: Efficient Zero-Shot 9DoF Pose Estimation of Boxes for Robotic Manipulation

AnyBox is a zero‑shot framework that estimates the full 9DoF pose (6D pose plus 3D dimensions) of boxes from a single RGB‑D image, leveraging the geometric regularity of boxes. It alternates between pose and scale estimation, using a binary search guided by the discrepancy between a reprojected template and the observed mask, and employs a depth‑consistency filter and an early‑stopping rule to prune implausible hypotheses. On public benchmarks and a warehouse dataset, AnyBox improves detection AP by up to 36 points and boosts robotic box‑shelving success by 28%.

By Yintao Ma, Sajjad Pakdamansavoji, Charles Eret, Rui Heng Yang, Xuan Zhao, Yingxue Zhang, Tongtong Cao, Amir Rasouli
arXiv Computer Vision
4d ago

GenCOPE: Syn2Real Generalized Category-Level Object Pose Estimation for Robotic Picking

GenCOPE introduces a synthetic-to-real (Syn2Real) approach for category-level object pose estimation (COPE) that eliminates the need for labor-intensive real-world data collection. By learning domain-invariant representations through 2D and 3D semantic consistency constraints and employing an end-to-end pose regression framework with 2D-3D cross consistency, the model achieves robust generalization across synthetic and real domains. The architecture relies solely on global features, resulting in a lightweight and efficient design validated on REAL275, Wild6D, and real-world robotic manipulation scenes.

By Jian Liu, Wei Sun, Zhenqi Dai, Hui Yang, Jian Xiao, Nicu Sebe, Na Zhao
arXiv Computer Vision
1d ago

Physical AI Smart Spaces: A Large-Scale Benchmark for Multi-Camera 3D Perception in Smart Spaces

Physical AI Smart Spaces is the first benchmark that offers large‑scale, multi‑class, multi‑camera 3D perception data for indoor smart spaces. It includes over 280 hours of synchronized 1080p footage from nearly 1,800 cameras in warehouses, hospitals, and retail venues, with automatic annotations for identities, 2D and 3D bounding boxes, camera calibration, and depth. The benchmark spans synthetic generation, appearance augmentation, and real‑world Sim2Real evaluation, and introduces a 3D version of Higher Order Tracking Accuracy (HOTA) for evaluating multi‑class 3D box tracking.

By Yuxing Wang, Yizhou Wang, Anqi Li, Shuo Wang, Sameer Satish Pusegaonkar, Haoquan Liang, Jiajun Li, Shenxin Jiang, Jianhe Yuan, Shangru Li, Tongwei Dai, Zihao Chen, David C. Anastasiu, Sujit Biswas, Xunlei Wu, Zheng Tang