arXiv AI

Leveraging Deep Learning for Object and Position Recognition of Load Carriers for Autonomous Logistics Vehicles

arXiv:2606. 16042v1 Announce Type: cross Abstract: This work explores the use of artificial intelligence in mobile robotics to achieve autonomous detection and pose estimation of load carriers for automated pickup.

arXiv Computer Vision
Oct 2

GenCOPE: Syn2Real Generalized Category-Level Object Pose Estimation for Robotic Picking

GenCOPE introduces a synthetic-to-real (Syn2Real) approach for category-level object pose estimation (COPE) that eliminates the need for labor-intensive real-world data collection. By learning domain-invariant representations through 2D and 3D semantic consistency constraints and employing an end-to-end pose regression framework with 2D-3D cross consistency, the model achieves robust generalization across synthetic and real domains. The architecture relies solely on global features, resulting in a lightweight and efficient design validated on REAL275, Wild6D, and real-world robotic manipulation scenes.

By Jian Liu, Wei Sun, Zhenqi Dai, Hui Yang, Jian Xiao, Nicu Sebe, Na Zhao
arXiv Computer Vision
Sep 4

An Ensemble-Based Self-Taught Learning Approach for Parking Space Classification Under Limited Data

The paper proposes an ensemble-based self‑taught learning framework for parking space classification that uses unsupervised convolutional autoencoders to learn transferable visual representations from unlabeled data. These learned encoders serve as fixed feature extractors for supervised classification with limited annotated samples, and an ensemble of heterogeneous autoencoders with independent classifier heads is employed to enhance robustness and reduce architectural bias. Experiments on PKLot and CNRPark benchmarks demonstrate that this approach significantly lowers annotation requirements while achieving high accuracies (93–96%) under cross‑dataset evaluation protocols.

By Lucas de Oliveira Cunha, Joelton Deonei Gotz, Paulo Lisboa de Almeida, Andre Gustavo Hochuli
arXiv AI
Jun 30

Tactile Gesture Recognition with Built-in Joint Sensors for Industrial Robots

arXiv:2508. 12435v2 Announce Type: replace-cross Abstract: While gesture recognition using vision or robot skins is an active research area in Human-Robot Collaboration (HRC), this paper explores deep learning methods relying solely on a robot's built-in joint sensors, eliminating the need for external sensors.

By Deqing Song, Weimin Yang, Maryam Rezayati, Hans Wernher van de Venn
arXiv Computer Vision
Aug 31

SUFLECA: Scaling Up Feature Learning for CAD-to-image Alignment

SUFLECA is a weakly supervised framework that improves zero‑shot CAD‑to‑image alignment by scaling geometry‑grounded feature learning using Normalized Object Coordinates across up to 12 real and synthetic datasets. It introduces a geometrically consistent matching algorithm that reliably establishes CAD‑to‑image correspondences, enabling accurate, sub‑second alignment without iterative pose refinement. On the ScanNet25k benchmark, SUFLECA achieves 32.8%/42.6% category/instance accuracy, outperforming the strongest zero‑shot baseline by 9.7/12.5 percentage points and surpassing existing pose‑supervised methods for the first time.

By Saad Ejaz, Miguel Fernandez-Cortizas, Javier Civera, Holger Voos, Jose Luis Sanchez-Lopez
arXiv AI
Jul 7

MemPose: Category-level Object Pose Estimation with Memory

arXiv:2607. 04930v1 Announce Type: cross Abstract: In the pursuit of robust and generalizable category-level object pose estimation, most existing methods adopt parametric formulations that learn effective representations from data, yet they primarily encode category-level patterns into fixed shape priors or static parameter weights, which limits their scalability to highly diverse instances.

By Xiao Lin, Minghao Zhu, Yun Peng, Liuyi Wang, Qiyi Wang, Chengju Liu, Qijun Chen
arXiv AI
2d ago

Supermarket Product Detection and Recognition: Utilizing Deep Learning with Rectified Imagery

The paper investigates how rectifying supermarket product images using homography estimation and the Hough transform can improve deep learning-based object detection. It evaluates the impact of angle variation and object density on detection accuracy, highlighting both benefits and limitations of image rectification. The authors advocate for a new dataset to further study these effects.

By Mayank Sah, Jimson Mathew