arXiv:2606. 26260v1 Announce Type: cross Abstract: In laser penetration welding, the assessment of penetration state and weld seam morphology plays a crucial role in determining the weld quality.
By Sen Li, Haichao Cui, Chendong Shao, Yaqi Wang, Xinhua Tang
The study evaluates automatic weld seam segmentation using RGB and polarimetric images, comparing convolutional neural networks (CNNs) and transformer architectures. In controlled RGB settings, CNNs achieve a mean mask mAP50 of up to 0.87, but performance drops significantly under uncontrolled conditions. Polarimetric imaging, combined with geometric augmentation, reaches a mean mask mAP50 of up to 093 even in uncontrolled settings, and transformer models, especially RF‑DETR, maintain high accuracy under viewpoint shifts while CNNs fail.
By Simone Garbin, Leonardo Venturoso, Marco Todescato
The paper presents a multi‑modal deep learning model that uses temporal attention to detect internal welding defects such as porosity, lack of penetration, fusion, undercut, and cold lap in fillet joints during real‑time Gas Metal Arc Welding. Trained on images and sound data from an industrial collaborative welding robot, the model achieves an F1 score of 0.99. Explainable AI techniques are applied to interpret the model’s behavior, highlighting key image and sound spectrogram regions and the most effective modality for each defect type, thereby enhancing trust and reliability in AI‑driven welding inspection.
By Mobina Mobaraki, Mahyar Asadi, Klaske Van Heusden, Guy A. Dumont
arXiv:2608.29475v1 Announce Type: new
Abstract: Surface material recognition from incomplete visual observations remains a challenging problem in robotic perception and environmental understanding. T...
By Sindhuja Penchala, Sudip Mittal, Noorbakhsh Amiri Golilarz
arXiv:2609.39738v1 Announce Type: new
Abstract: Learning to generate machining process plans and toolpaths from B-rep CAD requires coupling discrete operation decisions with continuous tool motion as...
By Xiaolei Zhou, Boyi Lin, Yuchao Feng, Jianwei Zheng
arXiv:2609.36844v1 Announce Type: new
Abstract: Transparent surfaces are ubiquitous in built environments, yet they remain a persistent failure case for robotic perception. RGB cameras perceive the b...
By Suhani Grover, Astik Srivastava, Viswas Dinesh, Avinash Sharma, K. Madhava Krishna
arXiv:2607. 05568v1 Announce Type: cross Abstract: Representing 3D shapes as compact sets of geometric primitives is fundamental to robotics, simulation, and scene understanding.
By Gregor Kobsik, Tim Elsner, Leif Kobbelt
arXiv:2607.05568v2 Announce Type: replace-cross
Abstract: Compact primitive abstractions represent 3D shapes with a few geometric primitives while preserving recognizable components. Learned methods...
By Gregor Kobsik, Tim Elsner, Leif Kobbelt
The paper introduces PICO, an end-to-end trainable model for 6DoF surgical tool pose estimation that uses multi-task learning to predict segmentation, depth, and pose parameters. It incorporates two geometry-aware proxy tasks—a projection loss and a point-to-point loss—to enforce consistency in 2D and 3D spaces, improving accuracy and robustness. Evaluated on the SurgRIPE dataset, PICO achieves strong performance, ranking second in rotation accuracy and maintaining competitive translation results, especially under occlusion.
By Lucy Fothergill, Pietro Valdastri, Dominic Jones, Duygu Sarikaya
Tactile-JEPA is a self‑supervised pre‑training method for distributed tactile sensors that leverages the sensors’ spatial topology to learn topology‑aware representations. It predicts embeddings of masked sensing elements using a sensor connectivity graph and dual‑scale masking to capture both local contact details and the global tactile surface state. Evaluated on three diverse datasets, it improves force estimation by 6.3 % and in‑hand orientation error by 20.8 % over previous state‑of‑the‑art methods, and yields consistent gains in downstream tasks such as policy learning.
By Elizaveta Kovtun, Matvey Konovalov, Andrey Sakhovskiy, Semen Budennyy
SenseFuse introduces a label‑free fusion approach that balances 2D image and 3D shape encoders for open‑vocabulary 3D instance segmentation. By selecting a scene‑level fusion weight through an adaptive, sensitivity‑based mechanism, it improves mask labeling accuracy across multiple datasets, recovering up to 93% of the potential gain from an oracle weight. The method demonstrates that image and shape encoders have complementary failure patterns, leading to higher instance AP in most evaluated settings.
By Euiseok Han, Tri Ton, Hwanhee Kim, Seungyeon Ryu, Chang D. Yoo
The paper introduces YOLOEZ, a no-code, GUI-based tool that streamlines the entire YOLO model workflow—data labeling, training, and inference—for automated structural defect detection. It demonstrates that YOLOEZ outperforms traditional image‑processing methods across most detection metrics while simplifying deployment for users without programming expertise. The tool aims to lower technical barriers in structural health monitoring, enabling broader adoption of AI-driven inspection for predictive maintenance and intelligent structural systems.
By Michael Holm, Tanner McElroy, Xinghang Zhang, Guang Lin