arXiv Machine Learning

Enhanced Seam Segmentation for Automated Welding Robot in Construction Through Transfer Learning: Addressing Limitations of Bilateral Segmentation Network

arXiv:2607. 06150v1 Announce Type: cross Abstract: Reliable seam segmentation is essential for autonomous robotic welding in construction, where harsh illumination, specular reflections, and thin weld geometries often degrade segmentation performance.

arXiv Computer Vision
Aug 27

Automatic weld seam segmentation for industrial quality control: a comparison of RGB and polarimetric imaging with CNN and transformer architectures

The study evaluates automatic weld seam segmentation using RGB and polarimetric images, comparing convolutional neural networks (CNNs) and transformer architectures. In controlled RGB settings, CNNs achieve a mean mask mAP50 of up to 0.87, but performance drops significantly under uncontrolled conditions. Polarimetric imaging, combined with geometric augmentation, reaches a mean mask mAP50 of up to 093 even in uncontrolled settings, and transformer models, especially RF‑DETR, maintain high accuracy under viewpoint shifts while CNNs fail.

By Simone Garbin, Leonardo Venturoso, Marco Todescato
arXiv AI
Sep 10

Explainable Temporal Attention-based Defect Detection For Fillet Joints in Real-Time Gas Metal Arc Welding Based on Multi-modal Data

The paper presents a multi‑modal deep learning model that uses temporal attention to detect internal welding defects such as porosity, lack of penetration, fusion, undercut, and cold lap in fillet joints during real‑time Gas Metal Arc Welding. Trained on images and sound data from an industrial collaborative welding robot, the model achieves an F1 score of 0.99. Explainable AI techniques are applied to interpret the model’s behavior, highlighting key image and sound spectrogram regions and the most effective modality for each defect type, thereby enhancing trust and reliability in AI‑driven welding inspection.

By Mobina Mobaraki, Mahyar Asadi, Klaske Van Heusden, Guy A. Dumont
arXiv Computer Vision
Sep 28

PICO: Projection-Informed Consistency Optimisation for 6DoF Surgical Tool Pose Estimation

The paper introduces PICO, an end-to-end trainable model for 6DoF surgical tool pose estimation that uses multi-task learning to predict segmentation, depth, and pose parameters. It incorporates two geometry-aware proxy tasks—a projection loss and a point-to-point loss—to enforce consistency in 2D and 3D spaces, improving accuracy and robustness. Evaluated on the SurgRIPE dataset, PICO achieves strong performance, ranking second in rotation accuracy and maintaining competitive translation results, especially under occlusion.

By Lucy Fothergill, Pietro Valdastri, Dominic Jones, Duygu Sarikaya
arXiv AI
Sep 23

Tactile-JEPA: Topology-Aware Self-Supervised Representation Learning for Distributed Tactile Sensors

Tactile-JEPA is a self‑supervised pre‑training method for distributed tactile sensors that leverages the sensors’ spatial topology to learn topology‑aware representations. It predicts embeddings of masked sensing elements using a sensor connectivity graph and dual‑scale masking to capture both local contact details and the global tactile surface state. Evaluated on three diverse datasets, it improves force estimation by 6.3 % and in‑hand orientation error by 20.8 % over previous state‑of‑the‑art methods, and yields consistent gains in downstream tasks such as policy learning.

By Elizaveta Kovtun, Matvey Konovalov, Andrey Sakhovskiy, Semen Budennyy
arXiv Computer Vision
Sep 18

SenseFuse: Label-Free Fusion of Image and Shape Encoders for Open-Vocabulary 3D Instance Segmentation

SenseFuse introduces a label‑free fusion approach that balances 2D image and 3D shape encoders for open‑vocabulary 3D instance segmentation. By selecting a scene‑level fusion weight through an adaptive, sensitivity‑based mechanism, it improves mask labeling accuracy across multiple datasets, recovering up to 93% of the potential gain from an oracle weight. The method demonstrates that image and shape encoders have complementary failure patterns, leading to higher instance AP in most evaluated settings.

By Euiseok Han, Tri Ton, Hwanhee Kim, Seungyeon Ryu, Chang D. Yoo
arXiv Machine Learning
Aug 27

Lowering the Barrier to AI-Driven Inspection: A No-Code Workflow for Automated Structural Defect Detection

The paper introduces YOLOEZ, a no-code, GUI-based tool that streamlines the entire YOLO model workflow—data labeling, training, and inference—for automated structural defect detection. It demonstrates that YOLOEZ outperforms traditional image‑processing methods across most detection metrics while simplifying deployment for users without programming expertise. The tool aims to lower technical barriers in structural health monitoring, enabling broader adoption of AI-driven inspection for predictive maintenance and intelligent structural systems.

By Michael Holm, Tanner McElroy, Xinghang Zhang, Guang Lin