arXiv AI

Improving Engine Sound Analysis in Hot-Test Environments via a RAB-U-Net (Residual Attention Block U-Net) Noise Removal Method

arXiv:2606. 21887v2 Announce Type: replace-cross Abstract: During hot tests on a production line, engine-sound analysis is crucial to ensuring product quality and performance.

arXiv AI
Sep 10

Explainable Temporal Attention-based Defect Detection For Fillet Joints in Real-Time Gas Metal Arc Welding Based on Multi-modal Data

The paper presents a multi‑modal deep learning model that uses temporal attention to detect internal welding defects such as porosity, lack of penetration, fusion, undercut, and cold lap in fillet joints during real‑time Gas Metal Arc Welding. Trained on images and sound data from an industrial collaborative welding robot, the model achieves an F1 score of 0.99. Explainable AI techniques are applied to interpret the model’s behavior, highlighting key image and sound spectrogram regions and the most effective modality for each defect type, thereby enhancing trust and reliability in AI‑driven welding inspection.

By Mobina Mobaraki, Mahyar Asadi, Klaske Van Heusden, Guy A. Dumont
arXiv Machine Learning
Aug 17

Deep Vision in Smart Manufacturing: MODERN Framework for Intelligent Quality Monitoring and Diagnosis

arXiv:2608. 13937v1 Announce Type: cross Abstract: Smart manufacturing processes are often installed with a large number of sensors, imaging devices and computers, which not only enable instant communication across various modules of a production system but also aid in intelligent manufacturing management.

By Yicheng Kang, Yuling Jiao, Xin Geng, Mahesh Nagarajan
arXiv Machine Learning
Sep 24

"What's That Sound?": A Versatile, Robust, and Lightweight Convolutional Transformer for Environment Sound Recognition

The paper introduces RALCT, a lightweight Convolutional Transformer that combines randomized audio augmentations, MFCCs, and log‑mel spectrograms to extract robust features for environmental sound recognition. With only about 310,000 parameters, RALCT achieves state‑of‑the‑art accuracy—over 93% on UrbanSound8K, peaking at 94.56%—making it suitable for deployment on mobile devices. The authors also develop a mobile app that integrates the model to provide real‑time safety alerts for hearing‑impaired users.

By Julia Huang
arXiv Machine Learning
Aug 24

Training DeepFilterNet with Accurate Room Acoustic Simulations Improves Single-Channel Speech Enhancement

The study examines how the realism of synthetic room impulse response (RIR) datasets influences the training of DeepFilterNet3 for single‑channel speech enhancement. By comparing a DNS4 image‑source‑method RIR set with a higher‑fidelity hybrid wave‑based and geometrical acoustics RIR set, the authors find that the more realistic dataset consistently improves objective speech enhancement metrics and significantly reduces ASR word error rates on unseen measured RIRs. The results suggest that overall realism in synthetic acoustic training data enhances DeepFilterNet3’s generalization to new environments.

By Alessia Milo, Georg G\"otz, Steinar Gu{\dh}j\'onsson, Daniel Gert Nielsen, Jesper Pedersen, Finnur Pind
arXiv Machine Learning
3d ago

AFA-Net: A Differential Attention Approach for Auditory Attention Detection

AFA‑Net introduces a differential attention mechanism to improve Auditory Attention Detection (AAD) from EEG signals, explicitly targeting noisy data. The framework outperforms existing deep learning models, achieving 96.8% accuracy within a 2‑second decision window while using fewer parameters. It represents one of the first approaches to actively mitigate EEG noise in AAD tasks.

By Philip H. Lee, Shreeram Suresh Chandra, Karan Thakkar, John H. L. Hansen