The paper presents a multi‑modal deep learning model that uses temporal attention to detect internal welding defects such as porosity, lack of penetration, fusion, undercut, and cold lap in fillet joints during real‑time Gas Metal Arc Welding. Trained on images and sound data from an industrial collaborative welding robot, the model achieves an F1 score of 0.99. Explainable AI techniques are applied to interpret the model’s behavior, highlighting key image and sound spectrogram regions and the most effective modality for each defect type, thereby enhancing trust and reliability in AI‑driven welding inspection.
By Mobina Mobaraki, Mahyar Asadi, Klaske Van Heusden, Guy A. Dumont
arXiv:2503. 00340v2 Announce Type: cross Abstract: Lightweight models are essential for real-time speech enhancement applications.
By Xiaobin Rong, Leyan Yang, Dahan Wang, Yuxiang Hu, Changbao Zhu, Kai Chen, Jing Lu
arXiv:2608. 13937v1 Announce Type: cross Abstract: Smart manufacturing processes are often installed with a large number of sensors, imaging devices and computers, which not only enable instant communication across various modules of a production system but also aid in intelligent manufacturing management.
By Yicheng Kang, Yuling Jiao, Xin Geng, Mahesh Nagarajan
The paper introduces RALCT, a lightweight Convolutional Transformer that combines randomized audio augmentations, MFCCs, and log‑mel spectrograms to extract robust features for environmental sound recognition. With only about 310,000 parameters, RALCT achieves state‑of‑the‑art accuracy—over 93% on UrbanSound8K, peaking at 94.56%—making it suitable for deployment on mobile devices. The authors also develop a mobile app that integrates the model to provide real‑time safety alerts for hearing‑impaired users.
By Julia Huang
arXiv:2511. 21325v2 Announce Type: replace-cross Abstract: Deepfake (DF) audio detectors still struggle to generalize to out of distribution inputs.
By Ido Nitzan Hidekel, Gal lifshitz, Khen Cohen, Dan Raviv
The study examines how the realism of synthetic room impulse response (RIR) datasets influences the training of DeepFilterNet3 for single‑channel speech enhancement. By comparing a DNS4 image‑source‑method RIR set with a higher‑fidelity hybrid wave‑based and geometrical acoustics RIR set, the authors find that the more realistic dataset consistently improves objective speech enhancement metrics and significantly reduces ASR word error rates on unseen measured RIRs. The results suggest that overall realism in synthetic acoustic training data enhances DeepFilterNet3’s generalization to new environments.
By Alessia Milo, Georg G\"otz, Steinar Gu{\dh}j\'onsson, Daniel Gert Nielsen, Jesper Pedersen, Finnur Pind
AFA‑Net introduces a differential attention mechanism to improve Auditory Attention Detection (AAD) from EEG signals, explicitly targeting noisy data. The framework outperforms existing deep learning models, achieving 96.8% accuracy within a 2‑second decision window while using fewer parameters. It represents one of the first approaches to actively mitigate EEG noise in AAD tasks.
By Philip H. Lee, Shreeram Suresh Chandra, Karan Thakkar, John H. L. Hansen
Smart manufacturing processes are often installed with a large number of sensors, imaging devices and computers, which not only enable instant communication across various modules of a production syst...
arXiv:2608. 14287v1 Announce Type: cross Abstract: Passive acoustic sensing offers a critical, cost-efficient, and, crucially, passive alternative for detecting small unmanned aerial vehicles.
By Vadym Vilhurin, Volodymyr Sydorskyi, Andrii Shevtsov
arXiv:2602. 20967v2 Announce Type: replace-cross Abstract: Automatic speech recognition (ASR) degrades severely in noisy environments.
By Haoyang Li, Changsong Liu, Wei Rao, Hao Shi, Sakriani Sakti, Eng Siong Chng
arXiv:2607. 13330v1 Announce Type: cross Abstract: Diffusion-based text-to-audio generative models such as AudioLDM achieve high perceptual quality and strong semantic consistency; however, their practical deployment is hindered by the substantial computational cost of the U-Net denoising backbone.
By Arshdeep Singh, Yi Yuan, Yun Chen, Wenwu Wang, Mark D. Plumbley
arXiv:2606. 00852v1 Announce Type: cross Abstract: Printed circuit board (PCB) defect detection is challenging because many defects are small and difficult to distinguish from complex background patterns.
By Vinay Edula, Nilesh Badwe, Priyanka Bagade