arXiv:2606. 10972v1 Announce Type: cross Abstract: This study aims to explore the performance of the VAR model in comparison with mel-frequency cepstral coefficient (MFCC) matrices and log-mel spectrograms using deep learning.
By Ipek Sen, Ozgur Ozdemir, Elena Battini Sonmez
arXiv:2606. 26723v1 Announce Type: cross Abstract: Respiratory activity is a direct and interpretable physiological channel for wearable stress and affective-state recognition, yet many studies emphasize classification accuracy without identifying which respiratory properties separate different states.
By Andrei Velichko, Mehmet Tahir Huyut
WIPSNet is a 3D ResNet that processes stacked continuous wavelet transform scalograms of overnight impedance pneumography (IP) signals to detect paediatric wheeze. In a study of 15 patients (60 nights, 281 hours), it achieved an AUC of 0.783 ± 0.026, outperforming the traditional Expiratory Variability Index, a state‑space model, and two modern sleep‑staging architectures. The model’s best performance occurs with a 32‑minute temporal context, highlighting the importance of multi‑scale temporal aggregation for nocturnal respiratory dynamics.
By Felix Oury, Harley Day, Karina Mayoral, Ville-Pekka Sepp\"a, Sejal Saglani, Reiko J. Tanaka
The study investigates respiratory signals from the WESAD dataset to detect stress and other affective states. It compares compact 1‑D CNN models trained on raw 60‑second signals with handcrafted respiratory signatures that capture timing, variability, waveform, spectral, and autocorrelation features. While the CNN achieves the highest accuracy for stress detection, the handcrafted signatures provide stronger, physiologically interpretable markers for baseline, amusement, and especially meditation states.
By Andrei Velichko, Mehmet Tahir Huyut
BreathGRU is a semi‑supervised Bidirectional Gated Recurrent Unit framework designed to segment speech and breath events in respiratory audio. It combines acoustic feature extraction, bidirectional recurrent modeling, pseudo‑label refinement, and duration‑constrained Segmental Viterbi decoding to produce accurate speech‑breath segmentation. In evaluations against existing methods, BreathGRU achieved the highest breath event recall, lowest onset‑localisation error, and highest Mean Match Intersection over Union, outperforming large pretrained VAD models such as Silero.
By Sania Fatima Sayed, John W. Holloway, Reyer Zwiggelaar, Faisal I. Rezwan
The paper presents a factorized-axis convolutional gated recurrent unit (GRU) with dynamic adaptive pooling (DAP) for predicting the remaining useful life (RUL) of rolling bearings from time‑frequency representations (TFRs). It introduces multiscale anisotropic convolution, a dual‑axis convolution block attention module, and Monte Carlo dropout for uncertainty estimation, addressing the directional structure challenges in TFRs. Experiments on two public bearing datasets show that this approach outperforms existing RUL prediction methods and that the factorized axis design and adaptive pooling contribute to lower mean errors.
By Hanbyeol Park, Jungho Choo, Hyerim Bae
Convolutional neural networks (CNN) are widely used to predict the remaining useful life (RUL) of rolling bearings from time-frequency representations (TFRs) of vibration signals. However, during degr...
arXiv:2606. 08247v1 Announce Type: cross Abstract: Acute asthma risk assessment requires rapid interpretation of respiratory sounds, oxygenation, airflow limitation, speech ability, work of breathing, mental status, and response to reliever therapy.
By Aueaphum Aueawatthanaphisut
Apnoea of prematurity is characterised by recurrent episodes of cessation of breathing and remains difficult to detect reliably using routinely monitored physiological signals in the Neonatal Intensive Care Unit (NICU). Existing bedside monitors rely primarily on respiratory rate and oxygen saturation thresholds, often generating high false-positive alarm rates and missing short or irregular events.
arXiv:2606. 11922v1 Announce Type: cross Abstract: Recent respiratory sound classification (RSC) studies largely rely on CLS-token driven self-attention architectures such as the Audio Spectrogram Transformer (AST).
By Hemansh Shridhar, Miika Toikkanen, June-Woo Kim
The paper introduces a deep learning framework that uses a CNN‑GRU architecture to classify EEG recordings into resting or cognitive states. Time‑frequency analysis extracts salient signal features, which are then evaluated with both deep learning and traditional machine learning classifiers. The proposed method achieves accuracies of 83.177% for resting vs. mathematical tasks, 76.107% for resting vs. memory tasks, and 83.432% for resting vs. music tasks, outperforming comparative approaches.
By K. A. Januka S. Fernando, Harshit Srivastava
The paper introduces a deep learning framework that uses a CNN‑GRU architecture to classify EEG recordings into resting and various cognitive states. Time‑frequency analysis is applied to extract salient signal features, which are then evaluated with both deep learning and traditional machine learning classifiers, including a proposed 2D‑Net. The method achieves accuracies of 83.177% for resting vs. mathematical tasks, 76.107% for resting vs. memory tasks, and 83.432% for resting vs. music tasks, outperforming comparative approaches.