arXiv:2606. 10972v1 Announce Type: cross Abstract: This study aims to explore the performance of the VAR model in comparison with mel-frequency cepstral coefficient (MFCC) matrices and log-mel spectrograms using deep learning.
By Ipek Sen, Ozgur Ozdemir, Elena Battini Sonmez
arXiv:2606. 26723v1 Announce Type: cross Abstract: Respiratory activity is a direct and interpretable physiological channel for wearable stress and affective-state recognition, yet many studies emphasize classification accuracy without identifying which respiratory properties separate different states.
By Andrei Velichko, Mehmet Tahir Huyut
WIPSNet is a 3D ResNet that processes stacked continuous wavelet transform scalograms of overnight impedance pneumography (IP) signals to detect paediatric wheeze. In a study of 15 patients (60 nights, 281 hours), it achieved an AUC of 0.783 ± 0.026, outperforming the traditional Expiratory Variability Index, a state‑space model, and two modern sleep‑staging architectures. The model’s best performance occurs with a 32‑minute temporal context, highlighting the importance of multi‑scale temporal aggregation for nocturnal respiratory dynamics.
By Felix Oury, Harley Day, Karina Mayoral, Ville-Pekka Sepp\"a, Sejal Saglani, Reiko J. Tanaka
The study investigates respiratory signals from the WESAD dataset to detect stress and other affective states. It compares compact 1‑D CNN models trained on raw 60‑second signals with handcrafted respiratory signatures that capture timing, variability, waveform, spectral, and autocorrelation features. While the CNN achieves the highest accuracy for stress detection, the handcrafted signatures provide stronger, physiologically interpretable markers for baseline, amusement, and especially meditation states.
By Andrei Velichko, Mehmet Tahir Huyut
BreathGRU is a semi‑supervised Bidirectional Gated Recurrent Unit framework designed to segment speech and breath events in respiratory audio. It combines acoustic feature extraction, bidirectional recurrent modeling, pseudo‑label refinement, and duration‑constrained Segmental Viterbi decoding to produce accurate speech‑breath segmentation. In evaluations against existing methods, BreathGRU achieved the highest breath event recall, lowest onset‑localisation error, and highest Mean Match Intersection over Union, outperforming large pretrained VAD models such as Silero.
By Sania Fatima Sayed, John W. Holloway, Reyer Zwiggelaar, Faisal I. Rezwan
The paper presents a factorized-axis convolutional gated recurrent unit (GRU) with dynamic adaptive pooling (DAP) for predicting the remaining useful life (RUL) of rolling bearings from time‑frequency representations (TFRs). It introduces multiscale anisotropic convolution, a dual‑axis convolution block attention module, and Monte Carlo dropout for uncertainty estimation, addressing the directional structure challenges in TFRs. Experiments on two public bearing datasets show that this approach outperforms existing RUL prediction methods and that the factorized axis design and adaptive pooling contribute to lower mean errors.
By Hanbyeol Park, Jungho Choo, Hyerim Bae