arXiv AI

Optimizing 2D Input Representations and Sub-phase Fusion Strategies for Differential Diagnosis of Asthma and COPD Using CNN- and GRU-Based Networks

arXiv:2606. 10972v1 Announce Type: cross Abstract: This study aims to explore the performance of the VAR model in comparison with mel-frequency cepstral coefficient (MFCC) matrices and log-mel spectrograms using deep learning.

Hugging Face Trending Papers
Jun 9

Optimizing 2D Input Representations and Sub-phase Fusion Strategies for Differential Diagnosis of Asthma and COPD Using CNN- and GRU-Based Networks

This study aims to explore the performance of the VAR model in comparison with mel-frequency cepstral coefficient (MFCC) matrices and log-mel spectrograms using deep learning. In pulmonary sound classification, spectrogram-based representations suffer from inconsistent temporal dimensions due to varying respiratory cycle durations.

arXiv Machine Learning
Jun 26

State-Specific Respiratory Signatures for Affective and Stress Recognition: Interpretable Respiratory Markers, Autocorrelation Lags, and Compact CNN Models

arXiv:2606. 26723v1 Announce Type: cross Abstract: Respiratory activity is a direct and interpretable physiological channel for wearable stress and affective-state recognition, yet many studies emphasize classification accuracy without identifying which respiratory properties separate different states.

By Andrei Velichko, Mehmet Tahir Huyut
arXiv Machine Learning
5d ago

BreathGRU: A Novel Semi-Supervised Bidirectional Gated Recurrent Unit Framework for Speech and Breath Segmentation for Respiratory Audio

BreathGRU is a semi‑supervised Bidirectional Gated Recurrent Unit framework designed to segment speech and breath events in respiratory audio. It combines acoustic feature extraction, bidirectional recurrent modeling, pseudo‑label refinement, and duration‑constrained Segmental Viterbi decoding to produce accurate speech‑breath segmentation. In evaluations against existing methods, BreathGRU achieved the highest breath event recall, lowest onset‑localisation error, and highest Mean Match Intersection over Union, outperforming large pretrained VAD models such as Silero.

By Sania Fatima Sayed, John W. Holloway, Reyer Zwiggelaar, Faisal I. Rezwan
arXiv Machine Learning
1d ago

WIPSNet: Deep Learning for Paediatric Wheeze Detection from Overnight Impedance Pneumography

WIPSNet is a 3D ResNet that processes stacked continuous wavelet transform scalograms of overnight impedance pneumography (IP) signals to detect paediatric wheeze. In a study of 15 patients (60 nights, 281 hours), it achieved an AUC of 0.783 ± 0.026, outperforming the traditional Expiratory Variability Index, a state‑space model, and two modern sleep‑staging architectures. The model’s best performance occurs with a 32‑minute temporal context, highlighting the importance of multi‑scale temporal aggregation for nocturnal respiratory dynamics.

By Felix Oury, Harley Day, Karina Mayoral, Ville-Pekka Sepp\"a, Sejal Saglani, Reiko J. Tanaka
arXiv Machine Learning
Sep 14

State-specific respiratory signatures for affective and stress recognition: Interpretable respiratory markers, autocorrelation lags, and compact CNN models

The study investigates respiratory signals from the WESAD dataset to detect stress and other affective states. It compares compact 1‑D CNN models trained on raw 60‑second signals with handcrafted respiratory signatures that capture timing, variability, waveform, spectral, and autocorrelation features. While the CNN achieves the highest accuracy for stress detection, the handcrafted signatures provide stronger, physiologically interpretable markers for baseline, amusement, and especially meditation states.

By Andrei Velichko, Mehmet Tahir Huyut
arXiv AI
Jun 9

AeroSpectra Sentinel: An Auditable LLM Prompt-Chaining Decision-Support Workflow for Acute Asthma Risk Assessment from Respiratory Sounds and Clinical Signals

arXiv:2606. 08247v1 Announce Type: cross Abstract: Acute asthma risk assessment requires rapid interpretation of respiratory sounds, oxygenation, airflow limitation, speech ability, work of breathing, mental status, and response to reliever therapy.

By Aueaphum Aueawatthanaphisut
arXiv AI
6d ago

Factorized axis convolutional gated recurrent unit with dynamic adaptive pooling for remaining useful life prediction of rolling bearings

The paper presents a factorized-axis convolutional gated recurrent unit (GRU) with dynamic adaptive pooling (DAP) for predicting the remaining useful life (RUL) of rolling bearings from time‑frequency representations (TFRs). It introduces multiscale anisotropic convolution, a dual‑axis convolution block attention module, and Monte Carlo dropout for uncertainty estimation, addressing the directional structure challenges in TFRs. Experiments on two public bearing datasets show that this approach outperforms existing RUL prediction methods and that the factorized axis design and adaptive pooling contribute to lower mean errors.

By Hanbyeol Park, Jungho Choo, Hyerim Bae
Hugging Face Trending Papers
Jun 22

Deep learning-based detection of cessation of breathing in pre-term infants

Apnoea of prematurity is characterised by recurrent episodes of cessation of breathing and remains difficult to detect reliably using routinely monitored physiological signals in the Neonatal Intensive Care Unit (NICU). Existing bedside monitors rely primarily on respiratory rate and oxygen saturation thresholds, often generating high false-positive alarm rates and missing short or irregular events.

arXiv AI
Sep 10

AudioFuse: Unified Spectral-Temporal Learning via a Hybrid ViT-1D CNN Architecture for Robust Phonocardiogram Classification

AudioFuse is a hybrid architecture that jointly learns from spectrograms and raw waveforms to classify phonocardiograms. It combines a wide-and-shallow Vision Transformer for spectral features with a shallow 1D CNN for temporal waveforms, reducing overfitting while capturing complementary information. On the PhysioNet 2016 dataset, AudioFuse achieves a state‑of‑the‑art ROC‑AUC of 0.8608 and shows superior robustness to domain shift on the PASCAL dataset, outperforming both spectrogram‑only and waveform‑only baselines.

By Md. Saiful Bari Siddiqui, Utsab Saha