arXiv Machine Learning

Meta-Learning Approaches for Speaker-Dependent Voice Fatigue Models

arXiv:2505. 23378v3 Announce Type: replace Abstract: Speaker-dependent modelling can substantially improve performance in speech-based health monitoring applications.

arXiv AI
3d ago

Automatic estimation of verbal fluency index in people with Motor Neuron Disease using ASR alignment and pause modelling

The study introduces an automated system for estimating the Verbal Fluency Index (VFI) in individuals with Motor Neuron Disease (MND) by combining ASR (WhisperX) and VAD (Silero) with precise timestamping. Using a unique MND dataset, the approach outperformed traditional acoustic and self‑supervised embedding methods, achieving high predictive accuracy (R² up to 0.9 for P‑words and 0.8 for S‑words). Clinically inspired features were consistently superior, demonstrating the feasibility of automated VFI estimation for monitoring cognitive impairment in MND.

By Bahman Mirheidari, Leslie Ing, Daniel Blackburn, Sharon Abrahams, Christopher McDermott, Heidi Christensen
arXiv Computation and Language
Sep 21

Cross-Lingual Parkinson's Disease Severity Assessment Using Pre-trained Speech Embeddings: A Multi-Class Evaluation

The paper evaluates pre‑trained speech embeddings from four state‑of‑the‑art speech foundation models for cross‑lingual Parkinson's disease severity assessment. Experiments span three datasets in zero‑shot and k‑shot settings, showing that these embeddings can transfer meaningfully across languages, though performance varies with dataset characteristics, preprocessing, and adaptation strategy. Misclassifications linked to inter‑speaker variability and atypical speech patterns underscore the need for more robust feature extraction, modeling, and explainability to support reliable clinical insights.

By Simon Pals, Cristian Tejedor-Garcia
arXiv AI
Sep 18

Mitigating Stethoscope-Induced Shortcuts in Respiratory Sound Classification under Federated Domain Generalization with Causality-Inspired Interventions

The paper introduces BTS-CAFE, a federated domain generalization framework for respiratory sound classification that addresses stethoscope-induced shortcuts. It combines causality-inspired device-style interventions, counterfactual metadata augmentation, and gradient alignment to reduce style–content entanglement and promote device-invariant decision boundaries. Experiments on ICBHI and SPRSound datasets show a 3.69‑point improvement in out-of-distribution performance over the baseline and outperform conventional data augmentation and federated learning methods.

By Heejoon Koo, Yoon Tae Kim, Miika Toikkanen, June-Woo Kim