arXiv Machine Learning

BenSParX: A Robust Explainable Machine Learning Framework for Parkinson's Disease Detection from Bengali Conversational Speech

BenSparX introduces the first Bengali conversational speech dataset for Parkinson’s disease detection and pairs it with a robust, explainable machine learning framework. The framework uses diverse acoustic features, systematic feature selection, and advanced classifiers, achieving 95.67% accuracy, 95.62% F1, and 0.990 AUC. SHAP analysis is employed to explain feature contributions, and the model outperforms state‑of‑the‑art methods on other language datasets.

arXiv Computation and Language
2d ago

Cross-Lingual Parkinson's Disease Severity Assessment Using Pre-trained Speech Embeddings: A Multi-Class Evaluation

The paper evaluates pre‑trained speech embeddings from four state‑of‑the‑art speech foundation models for cross‑lingual Parkinson's disease severity assessment. Experiments span three datasets in zero‑shot and k‑shot settings, showing that these embeddings can transfer meaningfully across languages, though performance varies with dataset characteristics, preprocessing, and adaptation strategy. Misclassifications linked to inter‑speaker variability and atypical speech patterns underscore the need for more robust feature extraction, modeling, and explainability to support reliable clinical insights.

By Simon Pals, Cristian Tejedor-Garcia
arXiv AI
Jul 21

A Benchmark for Early-stage Parkinson's Disease Detection from Speech

arXiv:2605. 14066v2 Announce Type: replace-cross Abstract: Early-stage Parkinson's disease (EarlyPD) detection from speech is clinically meaningful yet underexplored, and published results are hard to compare because studies differ in datasets, languages, tasks, evaluation protocols, and EarlyPD definitions.

By Terry Yi Zhong, Cristian Tejedor-Garcia, Khiet P. Truong, Janna Maas, Louis ten Bosch, Bastiaan R. Bloem
arXiv AI
Aug 25

Multi-Task Learning for Non-Canonical Phoneme Recognition via Articulatory Feature Decomposition

The paper proposes a linguistically structured multi‑task learning framework for recognizing non‑canonical phonemes by decomposing phoneme prediction into articulatory feature dimensions such as manner, place, and voicing. A hierarchical architecture with task‑specific heads and a cross‑attention fusion module is combined with semi‑supervised Momentum Pseudo‑Labeling and a cascaded training strategy that gradually introduces articulatory tasks. Experiments on the L2‑ARCTIC dataset demonstrate significant improvements over baseline models and produce interpretable error patterns aligned with phonological feature structure.

By Sophia Riaz, Haoze Zheng, Amos Roche, Miyu Zhang, Anamika Ragu, Salvatore Penachio, Kaustav Mukherjee, Aneesh Jonelagadda