Hugging Face Trending Papers

Structured Phonological Representations for Audio-Articulatory rtMRI Speech Classification

Real-time MRI makes it possible to observe vocal-tract articulation during speech, but mapping these articulatory patterns to phonetic and phonological categories remains challenging. We investigate whether PhonoQ, an audio-based model trained to recognize structured phonological features, provides useful information for audio--articulatory modeling.

arXiv AI
Aug 25

Multi-Task Learning for Non-Canonical Phoneme Recognition via Articulatory Feature Decomposition

The paper proposes a linguistically structured multi‑task learning framework for recognizing non‑canonical phonemes by decomposing phoneme prediction into articulatory feature dimensions such as manner, place, and voicing. A hierarchical architecture with task‑specific heads and a cross‑attention fusion module is combined with semi‑supervised Momentum Pseudo‑Labeling and a cascaded training strategy that gradually introduces articulatory tasks. Experiments on the L2‑ARCTIC dataset demonstrate significant improvements over baseline models and produce interpretable error patterns aligned with phonological feature structure.

By Sophia Riaz, Haoze Zheng, Amos Roche, Miyu Zhang, Anamika Ragu, Salvatore Penachio, Kaustav Mukherjee, Aneesh Jonelagadda
arXiv AI
Sep 25

Anatomy-aware cross-speaker adaptation of complete vocal-tract acoustic-to-articulatory inversion

The paper introduces a geometric adaptation framework for cross‑speaker acoustic‑to‑articulatory inversion that leverages anatomical landmarks on vertebrae and dental structures. By applying an affine transformation followed by a thin‑plate spline deformation, the method maps predicted vocal‑tract contours from a fixed model to unseen speakers without retraining. Experiments on a single‑speaker rt‑MRI database and eight additional speakers show that the combined affine‑plus‑TPS approach with 12 and 14 landmarks yields the lowest mean point‑to‑closest‑point error of 3.19 mm.

By Nhat-Nam Nguyen, Pierre-Andre Vuissoz, Yves Laprie
arXiv AI
Sep 2

Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models

The paper investigates whether audio‑language models capture paralinguistic cues beyond spoken content. Using the Expresso dataset and four open‑source models, the authors trace how speaking style information is encoded in the late layers of the audio encoder but is degraded before reaching the final output. They find that some models are content‑driven while others are acoustic‑driven, revealing a gap between what is encoded and what is utilized in current audio‑language models.

By Bhuvan Koduru, Dareen Safar B Alharthi, Rita Singh, Bhiksha Raj
arXiv Computation and Language
Sep 1

Vocal Music under Phoneme-Conditional Analysis

arXiv:2608.30823v1 Announce Type: cross Abstract: The vocal music of each language carries a distinctive sonic identity, even without instrumental accompaniment. We ask whether these differences are...

By Hayoon Kim, Kyogu Lee