Hugging Face Trending Papers

Structured Phonological Representations for Audio-Articulatory rtMRI Speech Classification

Read the original on Hugging Face Trending Papers →

Real-time MRI makes it possible to observe vocal-tract articulation during speech, but mapping these articulatory patterns to phonetic and phonological categories remains challenging. We investigate whether PhonoQ, an audio-based model trained to recognize structured phonological features, provides useful information for audio--articulatory modeling.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Aug 25

Multi-Task Learning for Non-Canonical Phoneme Recognition via Articulatory Feature Decomposition

The paper proposes a linguistically structured multi‑task learning framework for recognizing non‑canonical phonemes by decomposing phoneme prediction into articulatory feature dimensions such as manner, place, and voicing. A hierarchical architecture with task‑specific heads and a cross‑attention fusion module is combined with semi‑supervised Momentum Pseudo‑Labeling and a cascaded training strategy that gradually introduces articulatory tasks. Experiments on the L2‑ARCTIC dataset demonstrate significant improvements over baseline models and produce interpretable error patterns aligned with phonological feature structure.

By Sophia Riaz, Haoze Zheng, Amos Roche, Miyu Zhang, Anamika Ragu, Salvatore Penachio, Kaustav Mukherjee, Aneesh Jonelagadda
arXiv AI
Sep 25

Anatomy-aware cross-speaker adaptation of complete vocal-tract acoustic-to-articulatory inversion

The paper introduces a geometric adaptation framework for cross‑speaker acoustic‑to‑articulatory inversion that leverages anatomical landmarks on vertebrae and dental structures. By applying an affine transformation followed by a thin‑plate spline deformation, the method maps predicted vocal‑tract contours from a fixed model to unseen speakers without retraining. Experiments on a single‑speaker rt‑MRI database and eight additional speakers show that the combined affine‑plus‑TPS approach with 12 and 14 landmarks yields the lowest mean point‑to‑closest‑point error of 3.19 mm.

By Nhat-Nam Nguyen, Pierre-Andre Vuissoz, Yves Laprie