arXiv AI

Phonetic Error Analysis of Raw Waveform Acoustic Models

arXiv:2606. 07030v1 Announce Type: cross Abstract: We analyse error patterns of raw waveform acoustic models on TIMIT phone recognition beyond the overall phone error rate (PER).

arXiv Machine Learning
Jul 14

An Empirical Recipe for Universal Phone Recognition

arXiv:2603. 29042v2 Announce Type: replace-cross Abstract: Phone recognition (PR) is a key enabler of multilingual and low-resource speech processing tasks, yet robust performance remains elusive.

By Shikhar Bharadwaj, Chin-Jou Li, Kwanghee Choi, Eunjung Yeo, William Chen, Shinji Watanabe, David R. Mortensen
arXiv Computation and Language
Sep 25

DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units

DiscoPhon is a multilingual benchmark designed to evaluate unsupervised phoneme discovery from discrete speech units. It includes 6 development and 6 test languages that cover a wide range of phonemic contrasts, and requires systems to generate discrete units mapped to a predefined phoneme inventory using only 10 hours of speech from an unseen language. The benchmark assesses unit quality, recognition, and segmentation, and provides four pretrained multilingual HuBERT and SpidR baselines that demonstrate current models can produce units that correlate well with phonemes, though performance varies across languages.

By Maxime Poli, Manel Khentout, Angelo Ortiz Tandazo, Ewan Dunbar, Emmanuel Chemla, Emmanuel Dupoux
arXiv AI
6d ago

HuPER: A Human-Inspired Framework for Phonetic Perception

HuPER is a human-inspired framework that models phonetic perception as adaptive inference over acoustic‑phonetics evidence and linguistic knowledge. Using only 100 hours of training data, it achieves state‑of‑the‑art phonetic error rates on five English benchmarks and demonstrates strong zero‑shot transfer to 95 unseen languages. It uniquely enables adaptive, multi‑path phonetic perception across diverse acoustic conditions, and all training data, models, and code are open‑sourced.

By Chenxu Guo, Jiachen Lian, Yisi Liu, Baihe Huang, Shriyaa Narayanan, Bixing Wu, Zoe Ezzes, Jet Vonk, Zachary Miller, Cheol Jun Cho, Maria Gorno-Tempini, Gopala Anumanchipalli
Hugging Face Trending Papers
Aug 10

Edge Phoneme Recognition for Children's Speech through Age-Aware Training

Detecting phonemes from children's speech has historically been difficult due to the scarcity of training data, and unique characteristics of children's speech. During a phoneme detection competition, we found that training a lightweight model to predict the age of the learner, as well as the phoneme sequence, enabled a 94M-parameter model to outperform WavLM Large models (317M) on the target DrivenData distribution, and fall within approximately 0.

arXiv AI
Sep 2

Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models

The paper investigates whether audio‑language models capture paralinguistic cues beyond spoken content. Using the Expresso dataset and four open‑source models, the authors trace how speaking style information is encoded in the late layers of the audio encoder but is degraded before reaching the final output. They find that some models are content‑driven while others are acoustic‑driven, revealing a gap between what is encoded and what is utilized in current audio‑language models.

By Bhuvan Koduru, Dareen Safar B Alharthi, Rita Singh, Bhiksha Raj