arXiv Computation and Language

Letters hide the truth from our eyes: English homophones have meaningfully different phonetic realizations

The study examines whether English homophones differ in phonetic realization beyond spoken word duration. Analyzing 14,000 homophone tokens from American television news, it finds that pairs such as "weight" and "wait" exhibit distinct phonetic patterns that can be predicted from their meanings in context. These differences persist even after controlling for duration, and time‑normalized spectrograms prove effective for detecting such fine‑grained phonetic variation without relying on phonetic transcriptions.

Hugging Face Trending Papers
Aug 27

Letters hide the truth from our eyes: English homophones have meaningfully different phonetic realizations

The study examines how English homophones—words that sound identical but have different meanings—are pronounced differently in spoken language. Analyzing 14,000 homophone tokens from American television news, researchers found that pairs such as "weight" and "wait" exhibit distinct phonetic realizations that can be predicted from their contextual meanings, even after controlling for word duration. Time‑normalized spectrograms proved effective for detecting these subtle differences without relying on phonetic transcriptions.

arXiv AI
Sep 12

Flexible and Interpretable Accent Distance Measurements

The paper introduces a method for measuring accent differences that balances interpretability and practicality. It proposes using articulatory representations obtained via articulatory inversion as an interpretable basis for accent comparison, while employing optimal transport to compare accents across any type of recording. This approach aims to overcome the limitations of traditional phonetic analyses and embedding‑based methods, which are either time‑consuming or non‑interpretable.

By Charles McGhee, Mark J. F. Gales, Kate M. Knill
arXiv Computation and Language
Sep 18

Evaluating Bias in Phoneme-Based Automatic Speech Recognition Systems: An Analysis of IPA Transcription Models

The paper evaluates bias in phoneme-based automatic speech recognition (ASR) systems, focusing on WhisperIPA and ZIPA, which produce International Phonetic Alphabet (IPA) transcriptions. Using multilingual speech corpora and demographically annotated English datasets, the authors compare model-generated IPA against grapheme-to-phoneme (G2P) outputs with both standard phoneme error rate (PER) and a new Soft PER metric that allows linguistically similar substitutions. The study finds persistent disparities across language, gender, accent, ethnicity, and age, even when accounting for acceptable phonemic variation.

By Maneesha Rani Saha, Catherine Bao, Neal Patwari
arXiv Computation and Language
Sep 14

Quantifying Consonant Contributions to Word Intelligibility via Acoustic Masking

The study introduces a scalable acoustic‑masking method to quantify how much each consonant contributes to word intelligibility. By silencing individual consonants in isolated words and measuring misrecognition rates with three ASR models, the authors define a mask‑induced misrecognition rate (MMR). Across English, Spanish, German, and Czech, MMR negatively correlates with phoneme frequency and positively with functional load, revealing that consonant importance varies by language.

By Eunjung Yeo, Kwanghee Choi, Krupaben Kothadia, Visar Berisha, Julie M. Liss, David R. Mortensen, David Harwath
arXiv Computation and Language
Sep 16

Speaker-Specific and Language-Dependent Temporal Organization in Bilingual Political Speech

The study investigates how bilingual politicians structure the timing of their speeches in Luxembourgish and French, analyzing 400 sentences from ten speakers. Rhythm metrics were computed for consonants and vowels, revealing that consonant patterns are largely speaker-specific while vowel patterns are strongly influenced by language choice. French tokens exhibited longer, more variable vowels and vocalic intervals, whereas consonant timing differences were smaller, with no significant language-by-gender interactions.

By Nina Hosseini-Kivanani, Nafiseh Taghva, Peter Gilles, Oliver Niebuhr
Hugging Face Trending Papers
Sep 24

Do Audio Language Models Hear and Read Distinctive Features Alike?

The paper investigates whether audio language models encode phonetic features similarly when processing spoken versus written input. By comparing mean representations of minimal phoneme pairs across six models, seven features, and 15 languages, the study finds that only voicing in two Qwen2.5-Omni models shows a significant shared direction, and that the model family—not size—determines feature representation. The analysis uses cosine similarity against a random-pair reference to assess alignment across modalities.

arXiv Computation and Language
Sep 7

Evaluation of Phonetic Encoding Algorithms on Transcription Datasets

The paper introduces a new evaluation framework for phonetic encoding algorithms, using a generalized Rand Index called the Hüllermeier‑Rifqi Index. It measures discordance by comparing pairwise similarity scores of ground‑truth IPA transcriptions with those of encoded strings, adjusted against a random generator. The method is applied to multilingual datasets, assessing recall via collision rate and demonstrating its use in evaluating orthographic transparency.

By Can \"Ozbey, Emre Kaplan, Berkin Deniz Kahya
arXiv Computation and Language
Sep 1

Vocal Music under Phoneme-Conditional Analysis

arXiv:2608.30823v1 Announce Type: cross Abstract: The vocal music of each language carries a distinctive sonic identity, even without instrumental accompaniment. We ask whether these differences are...

By Hayoon Kim, Kyogu Lee