The study examines how English homophones—words that sound identical but have different meanings—are pronounced differently in spoken language. Analyzing 14,000 homophone tokens from American television news, researchers found that pairs such as "weight" and "wait" exhibit distinct phonetic realizations that can be predicted from their contextual meanings, even after controlling for word duration. Time‑normalized spectrograms proved effective for detecting these subtle differences without relying on phonetic transcriptions.
Time-normalized f0 contours of Mandarin words in conversational speech have been shown to be predictable in part from their contextualized embeddings (CEs). The present study investigates whether CEs also predict spoken word duration for 7470 tokens of Mandarin monosyllabic CV words extracted from a Mandarin corpus of spontaneous speech.
arXiv:2607. 04154v1 Announce Type: cross Abstract: This paper explains the principles and provides examples of a new method for distinguishing between FAKE human speech synthesized by generative AI and natural speech.
By Yusei Tamura, Shigekazu Ishihara, Ken Ito
The paper introduces a method for measuring accent differences that balances interpretability and practicality. It proposes using articulatory representations obtained via articulatory inversion as an interpretable basis for accent comparison, while employing optimal transport to compare accents across any type of recording. This approach aims to overcome the limitations of traditional phonetic analyses and embedding‑based methods, which are either time‑consuming or non‑interpretable.
By Charles McGhee, Mark J. F. Gales, Kate M. Knill
The paper evaluates bias in phoneme-based automatic speech recognition (ASR) systems, focusing on WhisperIPA and ZIPA, which produce International Phonetic Alphabet (IPA) transcriptions. Using multilingual speech corpora and demographically annotated English datasets, the authors compare model-generated IPA against grapheme-to-phoneme (G2P) outputs with both standard phoneme error rate (PER) and a new Soft PER metric that allows linguistically similar substitutions. The study finds persistent disparities across language, gender, accent, ethnicity, and age, even when accounting for acceptable phonemic variation.
By Maneesha Rani Saha, Catherine Bao, Neal Patwari
The study introduces a scalable acoustic‑masking method to quantify how much each consonant contributes to word intelligibility. By silencing individual consonants in isolated words and measuring misrecognition rates with three ASR models, the authors define a mask‑induced misrecognition rate (MMR). Across English, Spanish, German, and Czech, MMR negatively correlates with phoneme frequency and positively with functional load, revealing that consonant importance varies by language.
By Eunjung Yeo, Kwanghee Choi, Krupaben Kothadia, Visar Berisha, Julie M. Liss, David R. Mortensen, David Harwath
The study investigates how bilingual politicians structure the timing of their speeches in Luxembourgish and French, analyzing 400 sentences from ten speakers. Rhythm metrics were computed for consonants and vowels, revealing that consonant patterns are largely speaker-specific while vowel patterns are strongly influenced by language choice. French tokens exhibited longer, more variable vowels and vocalic intervals, whereas consonant timing differences were smaller, with no significant language-by-gender interactions.
By Nina Hosseini-Kivanani, Nafiseh Taghva, Peter Gilles, Oliver Niebuhr
The paper investigates whether audio language models encode phonetic features similarly when processing spoken versus written input. By comparing mean representations of minimal phoneme pairs across six models, seven features, and 15 languages, the study finds that only voicing in two Qwen2.5-Omni models shows a significant shared direction, and that the model family—not size—determines feature representation. The analysis uses cosine similarity against a random-pair reference to assess alignment across modalities.
The paper introduces a new evaluation framework for phonetic encoding algorithms, using a generalized Rand Index called the Hüllermeier‑Rifqi Index. It measures discordance by comparing pairwise similarity scores of ground‑truth IPA transcriptions with those of encoded strings, adjusted against a random generator. The method is applied to multilingual datasets, assessing recall via collision rate and demonstrating its use in evaluating orthographic transparency.
By Can \"Ozbey, Emre Kaplan, Berkin Deniz Kahya
Neural TTS systems can sound natural across languages, but naturalness does not guarantee the preservation of sound contrasts that distinguish words from their grammatical forms. Standard metrics like MOS do not test for this.
The neutral, or floating, tone of Mandarin Chinese is a tone with an enigmatic set of properties. It has been described as a reduced tone, or as a tone that sometimes is lexically fixed but that can also be toneless.
arXiv:2608.30823v1 Announce Type: cross
Abstract: The vocal music of each language carries a distinctive sonic identity, even without instrumental accompaniment. We ask whether these differences are...
By Hayoon Kim, Kyogu Lee