The study examines whether English homophones differ in phonetic realization beyond spoken word duration. Analyzing 14,000 homophone tokens from American television news, it finds that pairs such as "weight" and "wait" exhibit distinct phonetic patterns that can be predicted from their meanings in context. These differences persist even after controlling for duration, and time‑normalized spectrograms prove effective for detecting such fine‑grained phonetic variation without relying on phonetic transcriptions.
By Yu-Hsiang Tseng, Mirjam T. C. Ernestus, Louis F. M. ten Bosch, R. Harald Baayen
Time-normalized f0 contours of Mandarin words in conversational speech have been shown to be predictable in part from their contextualized embeddings (CEs). The present study investigates whether CEs also predict spoken word duration for 7470 tokens of Mandarin monosyllabic CV words extracted from a Mandarin corpus of spontaneous speech.
arXiv:2607. 04154v1 Announce Type: cross Abstract: This paper explains the principles and provides examples of a new method for distinguishing between FAKE human speech synthesized by generative AI and natural speech.
By Yusei Tamura, Shigekazu Ishihara, Ken Ito
The study investigates how bilingual politicians structure the timing of their speeches in Luxembourgish and French, analyzing 400 sentences from ten speakers. Rhythm metrics were computed for consonants and vowels, revealing that consonant patterns are largely speaker-specific while vowel patterns are strongly influenced by language choice. French tokens exhibited longer, more variable vowels and vocalic intervals, whereas consonant timing differences were smaller, with no significant language-by-gender interactions.
By Nina Hosseini-Kivanani, Nafiseh Taghva, Peter Gilles, Oliver Niebuhr
The paper introduces a method for measuring accent differences that balances interpretability and practicality. It proposes using articulatory representations obtained via articulatory inversion as an interpretable basis for accent comparison, while employing optimal transport to compare accents across any type of recording. This approach aims to overcome the limitations of traditional phonetic analyses and embedding‑based methods, which are either time‑consuming or non‑interpretable.
By Charles McGhee, Mark J. F. Gales, Kate M. Knill
The neutral, or floating, tone of Mandarin Chinese is a tone with an enigmatic set of properties. It has been described as a reduced tone, or as a tone that sometimes is lexically fixed but that can also be toneless.