The study examined 400 utterances from 10 politicians speaking in both Luxembourgish and French to determine how much of charismatic prosody is due to speaker identity versus language. Mixed‑effects modeling revealed that speaker identity explained most of the variance, while language contributed less but still produced systematic differences: French speech had higher shimmer and phrase‑final F0, suggesting a polite, respectful tone, whereas Luxembourgish speech showed stronger mid‑frequency spectral energy, indicating a more vocally present profile. These acoustic patterns reflect the sociolinguistic roles of Luxembourgish as an informal identity language and French as a high‑prestige institutional variety.
By Nina Hosseini-Kivanani, Nafiseh Taghva, Peter Gilles, Oliver Niebuhr
arXiv:2608.30828v1 Announce Type: new
Abstract: We present three large-scale studies of spoken parliamentary speech across four Slavic languages (Croatian, Czech, Polish, Serbian), drawing on over 6,...
By Ivan Porupski, Nikola Ljube\v{s}i\'c
Filled pauses (FPs) are a universal feature of spontaneous speech, yet most studies rely on small, single-language corpora, limiting the generalisability of their findings. We analyse ~4,000 hours of parliamentary speech across four related Slavic languages (Croatian, Czech, Polish, Serbian).
The study analyzes oral political language in U.S. presidential debates from 1960 to 2024, focusing on 19 candidates. It finds a clear trend toward simplification: sentence length and complex terms have decreased, while emotional tone has risen and logical, rational content has diminished. The research also explores whether specific presidents exhibit unique stylistic traits and whether language patterns correlate with electoral success.
By Jacques Savoy
arXiv:2608.30260v1 Announce Type: cross
Abstract: While it is well-established that prosody carries crucial cues for syntactic structure, the degree and nature of correspondence between these two dom...
By Junghyun Min, Alex Warstadt, Tamar I. Regev, Tiago Pimentel, Ethan Gotlieb Wilcox
This paper investigates whether prosodic features—pitch, energy, and timing—are preserved when speech is translated between languages. Using multilingual dubbing data for English‑German, English‑Spanish, and English‑French pairs, the authors conduct a fine‑grained cross‑lingual analysis to quantify similarities and differences in prosody. The study identifies inherent cross‑lingual correlations in prosodic structure and explores how linguistic and alignment factors influence these patterns.
By Haopeng Xie, Ismail Rasim Ulgen, Sofia Son, Berrak Sisman, Philipp Koehn
arXiv:2608. 03507v1 Announce Type: cross Abstract: Historical language change affects morphology, syntax, semantics, and pragmatics, yet computational studies typically examine these levels with incompatible representations and therefore cannot determine whether they evolve together across languages.
By Gagan Bhatia, Julian Schlenker, Simone Paolo Ponzetto, Steffen Eger
arXiv:2606. 16753v1 Announce Type: cross Abstract: As Large Language Models (LLMs) become embedded in everyday communication, capturing regional linguistic variation is essential for reliable and equitable language use.
By Rafael Ferreira, In\^es Vieira, In\^es Calvo, James Furtado, Iago Paulo, Diogo Tavares, Diogo Gl\'oria-Silva, David Semedo, Jo\~ao Magalh\~aes
Automatic speech recognition (ASR) systems exhibit unequal error rates across speaker groups, motivating interventions on their internal representations. We ask whether speaker-linked attributes that...
The study examines whether English homophones differ in phonetic realization beyond spoken word duration. Analyzing 14,000 homophone tokens from American television news, it finds that pairs such as "weight" and "wait" exhibit distinct phonetic patterns that can be predicted from their meanings in context. These differences persist even after controlling for duration, and time‑normalized spectrograms prove effective for detecting such fine‑grained phonetic variation without relying on phonetic transcriptions.
By Yu-Hsiang Tseng, Mirjam T. C. Ernestus, Louis F. M. ten Bosch, R. Harald Baayen
Time-normalized f0 contours of Mandarin words in conversational speech have been shown to be predictable in part from their contextualized embeddings (CEs). The present study investigates whether CEs also predict spoken word duration for 7470 tokens of Mandarin monosyllabic CV words extracted from a Mandarin corpus of spontaneous speech.
The paper evaluates bias in phoneme-based automatic speech recognition (ASR) systems, focusing on WhisperIPA and ZIPA, which produce International Phonetic Alphabet (IPA) transcriptions. Using multilingual speech corpora and demographically annotated English datasets, the authors compare model-generated IPA against grapheme-to-phoneme (G2P) outputs with both standard phoneme error rate (PER) and a new Soft PER metric that allows linguistically similar substitutions. The study finds persistent disparities across language, gender, accent, ethnicity, and age, even when accounting for acceptable phonemic variation.
By Maneesha Rani Saha, Catherine Bao, Neal Patwari