Hugging Face Trending Papers

Umm... With Transformers? Insights from Filled Pause Use across Four Slavic Parliaments

Filled pauses (FPs) are a universal feature of spontaneous speech, yet most studies rely on small, single-language corpora, limiting the generalisability of their findings. We analyse ~4,000 hours of parliamentary speech across four related Slavic languages (Croatian, Czech, Polish, Serbian).

arXiv Computation and Language
Sep 16

Speaker-Specific and Language-Dependent Temporal Organization in Bilingual Political Speech

The study investigates how bilingual politicians structure the timing of their speeches in Luxembourgish and French, analyzing 400 sentences from ten speakers. Rhythm metrics were computed for consonants and vowels, revealing that consonant patterns are largely speaker-specific while vowel patterns are strongly influenced by language choice. French tokens exhibited longer, more variable vowels and vocalic intervals, whereas consonant timing differences were smaller, with no significant language-by-gender interactions.

By Nina Hosseini-Kivanani, Nafiseh Taghva, Peter Gilles, Oliver Niebuhr
arXiv Computation and Language
Sep 16

Speaker or Language? Explaining Variance in Charismatic Prosody Across Luxembourgish and French

The study examined 400 utterances from 10 politicians speaking in both Luxembourgish and French to determine how much of charismatic prosody is due to speaker identity versus language. Mixed‑effects modeling revealed that speaker identity explained most of the variance, while language contributed less but still produced systematic differences: French speech had higher shimmer and phrase‑final F0, suggesting a polite, respectful tone, whereas Luxembourgish speech showed stronger mid‑frequency spectral energy, indicating a more vocally present profile. These acoustic patterns reflect the sociolinguistic roles of Luxembourgish as an informal identity language and French as a high‑prestige institutional variety.

By Nina Hosseini-Kivanani, Nafiseh Taghva, Peter Gilles, Oliver Niebuhr
arXiv AI
Sep 17

Evolution of US Oral Political Language

The study analyzes oral political language in U.S. presidential debates from 1960 to 2024, focusing on 19 candidates. It finds a clear trend toward simplification: sentence length and complex terms have decreased, while emotional tone has risen and logical, rational content has diminished. The research also explores whether specific presidents exhibit unique stylistic traits and whether language patterns correlate with electoral success.

By Jacques Savoy
arXiv AI
Sep 1

Using Prosody to Predict Syntactic Structure

arXiv:2608.30260v1 Announce Type: cross Abstract: While it is well-established that prosody carries crucial cues for syntactic structure, the degree and nature of correspondence between these two dom...

By Junghyun Min, Alex Warstadt, Tamar I. Regev, Tiago Pimentel, Ethan Gotlieb Wilcox
arXiv AI
Sep 11

Voice or Stereotype? Disentangling Acoustic and Content-Based Gender in Speech-to-Speech Models

The study investigates how speech‑to‑speech (S2S) models handle gender, distinguishing between the acoustic voice and the content’s gender cues. Experiments across five models in English, Spanish, and Mandarin show that while the rendered voice remains unbiased, the models consistently attribute speaker gender based on textual content rather than voice. When content and voice disagree, misgendering rates soar to 90%, whereas agreement yields only 2% misgendering.

By Xiaoqun Liu, Tanu Mitra, Harshit Rajgarhia, Abhishek Mukherji
arXiv Machine Learning
Sep 11

When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text

Large language models (LLMs) are increasingly used to assess social bias in text, but the passages they evaluate often contain surface noise such as typos and broken punctuation. This study applied five realistic noise conditions at varying intensities to 3,822 stereotype‑related responses and compared bias judgments on noisy versus original text. The findings show that noise disproportionately turns neutral judgments into biased ones—up to 120 times more likely—while rarely converting biased judgments into neutral ones, and that the most fragile LLM judge exhibits the greatest distortion at mild noise levels. As LLMs become more robust, the bias distortion tends toward parity rather than reversal, meaning bias measured on noisy text is systematically overestimated, especially in fairness‑critical categories.

By DongHyun Ryu, Jaehyeok Lee, YeongJun Hwang, JinYeong Bak
Hugging Face Trending Papers
Sep 10

When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text

Large language models used as judges for social bias are affected by noisy text, such as typos and broken punctuation. In experiments with 3,822 stereotype-related responses, noise more often turns neutral judgments into biased ones than the reverse, with up to a 120‑fold difference. The effect is strongest at mild realistic noise levels and leads to systematic overestimation of bias, especially in fairness‑critical categories.