arXiv:2609. 21362v1 Announce Type: new Abstract: Conventional tokenizers represent text as characters or statistically derived subwords, overlooking the internal phonological structure of syllables and often requiring large vocabularies.
By Nghia Hieu Nguyen, Thai Bao Huynh, Binh-An Dinh-Le, Phu Gia Hoang, Dat Tien Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen
Time-normalized f0 contours of Mandarin words in conversational speech have been shown to be predictable in part from their contextualized embeddings (CEs). The present study investigates whether CEs also predict spoken word duration for 7470 tokens of Mandarin monosyllabic CV words extracted from a Mandarin corpus of spontaneous speech.
The paper investigates Mandarin Chinese reduplicative constructions that repeat two-character base words or their constituents, such as expressions meaning ‘in good health’ or ‘discuss a bit’. Using Tencent word embeddings, the study demonstrates that distributional semantics can recover known semantic and grammatical properties of these reduplications, revealing clear semantic and pragmatic differentiation between the two patterns. Procrustes analysis shows that the overall organization of the base-word space is largely preserved in the reduplication space, with local mismatches indicating discourse-pragmatic reorganization.
By Chaoyi Wu, Yu-Hsiang Tseng, R. Harald Baayen
The study investigates how multimodal large language models (MLLMs) use prosodic cues in sarcasm detection. By testing Qwen2.5‑Omni and Qwen3‑Omni on Mandarin Chinese and English across five modality conditions, the authors find that adding audio increases false positives without improving true positives. Acoustic error analysis shows that models rely on a stereotypical prosodic pattern—elevated pitch and irregular pausing—that does not align with genuine sarcasm cues, and manipulating these dimensions alone can raise false positive rates up to 60%. The same effect appears in Gemini 3 Flash Preview, indicating the heuristic is not limited to a single architecture.
By Yongjian Chen, Pengfei Wei, Yiqun Sun, Zhu Li, Lawrence B. Hsieh
arXiv:2609.24310v1 Announce Type: cross
Abstract: Text-to-speech models for Bantu tonal languages are challenged by a tonal system that is rooted in both the lexis (i.e., the inventory of words, stem...
By Antoine Nzeyimana
arXiv:2608.30823v1 Announce Type: cross
Abstract: The vocal music of each language carries a distinctive sonic identity, even without instrumental accompaniment. We ask whether these differences are...
By Hayoon Kim, Kyogu Lee
This paper investigates whether prosodic features—pitch, energy, and timing—are preserved when speech is translated between languages. Using multilingual dubbing data for English‑German, English‑Spanish, and English‑French pairs, the authors conduct a fine‑grained cross‑lingual analysis to quantify similarities and differences in prosody. The study identifies inherent cross‑lingual correlations in prosodic structure and explores how linguistic and alignment factors influence these patterns.
By Haopeng Xie, Ismail Rasim Ulgen, Sofia Son, Berrak Sisman, Philipp Koehn
Multimodal Large Language Models (MLLMs) process speech and text jointly, yet whether they exploit prosodic cues for pragmatic inference or rely on surface acoustic patterns has received little system...
T‑SANDHI is a lightweight Taiwanese Hokkien ASR model that addresses tone sandhi by decoupling surface acoustics from lexical intent on a frozen Whisper backbone. It uses a lexicon‑guided multi‑task learning framework with text‑derived pseudo labels and a hybrid injection module that dynamically gates independent citation and sandhi phonetic streams. Experiments on the TAT‑MOE corpus and two blind test sets show that this explicit disentanglement resolves tonal mapping confusion and outperforms baselines while keeping parameter usage low.
By Hung-Yang Sung, Chien-Chun Wang, Tien-Hong Lo, Yu-Sheng Tsao, Yung-Chang Hsu, Berlin Chen
arXiv:2601.09050v2 Announce Type: replace
Abstract: Tonal low-resource languages are widely spoken but remain underserved by modern speech technologies. A central challenge is learning speech represe...
By Tianyi Xu, Xuan Ouyang, Binwei Yao, Shoua Xiong, Sara Misurelli, Maichou Lor, Junjie Hu
arXiv:2609.24903v1 Announce Type: new
Abstract: Tone languages constitute over 50-70% of the world's languages, but the vast majority are low-resource, lacking the large transcribed corpora needed fo...
By Qisheng Liao, Youngah Do
arXiv:2606. 17835v1 Announce Type: cross Abstract: This study examines the extent to which the wav2vec2.
By James Kirby, Ioana Krehan, Michele Gubian