Time-normalized f0 contours of Mandarin words in conversational speech have been shown to be predictable in part from their contextualized embeddings (CEs). The present study investigates whether CEs also predict spoken word duration for 7470 tokens of Mandarin monosyllabic CV words extracted from a Mandarin corpus of spontaneous speech.
arXiv:2606. 17835v1 Announce Type: cross Abstract: This study examines the extent to which the wav2vec2.
By James Kirby, Ioana Krehan, Michele Gubian
arXiv:2606. 24093v1 Announce Type: cross Abstract: We ask whether the geographic origin of Tang-dynasty poets leaves a detectable linguistic trace in their work.
By Chi-Sheng Chen, Hung-Yun Liu
Recent advances in zero-shot text-to-speech (TTS) have substantially improved speech quality and voice cloning fidelity. However, many zero-shot TTS systems still depend on audio prompt transcripts at inference time.
arXiv:2608. 01281v1 Announce Type: cross Abstract: Phoneme-based multilingual automatic speech recognition (ASR) can share acoustic evidence across languages more directly than language-specific subword modeling.
By Saierdaer Yusuyin, Nanling Jiang, Hao Huang, Zhijian Ou
arXiv:2607. 04154v1 Announce Type: cross Abstract: This paper explains the principles and provides examples of a new method for distinguishing between FAKE human speech synthesized by generative AI and natural speech.
By Yusei Tamura, Shigekazu Ishihara, Ken Ito
arXiv:2606. 01016v1 Announce Type: cross Abstract: While End-to-End (E2E) Speech-Large Language Models (Speech-LLMs) are rapidly evolving, their evaluation methodologies remain limited to the era of simple transcription.
By Sicheng Yang, Shulan Ruan, Shiwei Wu, Yu Liu, Lu Fan, Zhi Li, You He
arXiv:2603. 26292v2 Announce Type: replace-cross Abstract: Syllable-level units offer compact and linguistically meaningful representations for spoken language modeling and unsupervised word discovery, but research on syllabification remains fragmented across disparate implementations, datasets, and evaluation protocols.
By H\'ector Javier V\'azquez Mart\'inez
arXiv:2406. 05335v3 Announce Type: replace-cross Abstract: Generation of text and speech in natural languages can be modeled as a stochastic process.
By Kai Nakaishi, Yoshihiko Nishikawa, Koji Hukushima
arXiv:2608. 01935v1 Announce Type: cross Abstract: Prior work in Ancient Greek NLP relies on corpora that do not disambiguate the phonemic vowel length of alpha, iota, and ypsilon, together known as the dichrona.
By Albin Th\"orn Cleland, Eric Cullhed
arXiv:2607. 21332v1 Announce Type: cross Abstract: Phonetic forced alignment is a key technique in phonetic research, yet existing alignment systems lack specialized models for low-resource language varieties.
By Zhiheng Qian, Aini Li, Hai Hu, Liang Zhao
Phonetic forced alignment is a key technique in phonetic research, yet existing alignment systems lack specialized models for low-resource language varieties. We address this by training text-dependent and text-independent aligners for Chengdu Mandarin using a 17-hour corpus and a custom G2P dictionary.