Marking Contour Tones in Yor\`{u}b\'{a}
arXiv:2609.38627v1 Announce Type: new Abstract: Yor\`ub\'a is a tonal language in which contour tones pose persistent orthographic challenges. These are especially notable for personal names and lexi...
arXiv:2609.38627v1 Announce Type: new Abstract: Yor\`ub\'a is a tonal language in which contour tones pose persistent orthographic challenges. These are especially notable for personal names and lexi...
arXiv:2609.38627v2 Announce Type: new Abstract: Yor\`ub\'a is a tonal language in which contour tones pose persistent orthographic challenges. These are especially notable for personal names and lexi...
arXiv:2609.22125v1 Announce Type: new Abstract: Standard tokenizers used in large language models produce malformed text when applied to Brahmic scripts. They are a family of abugidas, writing system...
arXiv:2609.14817v1 Announce Type: new Abstract: In Yor\`ub\'a, pitch alone separates \d{o}k\d{o} (husband, Mid), \d{o}k\d{\`o} (vehicle, Low), and \d{o}k\d{\'o} (hoe, High) -- the diacritics ARE the...
arXiv:2608. 01935v1 Announce Type: cross Abstract: Prior work in Ancient Greek NLP relies on corpora that do not disambiguate the phonemic vowel length of alpha, iota, and ypsilon, together known as the dichrona.
arXiv:2609.24310v1 Announce Type: cross Abstract: Text-to-speech models for Bantu tonal languages are challenged by a tonal system that is rooted in both the lexis (i.e., the inventory of words, stem...
arXiv:2608.29170v1 Announce Type: new Abstract: This paper proposes the Sinitic Romanization Ecosystem, a cross-lingual Sinitic romanization design framework with supporting digital infrastructure an...
The article examines how Byte‑Pair Encoding (BPE) tokenization handles Polish, an inflectional language, and finds that BPE tends to stabilize frequent surface fragments of grammatical exponents rather than true grammatical categories. It introduces the concept of grammatical form anchoring, showing that certain Polish verb forms can signal the speaking subject without an explicit pronoun, and highlights that language models may lack a stable grammatical "I" and can shift gender or mirror user forms. The study proposes Roclawski’s segmentation‑flexional forms as a diagnostic framework and suggests that more stable Polish modeling would require sublexical stabilization, anchoring grammatical form in the inflectional system, and maintaining the grammatical "I" in dialogue.
The paper introduces an ASCII‑only romanization system that covers all phonemic distinctions in Thai and Lao, including segmental contrasts, vowel length, and lexical tone. It ensures one‑symbol‑one‑phoneme transparency and systematic cross‑lingual correspondence between the two languages, while also aligning with Pinyin and Jyutping where possible. The design prioritizes synchronic phonetic clarity, offers optional historical tone annotations, and results in a readable, keyboard‑friendly, machine‑processable representation useful for language learning and cross‑lingual speech processing.
The paper introduces MoirfEolas, a dataset of over 35,000 Irish words annotated with their morphological components, and CríochScore, a metric that measures how well tokenization aligns with these morphological boundaries. Using CríochScore, the authors evaluate common tokenization algorithms and find that the Unigram Language Model best aligns with Irish morphology. They also discuss trade‑offs between morphological alignment, compression, and vocabulary efficiency, offering practical guidance for Irish NLP development.
Vagdhenu is a Sanskrit shloka‑to‑chant text‑to‑speech system that preserves meter (vrutta) and phonological nuances. It builds on an off‑the‑shelf flow‑matching backbone and a large‑scale neural vocoder, adding a Kannada‑based frontend to avoid schwa deletion, a phonology‑aware frontend handling visarga sandhi and sibilant distinctions, and a vrutta‑aware reference selection mechanism. The authors report that a text‑side prosody conditioner is ineffective in their architecture, while reference clips and voice‑steering retraining provide the necessary prosody control, and they demonstrate the system’s performance on a 32‑chapter video corpus and an audio app covering 18,000 verses. whyItMatters":"The system delivers high‑fidelity, meter‑aware Sanskrit chanting, enabling large‑scale deployment of authentic recitations for educational and cultural preservation purposes."
arXiv:2603. 26292v2 Announce Type: replace-cross Abstract: Syllable-level units offer compact and linguistically meaningful representations for spoken language modeling and unsupervised word discovery, but research on syllabification remains fragmented across disparate implementations, datasets, and evaluation protocols.