arXiv:2607. 22699v1 Announce Type: new Abstract: As Large Language Models (LLMs) grow more capable across diverse tasks, their (in)ability to generalize remains difficult to quantify and poorly understood beyond limited domains.
By Supantho Rakshit, Adele Goldberg, Henry Conklin
The paper introduces two new corpus‑level, reference‑free metrics—Phoneme‑Cluster Mutual Information (PCMI) and Word Acoustic Consistency Score (WACS)—that use self‑supervised speech representations to evaluate forced alignment quality. PCMI quantifies how well aligned phoneme labels agree with clusters derived from SSL representations, while WACS assesses consistency across repeated word realizations via dynamic time warping of word representation sequences. Experiments on 85 languages from FLEURS and 45 languages in DoReCo show that both metrics degrade predictably under alignment perturbations, effectively distinguish high‑ from low‑quality alignments, and correlate strongly with traditional timestamp‑based measures, enabling scalable, multilingual alignment evaluation without manual annotations.
By V. S. D. S. Mahesh Akavarapu, Michael Daniel, Gerhard J\"ager
CWoMP (Contrastive Word‑Morpheme Pretraining) is a new approach for generating interlinear glossed text that treats morphemes as atomic form‑meaning units with learned representations. It uses a contrastively trained encoder to align words in context with their constituent morphemes in a shared embedding space, and an autoregressive decoder that retrieves morpheme sequences from a mutable lexicon of these embeddings. The method yields interpretable predictions grounded in lexicon entries and allows users to improve results at inference time by expanding the lexicon without retraining, achieving superior performance and efficiency on diverse low‑resource languages, especially in extremely low‑resource settings.
By Morris Alper, Enora Rice, Bhargav Shandilya, Alexis Palmer, Lori Levin
arXiv:2608. 01935v1 Announce Type: cross Abstract: Prior work in Ancient Greek NLP relies on corpora that do not disambiguate the phonemic vowel length of alpha, iota, and ypsilon, together known as the dichrona.
By Albin Th\"orn Cleland, Eric Cullhed
KinyaEmbed is the first sentence‑embedding model specifically designed for Kinyarwanda, built on KinyaBERT‑large and trained through a four‑stage curriculum that incorporates paraphrase pairs, translated MNLI triplets, OPUS‑100 translation pairs, and high‑quality KinyaCOMET pairs. It outperforms existing multilingual embeddings on the SemRel2024‑rw benchmark, achieving a Spearman ψ of 0.7298, and introduces the Wiki‑RW‑STS benchmark of 300 contamination‑free Kinyarwanda sentence pairs. All model checkpoints, filtered pairs, and the new benchmark are publicly released.
By Ireddi Rakshitha, Devavarapu Yashwanth, Ntakirutimana Pierre
arXiv:2609.09974v1 Announce Type: new
Abstract: Grapheme-to-phoneme conversion (G2P) refers to the task of converting a sequence of graphemes to a corresponding sequence of phonemes. While Filipino G...
By Lorenz Bernard Marqueses, Paulo Grane Gabriel Silva, Chastine Cabatay, Ericson Adler Tan, Ann Franchesca Laguna