arXiv AI

CyrillicQA: The Influence of Phonetically Encoded Secret Language on LLM Performance

arXiv Computation and Language
Sep 28

Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs

Audio large language models (Audio LLMs) often fail to transcribe English‑Mandarin code‑switching speech, exhibiting language omission, translation‑instead‑of‑transcription, and hallucination. By applying Direct Preference Optimization (DPO) with 100 K preference pairs, the models learn to preserve mixed‑language content rather than translate, leading to significant reductions in mixed‑error rates (up to 89.6% in‑distribution). The study demonstrates that DPO can effectively align multilingual Audio LLMs for accurate code‑switching transcription.

By Trung Nguyen Quang, Cheng Yi Lewis Won, Minh Duc Pham, Yingxu He, Shuo Sun, Ai Ti Aw
arXiv Computation and Language
Sep 25

Closing the Quality Gap in Low-Resource Text-to-Speech: LoRA Fine-Tuning of VoxCPM2 for Khmer and Korean

The paper investigates how to close the quality gap in low‑resource text‑to‑speech for Khmer and Korean using the VoxCPM2 model. By training a single low‑rank adaptation (LoRA) adapter on a shared 25.5‑hour corpus, the authors improve Khmer’s mean opinion score from 3.85 to 4.23 with a rank‑64 adapter, while Korean shows no significant gain. The study highlights that adaptation benefits mainly when the base model is weak and that training loss does not always align with human ratings.

By Phannet Pov, Hyun Woo Park, Voneat Pen, Sovandara Chhoun, Wan-Sup Cho, Saksonita Khoeurn
arXiv Machine Learning
Jul 14

An Empirical Recipe for Universal Phone Recognition

arXiv:2603. 29042v2 Announce Type: replace-cross Abstract: Phone recognition (PR) is a key enabler of multilingual and low-resource speech processing tasks, yet robust performance remains elusive.

By Shikhar Bharadwaj, Chin-Jou Li, Kwanghee Choi, Eunjung Yeo, William Chen, Shinji Watanabe, David R. Mortensen
arXiv Computation and Language
Sep 7

Evaluation of Phonetic Encoding Algorithms on Transcription Datasets

The paper introduces a new evaluation framework for phonetic encoding algorithms, using a generalized Rand Index called the Hüllermeier‑Rifqi Index. It measures discordance by comparing pairwise similarity scores of ground‑truth IPA transcriptions with those of encoded strings, adjusted against a random generator. The method is applied to multilingual datasets, assessing recall via collision rate and demonstrating its use in evaluating orthographic transparency.

By Can \"Ozbey, Emre Kaplan, Berkin Deniz Kahya
arXiv AI
Sep 28

HuPER: A Human-Inspired Framework for Phonetic Perception

HuPER is a human-inspired framework that models phonetic perception as adaptive inference over acoustic‑phonetics evidence and linguistic knowledge. Using only 100 hours of training data, it achieves state‑of‑the‑art phonetic error rates on five English benchmarks and demonstrates strong zero‑shot transfer to 95 unseen languages. It uniquely enables adaptive, multi‑path phonetic perception across diverse acoustic conditions, and all training data, models, and code are open‑sourced.

By Chenxu Guo, Jiachen Lian, Yisi Liu, Baihe Huang, Shriyaa Narayanan, Bixing Wu, Zoe Ezzes, Jet Vonk, Zachary Miller, Cheol Jun Cho, Maria Gorno-Tempini, Gopala Anumanchipalli
arXiv AI
Aug 5

Should We Type or Talk to LLM Agents? A Comprehensive Study of Voice and Keyboard Input Perturbations

arXiv:2608. 03970v1 Announce Type: new Abstract: Human input reaches language models by typing or speaking, and each channel leaves a distinct signature: orthographic noise for keyboards; for voice, disfluency from conventional transcription and restructuring from AI-backed dictation tools.

By Zizhao Hu, Nathan Elijah Segura, Mohammad Rostami, Jesse Thomason