Audio large language models (Audio LLMs) often fail to transcribe English‑Mandarin code‑switching speech, exhibiting language omission, translation‑instead‑of‑transcription, and hallucination. By applying Direct Preference Optimization (DPO) with 100 K preference pairs, the models learn to preserve mixed‑language content rather than translate, leading to significant reductions in mixed‑error rates (up to 89.6% in‑distribution). The study demonstrates that DPO can effectively align multilingual Audio LLMs for accurate code‑switching transcription.
By Trung Nguyen Quang, Cheng Yi Lewis Won, Minh Duc Pham, Yingxu He, Shuo Sun, Ai Ti Aw
The paper investigates how to close the quality gap in low‑resource text‑to‑speech for Khmer and Korean using the VoxCPM2 model. By training a single low‑rank adaptation (LoRA) adapter on a shared 25.5‑hour corpus, the authors improve Khmer’s mean opinion score from 3.85 to 4.23 with a rank‑64 adapter, while Korean shows no significant gain. The study highlights that adaptation benefits mainly when the base model is weak and that training loss does not always align with human ratings.
By Phannet Pov, Hyun Woo Park, Voneat Pen, Sovandara Chhoun, Wan-Sup Cho, Saksonita Khoeurn
arXiv:2607. 23440v1 Announce Type: cross Abstract: In this paper, we push the boundary of LLM reasoning by testing them in a Chinese language game, xiehouyu, with novel xiehouyu created by linguists that had not existed before to avoid data contamination.
By Hai Hu, Siyuan Song, Chongtian Shao, Kejia Zhang, Tianjian Zhu, Xiaojing Zhao
A 2025 review of large language models, from DeepSeek R1 and RLVR to inference-time scaling, benchmarks, architectures, and predictions for 2026.
By Sebastian Raschka, PhD
arXiv:2607. 26751v1 Announce Type: cross Abstract: State-of-the-art intracortical brain-to-text systems pair a neural-sequence phone decoder with an external language model.
By Lucas Zamora Vera, Jose A. Gonzalez-Lopez
arXiv:2609.23825v1 Announce Type: new
Abstract: We present a comprehensive benchmark of Federated Learning (FL) for multilingual Automatic Speech Recognition (ASR), evaluating four Speech-LLM archite...
By Jordi Luque, Aleix Sant, Fernando L\'opez
arXiv:2603. 29042v2 Announce Type: replace-cross Abstract: Phone recognition (PR) is a key enabler of multilingual and low-resource speech processing tasks, yet robust performance remains elusive.
By Shikhar Bharadwaj, Chin-Jou Li, Kwanghee Choi, Eunjung Yeo, William Chen, Shinji Watanabe, David R. Mortensen
The paper introduces a new evaluation framework for phonetic encoding algorithms, using a generalized Rand Index called the Hüllermeier‑Rifqi Index. It measures discordance by comparing pairwise similarity scores of ground‑truth IPA transcriptions with those of encoded strings, adjusted against a random generator. The method is applied to multilingual datasets, assessing recall via collision rate and demonstrating its use in evaluating orthographic transparency.
By Can \"Ozbey, Emre Kaplan, Berkin Deniz Kahya
HuPER is a human-inspired framework that models phonetic perception as adaptive inference over acoustic‑phonetics evidence and linguistic knowledge. Using only 100 hours of training data, it achieves state‑of‑the‑art phonetic error rates on five English benchmarks and demonstrates strong zero‑shot transfer to 95 unseen languages. It uniquely enables adaptive, multi‑path phonetic perception across diverse acoustic conditions, and all training data, models, and code are open‑sourced.
By Chenxu Guo, Jiachen Lian, Yisi Liu, Baihe Huang, Shriyaa Narayanan, Bixing Wu, Zoe Ezzes, Jet Vonk, Zachary Miller, Cheol Jun Cho, Maria Gorno-Tempini, Gopala Anumanchipalli
arXiv:2608. 03970v1 Announce Type: new Abstract: Human input reaches language models by typing or speaking, and each channel leaves a distinct signature: orthographic noise for keyboards; for voice, disfluency from conventional transcription and restructuring from AI-backed dictation tools.
By Zizhao Hu, Nathan Elijah Segura, Mohammad Rostami, Jesse Thomason
Human input reaches language models by typing or speaking, and each channel leaves a distinct signature: orthographic noise for keyboards; for voice, disfluency from conventional transcription and restructuring from AI-backed dictation tools. How do they impact an LLM's performance?