VietPrism is a newly released, large‑scale Vietnamese speech corpus that combines 993.4 hours of real utterances from 1,262 verified speakers with 3.1 k hours of synthetic spoof speech. It uniquely offers transcripts, consistent speaker identities, five dialect groups, and extensive Vietnamese‑English code‑switching—nearly half of the corpus—while pairing each spoof with a matched bona fide utterance. The dataset enables controlled evaluation of deep‑fake detection models, revealing significant variability in detector performance across dialects and speaker similarity.
By Minh Hoang, Thai Le
The paper introduces a cross-dialect Named Entity Recognition (NER) framework for Bangla, leveraging the ANCHOLIK-NER dataset that covers five major regional dialects. Using a Leave-One-Dialect-Out Cross-Validation strategy, eight transformer-based models were evaluated, with Multilingual-E5 Large achieving the best performance (F1 up to 97.26% on Mymensingh, 82.38% on Chattogram). Local Interpretable Model-agnostic Explanations (LIME) revealed that the models rely mainly on the surface form of entity words rather than surrounding context, suggesting a direction for future improvement.
By Shamim Rahim Refat, Faika Fairuj Preotee, Shuvashis Sarker, Shifat Islam, Bidyarthi Paul, Mohammad Ashraful Hoque
The study investigates why language models exhibit systematic performance gaps across English dialects, a phenomenon termed the "dialect tax." Using parallel dialect corpora that preserve meaning while altering surface form, the authors confirm that models treat Standard American English and dialectal texts as semantically equivalent, yet find representational disparities that persist through tokenization, pre‑training, post‑training, and inference. Even a character‑level tokenizer does not eliminate input/output asymmetries or accuracy gaps, and dialect pairs produce more divergent gradient updates than unrelated Standard texts, indicating that dialectal content is harder for models to learn.
By Elle
The paper introduces a three-level evaluation framework—behavioral deployment, LM-head readout, and probe recoverability—to distinguish whether a language model fails a syntactic test by not encoding structure or by failing to use it. Using a trilingual control-dependency benchmark, the authors find that probe recoverability consistently exceeds LM-head readout, which in turn exceeds behavioral deployment across seven models and three languages, with the largest gap observed in Qwen3-0.6B Instruct. Layer-localized activation patching shows that instruction tuning shifts the decoded layer later, suggesting decoding favors surface shortcuts and that behavioral evaluation understates what models encode while probing alone overstates what they deploy.
By Zhenyan Lu, He Wang, Xiaohui Huang
arXiv:2606. 01016v1 Announce Type: cross Abstract: While End-to-End (E2E) Speech-Large Language Models (Speech-LLMs) are rapidly evolving, their evaluation methodologies remain limited to the era of simple transcription.
By Sicheng Yang, Shulan Ruan, Shiwei Wu, Yu Liu, Lu Fan, Zhi Li, You He
arXiv:2601. 22888v4 Announce Type: replace-cross Abstract: More than 80% of the 1.
By Jio Oh, Paul Vicinanza, Thomas Butler, Steven Euijong Whang, Dezhi Hong, Amani Namboori
arXiv:2610.09152v1 Announce Type: new
Abstract: Multilingual LLM benchmarks omit Slovak, a morphologically rich West Slavic language of five million speakers, or cover it only by machine translation....
By Marek \v{S}uppa, Ivan Vykopal, Andrej Ridzik, Kristi\'an Sopkovi\v{c}, Nat\'alia K\v{n}a\v{z}ekov\'a, Jaroslav Kop\v{c}an, Miroslav Bl\v{s}t\'ak, Vikt\'oria Ondrejov\'a, Daniel Hl\'adek, Michal Gregor, Martin Tamajka, Mari\'an \v{S}imko
arXiv:2609.18284v1 Announce Type: new
Abstract: In recent years, three initiatives have emerged to develop generative language models in Hungary. The motivation behind them is the same. For Hungarian...
By M\'aty\'as Osv\'ath, Enik\H{o} H\'eja, No\'emi Ligeti-Nagy
LuxIT is a monolingual instruction‑tuning dataset for Luxembourgish, created by synthesizing instruction‑answer pairs from native texts using the DeepSeek‑R1‑0528 model and a quality‑assurance LLM‑as‑judge process. The resulting 227,507 high‑quality pairs were used to fine‑tune 14 LLMs (≤15 B parameters), yielding an average accuracy increase of +5.37 percentage points on standardized Luxembourgish proficiency exams and improvements in macro‑averaged F1 on nine of the fourteen downstream NLP tasks. These findings demonstrate that synthetic monolingual data can effectively enhance LLM performance in low‑resource languages and reveal the complex relationship between exam performance and practical NLP gains.
By Julian Valline, Cedric Lothritz, Siwen Guo, Jordi Cabot
arXiv:2608. 03803v1 Announce Type: cross Abstract: Multilingual language models are deployed across a hundred or more languages, yet most benchmarks test whether a model can perform a task _in_ a language rather than whether it commands the language itself, conflating fluency with proficiency.
By Tom\'a\v{s} Burkert, Angelika Peljak-{\L}api\'nska, David Zelen\'y
The paper introduces Latent Space Refusal Anchoring (LSR‑Anchoring), a training‑free technique that extracts a refusal direction from English prompts and applies it to the residual stream of instruction‑tuned models at inference time. The primary variant, Mean‑Activation Steering (MAS), works across several architectures (Llama‑3‑8B, Llama‑3.1‑70B, Mistral‑7B‑Instruct, Qwen2.5‑7B), restoring safety for low‑resource African languages with minimal performance loss, while a refined SAE‑Derived Steering (SDS) further reduces KL divergence without degrading legitimate prompt performance. The method shows positive transfer for Yoruba, Igbo, Igala, and Hausa, but fails for Arabic, suggesting a geometric mismatch rather than a data scarcity issue.
By Godwin Abuh Faruna
arXiv:2609.36214v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly used by people whose first language is not English, yet these users have been shown to receive systematic...
By Yusheng Zhou, Eleanor Lin, David Jurgens