arXiv Computation and Language By Joanna Rado{\l}a, Josep Maria Crego, Fran\c{c}ois Yvon

Improving Language Identification for Code-Switched Utterances with Integer Linear Programming

Read the original on arXiv Computation and Language →

The paper addresses the challenge of identifying code‑switched utterances in language identification systems. It revisits the MaskLID approach, highlighting its overreliance on word‑level language association scores, and reformulates its optimization as an Integer Linear Program to incorporate clear, interpretable constraints. These enhancements significantly improve performance across ten diverse languages on code‑switching benchmarks, with the authors releasing code and data for reproducibility.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Aug 28

AfriSwitch: A Benchmark for In-the-Wild African Code-Switched Speech Recognition

AfriSwitch is a 61.36‑hour, human‑transcribed benchmark of in‑the‑wild code‑switched speech covering 16 African languages and varieties, annotated with switch‑level English span tags, per‑utterance Code‑Mixing Index (CMI), and switch‑point counts. The corpus reveals that code‑switching behaviour varies widely across languages, with no single metric fully capturing how code‑switched a language is. Benchmarking five open and commercial multilingual ASR systems in a zero‑shot setting shows high word error rates, with the best system averaging 35.93% WER and none dropping below 24% on any language, indicating that Africa‑targeted training rather than model scale or nominal language coverage best predicts performance.

By Gabrial Zencha Ashungafac, Busayo Awobade, Tobi Olatunji
arXiv Computation and Language
2d ago

Beyond Word Error Rate: A Switch Aware Evaluation of ASR and Audio Language Models on English Yoruba Code-Switched Speech

arXiv:2609. 11786v1 Announce Type: new Abstract: Automatic speech recognition (ASR) systems and audio language models (audio LMs) now report low error rates on monolingual benchmarks, but their behavior on code switched speech in low resource, diacritic rich languages remains poorly characterized.

By Chibuzor Okocha, Christan Earl Grant
arXiv Computation and Language
2d ago

IndicTriMix: Developing Language Identification Datasets and Models for Tri-Language Code-Mixing

IndicTriMix presents a new benchmark and models for token‑level language identification in tri‑language code‑mixed text involving Hindi, Gujarati, and Bengali. The authors reformulate the task as sequence labeling and fine‑tune transformer models MuRIL and XLM‑RoBERTa, evaluating them on manually annotated test sets. They also introduce two code‑mixing generation methods using parallel sentences and release the datasets and fine‑tuned models for public use.

By Pruthwik Mishra, Rudra Trivedi, Avi Patel, Ashok Urlana, Shrikant Malviya
arXiv Computation and Language
Sep 2

Phrase-Localized Language-Contrastive Guidance: Training-Free Localized Accent Control for Code-Switching Text-to-Speech

arXiv:2609.01016v1 Announce Type: new Abstract: Current speech synthesis struggles with code-switching, which mixes a foreign language phrase into a primary language utterance, causing the phrase to...

By Che Hyun Lee, Sangkwon Park, Donghun Kang, Dongwook Lee, Youngho Cho, Heeseung Kim, Sungroh Yoon
arXiv Computation and Language
Sep 2

Latent Mechanisms of Language Control in Multilingual Language Models

The paper investigates how multilingual large language models can unintentionally switch languages during generation. It compares three techniques—ValSel, FreqSel, and AnnSel—for pinpointing latent variables that control language choice in cross‑layer transcoders. Using new multilingual benchmarks and targeted interventions on Gemma‑2‑2B and Qwen3‑4B, the study finds all methods can steer output language, with FreqSel performing best and AnnSel providing interpretable selections via explicit annotations.

By Ryo Mitsuhashi, Sabri Boughorbel, Majd Hawasly