arXiv Computation and Language By Minh Hoang, Thai Le

VietPrism: A large-scale Vietnamese speech and deepfake corpus with diverse dialects and code-switching

Read the original on arXiv Computation and Language →

VietPrism is a newly released, large‑scale Vietnamese speech corpus that combines 993.4 hours of real utterances from 1,262 verified speakers with 3.1 k hours of synthetic spoof speech. It uniquely offers transcripts, consistent speaker identities, five dialect groups, and extensive Vietnamese‑English code‑switching—nearly half of the corpus—while pairing each spoof with a matched bona fide utterance. The dataset enables controlled evaluation of deep‑fake detection models, revealing significant variability in detector performance across dialects and speaker similarity.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 24

Lost in Speech: Trilingual Spoken Hallucination Detection Across Audio and Transcripts

The paper introduces a trilingual spoken hallucination detection benchmark covering English, Russian, and Kazakh news, with 12,013 samples that include synthetic alterations and severity levels, as well as 290 fact‑checked misinformation items. Detectors are evaluated in a reference‑free setting on text, ASR transcripts, and audio, revealing that most models underperform baseline classifiers, except Gemma‑3n on transcripts. Synthetic‑trained detectors achieve high macro‑F1 scores on real‑world misinformation, but Russian provenance analysis highlights model‑dependent signals that confound synthetic benchmarks.

By Meruyert Aristombayeva, Jason S. Lucas, Chaewan Chun, Dongwon Lee
arXiv Computation and Language
3d ago

Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs

Audio large language models (Audio LLMs) often fail to transcribe English‑Mandarin code‑switching speech, exhibiting language omission, translation‑instead‑of‑transcription, and hallucination. By applying Direct Preference Optimization (DPO) with 100 K preference pairs, the models learn to preserve mixed‑language content rather than translate, leading to significant reductions in mixed‑error rates (up to 89.6% in‑distribution). The study demonstrates that DPO can effectively align multilingual Audio LLMs for accurate code‑switching transcription.

By Trung Nguyen Quang, Cheng Yi Lewis Won, Minh Duc Pham, Yingxu He, Shuo Sun, Ai Ti Aw
arXiv Machine Learning
Aug 12

How Robust Are LLMs to Vietnamese Dialects?

arXiv:2608. 10414v1 Announce Type: cross Abstract: Large Language Models (LLMs) are typically evaluated on standard written Vietnamese, yet everyday communication frequently involves regional dialects that preserve meaning but differ in surface form.

By Minh Tran, Trinh Chau, Thanh-Nhan Le, Nam Tran, Luan Thanh Nguyen, Cuong Dang, Duc Hoang