arXiv Computation and Language
Sep 24

Long-Tail Rebalancing for Non-Verbal Vocalization-Aware ASR: A Track 1 System for the NVVSpeech Challenge

The paper presents a data‑centric approach to improve automatic speech recognition for non‑verbal vocalizations (NVVs) in the ISCSLP NVVSpeech Challenge. It introduces cross‑dataset label harmonization and a two‑stage sampling schedule—first square‑root category sampling to address long‑tailed distributions, then uniform‑category fine‑tuning—to jointly transcribe lexical content and 16 NVV categories. The final system achieved an official score of 63.86, ranking fourth in Track 1.

By Shangyue Jia, Jingru Ma, Yangzhuo Li, Daoping Luo, Bowen Tian, Hanchen Lu, Wenze Ren, Yunxiang Chen, Houdun Liu, Su Feng, Lei Xie, Liumeng Xue
Hugging Face Trending Papers
Jun 2

Efficient ASR Training with Conversations that Never Happened

Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that generates scenario-level dialogues with participant metadata, maps speaker attributes to TTS voice profiles, and assembles synthesized utterances into speaker-aware simulated conversations.

arXiv AI
Aug 19

Language Family Matters: Evaluating LLM-Based ASR Across Linguistic Boundaries

The paper introduces a new strategy for connecting large language models (LLMs) to speech encoders in automatic speech recognition (ASR) systems by sharing a single connector across languages within the same linguistic family. This approach reduces the number of parameters needed compared to training a separate connector for each language, while improving generalization across different domains and real‑world corpora. Experiments with two multilingual LLMs and two speech datasets demonstrate that family‑based connectors are both efficient and effective for multilingual ASR deployment.

By Yuchen Zhang, Ravi Shekhar, Haralambos Mouratidis