arXiv Computation and Language By Shangyue Jia, Jingru Ma, Yangzhuo Li, Daoping Luo, Bowen Tian, Hanchen Lu, Wenze Ren, Yunxiang Chen, Houdun Liu, Su Feng, Lei Xie, Liumeng Xue

Long-Tail Rebalancing for Non-Verbal Vocalization-Aware ASR: A Track 1 System for the NVVSpeech Challenge

Read the original on arXiv Computation and Language →

The paper presents a data‑centric approach to improve automatic speech recognition for non‑verbal vocalizations (NVVs) in the ISCSLP NVVSpeech Challenge. It introduces cross‑dataset label harmonization and a two‑stage sampling schedule—first square‑root category sampling to address long‑tailed distributions, then uniform‑category fine‑tuning—to jointly transcribe lexical content and 16 NVV categories. The final system achieved an official score of 63.86, ranking fourth in Track 1.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 22

Long-Tail Rebalancing for Non-Verbal Vocalization-Aware ASR: A Track~1 System for the NVVSpeech Challenge

arXiv:2609.23462v1 Announce Type: cross Abstract: Non-verbal vocalizations (NVVs) carry important paralinguistic information but are often omitted by conventional automatic speech recognition (ASR) s...

By Shangyue Jia, Jingru Ma, Yangzhuo Li, Daoping Luo, Bowen Tian, Hanchen Lu, Wenze Ren, Yunxiang Chen, Houdun Liu, Shuo Feng, Lei Xie, Liumeng Xue
Hugging Face Trending Papers
Jun 2

Efficient ASR Training with Conversations that Never Happened

Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that generates scenario-level dialogues with participant metadata, maps speaker attributes to TTS voice profiles, and assembles synthesized utterances into speaker-aware simulated conversations.

arXiv Computation and Language
Sep 25

A Training Criterion with Token-Level Tolerance to Transcription Ambiguity for Automatic Speech Recognition

The paper introduces a token‑level extension of Omni‑Temporal Classification (OTC) for automatic speech recognition, allowing unsupported tokens to be bypassed while preserving supervision for the rest of the word. Across 19 languages and three corpora, this token‑level OTC consistently outperforms standard CTC, achieving the lowest mean word error rate on every dataset and a 9.45% average relative WER reduction. A predictive‑entropy‑indexed schedule replaces epoch‑based relaxation, reducing training‑length dependence while maintaining performance.

By Saurabh Kumar, Diptiman Mohanta, Prasanta Kumar Ghosh