arXiv AI By Kaavya Chaparala, Su Huang, Stephen L. Miller, Rhiannon N. Miller, Anjalie Field

Pretrained ASR Pseudo-labeling for Noisy Police Audio

Read the original on arXiv AI →

The paper evaluates pseudo‑labeling to adapt pretrained ASR models (Whisper and Qwen3‑ASR) for noisy Broadcast Police Communication (BPC) from Baltimore and Chicago. It finds that internal confidence metrics cannot reliably separate high‑ and low‑quality pseudo‑labels, and proposes an external LLM‑as‑a‑judge filtering approach that more aggressively removes implausible transcripts, reducing WER. Additionally, a cross‑model pseudo‑labeling strategy is introduced, where one model is fine‑tuned with pseudo‑labels from the other, showing promise for future work.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 22

Long-Tail Rebalancing for Non-Verbal Vocalization-Aware ASR: A Track~1 System for the NVVSpeech Challenge

arXiv:2609.23462v1 Announce Type: cross Abstract: Non-verbal vocalizations (NVVs) carry important paralinguistic information but are often omitted by conventional automatic speech recognition (ASR) s...

By Shangyue Jia, Jingru Ma, Yangzhuo Li, Daoping Luo, Bowen Tian, Hanchen Lu, Wenze Ren, Yunxiang Chen, Houdun Liu, Shuo Feng, Lei Xie, Liumeng Xue
arXiv Computation and Language
Sep 24

Long-Tail Rebalancing for Non-Verbal Vocalization-Aware ASR: A Track 1 System for the NVVSpeech Challenge

The paper presents a data‑centric approach to improve automatic speech recognition for non‑verbal vocalizations (NVVs) in the ISCSLP NVVSpeech Challenge. It introduces cross‑dataset label harmonization and a two‑stage sampling schedule—first square‑root category sampling to address long‑tailed distributions, then uniform‑category fine‑tuning—to jointly transcribe lexical content and 16 NVV categories. The final system achieved an official score of 63.86, ranking fourth in Track 1.

By Shangyue Jia, Jingru Ma, Yangzhuo Li, Daoping Luo, Bowen Tian, Hanchen Lu, Wenze Ren, Yunxiang Chen, Houdun Liu, Su Feng, Lei Xie, Liumeng Xue
arXiv AI
Aug 24

Do SpeechLMs Hear Their Own Opinions? Diagnosing and Mitigating Previous-Belief Contamination in Streaming Emotion Understanding

The paper investigates how streaming emotion recognition models can be misled by their own prior predictions, a problem termed previous-belief contamination (PBC). Using a counterfactual diagnostic on CREMA-D-Stream, the authors show that feeding a model’s previous emotion label into its current prediction can drastically lower accuracy and flip many predictions, with the effect varying by label. To mitigate PBC, they propose EmoUpdate, a training‑free framework that isolates current audio perception from historical context through a prior‑blind firewall, a causal belief filter, and a decontamination operator, achieving significant gains across multiple SpeechLMs and benchmarks.

By Haoyue Liu, Zhichao Wang, Ye Chen, Haonan Deng, Xiaoying Tang
arXiv Computation and Language
Sep 28

Asymmetric Classifier-Free Guidance for Target-Speaker ASR

The paper introduces asymmetric classifier‑free guidance (CFG) for target‑speaker ASR using Whisper, where a speaker‑conditioned branch predicts the target transcript and a speaker‑unconditioned branch predicts serialized multi‑speaker transcripts. CFG modulates the influence of speaker conditioning during decoding via a single guidance scale, which is first set globally on development data and then refined per utterance by a lightweight encoder‑based predictor while keeping the recognition model fixed. The resulting system yields up to 21.8% relative WER reduction over a condition‑only baseline and 5.6% over standard conditional decoding under domain shifts.

By Yiwen Guan, Jacob Whitehill