arXiv Computation and Language

HearInContext: A Benchmark for Implicit Context in Speech Recognition

arXiv AI
Sep 2

VoiceLongMemEval: Do Assistants Remember How You Sounded?

VoiceLongMemEval (VLME) is a new benchmark that tests AI assistants on their ability to remember how users sounded by incorporating paralinguistic metadata—such as emotion labels, prosody descriptors, and voice events—into each conversational turn. The benchmark uses a three‑stage adversarial gate to ensure that models cannot succeed with transcript alone, revealing a significant affect gap: models gain 0.09 to 0.38 accuracy when provided with paralinguistic cues, and audio‑native models outperform standard ASR pipelines in extracting these signals. The dataset and code will be released upon acceptance.

By Ramit Pahwa, Parivesh Priye, Apoorva Beedu
arXiv Computation and Language
5d ago

PTC-Bias: Phoneme-Level Temporal Competition for Bias Retrieval and Post-Decoding Correction in Speech LLMs

PTC-Bias is a two-stage framework that improves contextual biasing in speech large language models by using phoneme-level temporal competition. In the first stage, PTC Retrieval performs frame-synchronous phoneme decoding to generate a compact shortlist of bias words and their speech intervals. The second stage, PTC Correction, applies a local competition between retrieved candidates and mismatched transcript spans within those intervals, reducing near-homophone and word-segmentation errors without extra SpeechLLM passes. Experiments on LibriSpeech demonstrate consistent gains across two SpeechLLMs, with PTC-Bias reducing B-WER by up to 23.9% relative to CTC-Filter while keeping U-WER nearly unchanged.

By Zhiqi Ai, Han Cheng, Shiyi Mu, Yongjin Zhou, Shugong Xu
arXiv AI
Sep 4

Listen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State Interactions

The paper introduces Hybrid Search, a method that refines warm-initialized large language model (LLM) based automatic speech recognition (ASR) systems by exploiting interactions between ASR hidden states and the base LLM’s hidden states. By identifying tokens with high semantic dependence and selectively correcting them, the approach surpasses traditional global LLM‑correction techniques such as rescoring and late fusion. The study demonstrates that even after warm initialization, LLM‑based ASR models can further benefit from their base LLM during inference.

By Chan-Jan Hsu, Jaeyeon Kim, Chao-Han Huck Yang, Shinji Watanabe, Hung-yi Lee, Carlos Busso
arXiv Computation and Language
Aug 28

Scaling phoneme-based TTS augmentation for ASR: A unified pipeline and controlled study

The paper introduces a unified phoneme‑based TTS‑to‑ASR augmentation pipeline that uses a multilingual TTS model with language‑ID conditioning and incorporates grapheme‑to‑phoneme conversion, reference‑speech filtering, and candidate‑text selection. It proposes phoneme‑frequency‑guided selection (PFGS) to rank sentences based on phoneme frequencies from real ASR labels, and demonstrates that random augmentation and PFGS both improve ASR performance across Arabic, French, Italian, and Portuguese test sets, with PFGS yielding up to a 19.3% relative WER reduction. The study also shows that filtering reference speech can further lower WER by up to 0.59 points on certain datasets.

By Zhen Wang, TianRui Wu, RongQi Han, Hao Wu, Wei Liang
arXiv AI
Aug 25

Unlearning Is Not Just Erasing: Temporal Decoupling via Generation Inequality

The paper introduces ADU, a fine‑grained training framework that unlearns sensitive information from large language models by decoupling contextual attention pathways instead of erasing tokens. ADU exploits the distinction between local and global attention heads to identify and suppress attention paths that retrieve persistent sensitive anchors, while preserving local‑attention structure and overall language modeling performance. Evaluation on the TOFU and WMDP benchmarks shows ADU achieves superior forget quality (0.93 on TOFU) and retains 92.9% of model utility compared to 81.9% for existing baselines, with fewer side effects in benign contexts.

By Xunlei Chen, Qirui Ye, Yuang Li, Yi Gong, Zhaokun Wang, Wenyi Li, Shiyao Guo, Jinyu Guo
arXiv Computation and Language
Sep 14

Not All Speech Is Intent: Adaptive Self-Correcting Inference Layer for Post-ASR False Wake-Up

The paper introduces ASCIL, a post‑ASR correction framework that re‑evaluates wake‑up intent by combining acoustic embeddings, linguistic cues, device context, and past misclassifications. ASCIL interprets both implicit (hesitation, disengagement, silence) and explicit (cancellation, repetition) signals as noisy indicators of misclassification, enabling online pattern updates without manual annotation. On a proprietary dataset of 3,667 interactions, ASCIL reduces errors by up to 54.27% relative on a session‑disjoint subset and 24.39% at a 0.90 threshold, while adding less than 60 ms of latency and improving intentional acceptance rates.

By Preeti Saraswat, Divya Neelagiri, Anil Yadav