Agentic-GER is an LLM-based agent designed to improve terminology accuracy in long‑form speech transcription. It leverages global context from the entire transcript to flag suspicious terms, selectively re‑transcribes the source audio to verify candidate corrections, and uses accepted edits to inform future decisions. Experiments on GigaSpeechBench with four LLMs and two ASR systems show consistent terminology improvements in both Chinese and English, achieving up to a 36.8% relative reduction in biased character error rate over the Whisper baseline for Chinese speech.
By Yanqiao Zhu, Wupeng Wang, Zhifu Gao, Xiangang Li, Xie Chen
The paper introduces LOGIC (Logit‑Space Integration for Contextual Biasing), a new framework that injects contextual entity information directly into the decoding layer of Speech Large Language Models, bypassing the limitations of prompt‑based methods. LOGIC operates with constant‑time complexity regardless of the size of the entity list, and experiments with the Phi‑4‑MM model across 11 multilingual locales show an average 9% relative reduction in Entity WER while adding only a 0.30% increase in False Alarm Rate.
By Peidong Wang, Jian Xue, Jinyu Li
The paper introduces Hybrid Search, a method that refines warm-initialized large language model (LLM) based automatic speech recognition (ASR) systems by exploiting interactions between ASR hidden states and the base LLM’s hidden states. By identifying tokens with high semantic dependence and selectively correcting them, the approach surpasses traditional global LLM‑correction techniques such as rescoring and late fusion. The study demonstrates that even after warm initialization, LLM‑based ASR models can further benefit from their base LLM during inference.
By Chan-Jan Hsu, Jaeyeon Kim, Chao-Han Huck Yang, Shinji Watanabe, Hung-yi Lee, Carlos Busso
The paper introduces KVoiceBench, KOpenAudioBench, and KMMAU—three Korean speech benchmarks created through agent-driven frameworks that adapt existing SpokenQA and ASR resources into Korean SpokenQA and audio understanding tasks. These benchmarks total 12,345 samples and are publicly released to evaluate SpeechLMs beyond English. The authors benchmark eight recent SpeechLMs, revealing significant English‑Korean performance gaps and divergent rankings between SpokenQA and audio understanding, highlighting multilingual weaknesses not apparent in English-only tests.
By Haechan Kim, Seungjun Chung, Inkyu Park, Jihoo Lee, Jonghyun Lee
PTC-Bias is a two-stage framework that improves contextual biasing in speech large language models by using phoneme-level temporal competition. In the first stage, PTC Retrieval performs frame-synchronous phoneme decoding to generate a compact shortlist of bias words and their speech intervals. The second stage, PTC Correction, applies a local competition between retrieved candidates and mismatched transcript spans within those intervals, reducing near-homophone and word-segmentation errors without extra SpeechLLM passes. Experiments on LibriSpeech demonstrate consistent gains across two SpeechLLMs, with PTC-Bias reducing B-WER by up to 23.9% relative to CTC-Filter while keeping U-WER nearly unchanged.
By Zhiqi Ai, Han Cheng, Shiyi Mu, Yongjin Zhou, Shugong Xu
Text-to-speech systems increasingly process user-generated text (UGT) such as ppl and imo, whose pronunciation must be inferred from the canonical rather than surface form. We introduce UGTPhon, the f...
LOGIC is a framework that performs contextual biasing for speech large language models by integrating directly in logit space. It decouples context injection from input processing, allowing explicit control over biasing strength and reducing entity word error rates by an average of 9% relative to baseline methods. When combined with prompting, LOGIC further lowers entity word error rates by 5% relative to prompt-only approaches, with only a modest 2.8% runtime overhead.
By Peidong Wang, Jian Xue, Jinyu Li
arXiv:2609.09889v1 Announce Type: new
Abstract: Automatic Speech Recognition (ASR) technology is fundamental to customer service automation and large-scale transcription. However, even advanced ASR m...
By Yonghyun Jun, Jimin Lee, Hwan Chang, Dongho Shin, Seolah Kim, Hwanhee Lee
We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory. md and decides per utterance whether to act on the hypothesis or abstain and keep the 1-best.
Dual-Form ASR (DF-ASR) is a framework that unifies spoken-form ASR and semantics-aware written-form inverse text normalization (ITN) for Chinese speech recognition. It uses paired spoken- and written-form supervision generated and judged by a large language model, and introduces an ITN-MWER objective to penalize errors on normalization-sensitive spans. DF-ASR also employs a REQUIRE-ITN/FORBID-ITN protocol to separately evaluate required normalization and forbidden-span preservation, achieving superior performance over open-source ASR-ITN systems while maintaining prompt-level control between transcript forms.
By Fengrun Zhang, Li Fu, Wangjin Zhou, Lu Fan, Youzheng Wu, Xiaodong He
arXiv:2605. 19266v2 Announce Type: replace-cross Abstract: Automatic speech recognition (ASR) systems are typically optimized for verbatim transcription, which preserves disfluencies, filler words, and informal spoken structures that are often unsuitable for downstream writing-oriented applications.
By Wanyi Ning, Yinshang Guo, Haitao Qian, Jiyuan Cheng, Weiyuan Feng, Yufei Zhang
The paper introduces UGTPhon, a grapheme-to-phoneme benchmark for user‑generated text in English, Vietnamese, and Korean, and presents a taxonomy for diagnosing pronunciation errors. It shows that existing G2P models and large language models struggle with canonical‑to‑non‑canonical text, with errors up to 66.8 PER points. A compositional G2P approach that uses exact‑match lookup and staged decoding reduces these errors and performs competitively with larger few‑shot LLMs.
By MinJu Jeon, Younghan Park, Han Sung Park, Jong-Hwan Kim, Dong-Jin Kim, Hoyeon Lee