arXiv AI

agentic-ger: terminology recovery in long-form speech using global context

Agentic-GER is an LLM-based agent designed to improve terminology accuracy in long‑form speech transcription. It leverages global context from the entire transcript to flag suspicious terms, selectively re‑transcribes the source audio to verify candidate corrections, and uses accepted edits to inform future decisions. Experiments on GigaSpeechBench with four LLMs and two ASR systems show consistent terminology improvements in both Chinese and English, achieving up to a 36.8% relative reduction in biased character error rate over the Whisper baseline for Chinese speech.

Hugging Face Trending Papers
6d ago

agentic-ger: terminology recovery in long-form speech using global context

Agentic-GER is an LLM-based agent designed to correct terminology in long‑form speech transcripts. It leverages global context from the full transcript to flag suspicious terms, selectively re‑transcribes the source audio to verify candidate corrections, and uses accepted edits to inform future decisions. Experiments on GigaSpeechBench with four LLMs and two ASR systems show consistent terminology improvements in both Chinese and English, achieving up to a 36.8% relative reduction in biased character error rate over the Whisper baseline.

arXiv Computation and Language
Sep 11

Beyond Prompting: Efficient and Robust Contextual Biasing for Speech LLMs via Logit-Space Integration (LOGIC)

The paper introduces LOGIC (Logit‑Space Integration for Contextual Biasing), a new framework that injects contextual entity information directly into the decoding layer of Speech Large Language Models, bypassing the limitations of prompt‑based methods. LOGIC operates with constant‑time complexity regardless of the size of the entity list, and experiments with the Phi‑4‑MM model across 11 multilingual locales show an average 9% relative reduction in Entity WER while adding only a 0.30% increase in False Alarm Rate.

By Peidong Wang, Jian Xue, Jinyu Li
arXiv AI
Sep 1

KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs

The paper introduces KVoiceBench, KOpenAudioBench, and KMMAU—three Korean speech benchmarks created through agent-driven frameworks that adapt existing SpokenQA and ASR resources into Korean SpokenQA and audio understanding tasks. These benchmarks total 12,345 samples and are publicly released to evaluate SpeechLMs beyond English. The authors benchmark eight recent SpeechLMs, revealing significant English‑Korean performance gaps and divergent rankings between SpokenQA and audio understanding, highlighting multilingual weaknesses not apparent in English-only tests.

By Haechan Kim, Seungjun Chung, Inkyu Park, Jihoo Lee, Jonghyun Lee
arXiv AI
Sep 4

Listen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State Interactions

The paper introduces Hybrid Search, a method that refines warm-initialized large language model (LLM) based automatic speech recognition (ASR) systems by exploiting interactions between ASR hidden states and the base LLM’s hidden states. By identifying tokens with high semantic dependence and selectively correcting them, the approach surpasses traditional global LLM‑correction techniques such as rescoring and late fusion. The study demonstrates that even after warm initialization, LLM‑based ASR models can further benefit from their base LLM during inference.

By Chan-Jan Hsu, Jaeyeon Kim, Chao-Han Huck Yang, Shinji Watanabe, Hung-yi Lee, Carlos Busso
arXiv Computation and Language
5d ago

PTC-Bias: Phoneme-Level Temporal Competition for Bias Retrieval and Post-Decoding Correction in Speech LLMs

PTC-Bias is a two-stage framework that improves contextual biasing in speech large language models by using phoneme-level temporal competition. In the first stage, PTC Retrieval performs frame-synchronous phoneme decoding to generate a compact shortlist of bias words and their speech intervals. The second stage, PTC Correction, applies a local competition between retrieved candidates and mismatched transcript spans within those intervals, reducing near-homophone and word-segmentation errors without extra SpeechLLM passes. Experiments on LibriSpeech demonstrate consistent gains across two SpeechLLMs, with PTC-Bias reducing B-WER by up to 23.9% relative to CTC-Filter while keeping U-WER nearly unchanged.

By Zhiqi Ai, Han Cheng, Shiyi Mu, Yongjin Zhou, Shugong Xu
arXiv Computation and Language
Sep 4

Dual-Form ASR: Semantics-Aware Inverse Text Normalization for Chinese Speech Recognition

Dual-Form ASR (DF-ASR) is a framework that unifies spoken-form ASR and semantics-aware written-form inverse text normalization (ITN) for Chinese speech recognition. It uses paired spoken- and written-form supervision generated and judged by a large language model, and introduces an ITN-MWER objective to penalize errors on normalization-sensitive spans. DF-ASR also employs a REQUIRE-ITN/FORBID-ITN protocol to separately evaluate required normalization and forbidden-span preservation, achieving superior performance over open-source ASR-ITN systems while maintaining prompt-level control between transcript forms.

By Fengrun Zhang, Li Fu, Wangjin Zhou, Lu Fan, Youzheng Wu, Xiaodong He
arXiv AI
5d ago

LOGIC: Efficient and Robust Contextual Biasing for Speech LLMs via Logit-Space Integration

LOGIC is a framework that performs contextual biasing for speech large language models by integrating directly in logit space. It decouples context injection from input processing, allowing explicit control over biasing strength and reducing entity word error rates by an average of 9% relative to baseline methods. When combined with prompting, LOGIC further lowers entity word error rates by 5% relative to prompt-only approaches, with only a modest 2.8% runtime overhead.

By Peidong Wang, Jian Xue, Jinyu Li
arXiv AI
Jun 9

FormalASR: End-to-End Spoken Chinese to Formal Text

arXiv:2605. 19266v2 Announce Type: replace-cross Abstract: Automatic speech recognition (ASR) systems are typically optimized for verbatim transcription, which preserves disfluencies, filler words, and informal spoken structures that are often unsuitable for downstream writing-oriented applications.

By Wanyi Ning, Yinshang Guo, Haitao Qian, Jiyuan Cheng, Weiyuan Feng, Yufei Zhang