arXiv AI By Lucas Zamora Vera, Jose A. Gonzalez-Lopez

Phoneme- vs. Character-Level Targets and Selective State-Space Models for Intracortical Brain-to-Text

Read the original on arXiv AI →

arXiv:2607. 26751v1 Announce Type: cross Abstract: State-of-the-art intracortical brain-to-text systems pair a neural-sequence phone decoder with an external language model.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 25

PTC-Bias: Phoneme-Level Temporal Competition for Bias Retrieval and Post-Decoding Correction in Speech LLMs

PTC-Bias is a two-stage framework that improves contextual biasing in speech large language models by using phoneme-level temporal competition. In the first stage, PTC Retrieval performs frame-synchronous phoneme decoding to generate a compact shortlist of bias words and their speech intervals. The second stage, PTC Correction, applies a local competition between retrieved candidates and mismatched transcript spans within those intervals, reducing near-homophone and word-segmentation errors without extra SpeechLLM passes. Experiments on LibriSpeech demonstrate consistent gains across two SpeechLLMs, with PTC-Bias reducing B-WER by up to 23.9% relative to CTC-Filter while keeping U-WER nearly unchanged.

By Zhiqi Ai, Han Cheng, Shiyi Mu, Yongjin Zhou, Shugong Xu
arXiv Machine Learning
Jul 14

An Empirical Recipe for Universal Phone Recognition

arXiv:2603. 29042v2 Announce Type: replace-cross Abstract: Phone recognition (PR) is a key enabler of multilingual and low-resource speech processing tasks, yet robust performance remains elusive.

By Shikhar Bharadwaj, Chin-Jou Li, Kwanghee Choi, Eunjung Yeo, William Chen, Shinji Watanabe, David R. Mortensen
arXiv Computation and Language
Sep 7

The Anatomy of an ASR Hallucination

The paper investigates why automatic speech recognition (ASR) systems sometimes generate fluent text that does not correspond to the input audio, a phenomenon termed hallucination. By examining two independently trained Conformer‑Large models—one using CTC and the other RNN‑T—under conditions of environmental noise and speaker‑background shift, the authors identify the final encoder stage as a critical boundary. Bypassing this final block leads to divergence on almost all utterances, while bypassing earlier blocks has minimal effect; at this stage, representations become compact, the decoder can read the text, and grapheme information becomes explicit, yet the output is garbled or repetitive rather than fluent fabrication. The study thus pinpoints a mechanistic precondition for hallucination—failure to produce adequately grounded output—though it does not fully explain naturally occurring hallucinations, and it highlights a consistent terminal‑stage dependency across decoder families and distribution shifts.

By Hamees Sayed, Apoorv Singh, Kumar Aman, Akshat Mandloi