arXiv Computation and Language By Ki Woong Moon, Daniel Brenner

Does a prosody-trained representation help beyond trainable fusion? A parameter-matched study with frozen HuBERT

Read the original on arXiv Computation and Language →

The study investigates whether a prosody‑trained representation can improve automatic speech recognition beyond the effect of a trainable fusion mechanism. Using a frozen HuBERT backbone and a 64‑dimensional prosodic representation, the authors compare a baseline recognizer, a fusion model with no auxiliary input, and a fusion model with the learned representation across Buckeye, Switchboard, and AMI IHM datasets. While the fusion model without auxiliary input reduces WER relative to the baseline, adding the prosodic representation yields no significant WER improvement, though the model still depends on the representation for optimal performance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jun 9

Liberating LLM Capabilities in Full-Duplex Speech Models

arXiv:2606. 07547v1 Announce Type: cross Abstract: Speech-based large language models are typically constrained to spoken replies, which limits their user-facing outputs to what can be verbalized and suppresses text-native capabilities such as code generation, structured analysis, and multi-step reasoning in realtime interaction, for tasks that require persistent, structured, and inspectable intermediate outputs.

By Luoyuan Zhang, Bokai Xu, Junbo Cui, Weiyue Sun, Yingjing Xu, Hanyu Liu, Yuan Yao
arXiv Computation and Language
Sep 28

Asymmetric Classifier-Free Guidance for Target-Speaker ASR

The paper introduces asymmetric classifier‑free guidance (CFG) for target‑speaker ASR using Whisper, where a speaker‑conditioned branch predicts the target transcript and a speaker‑unconditioned branch predicts serialized multi‑speaker transcripts. CFG modulates the influence of speaker conditioning during decoding via a single guidance scale, which is first set globally on development data and then refined per utterance by a lightweight encoder‑based predictor while keeping the recognition model fixed. The resulting system yields up to 21.8% relative WER reduction over a condition‑only baseline and 5.6% over standard conditional decoding under domain shifts.

By Yiwen Guan, Jacob Whitehill