arXiv AI By Xingyu Shen, Runze Wang, Wei-Ping Zhu, Benoit Champagne

Parallel Time-Band Mixing with Learned Observation-Adding for Robust ASR Front-Ends

Read the original on arXiv AI →

The paper introduces a sequence‑parallel band‑split enhancement front‑end called Parallel Time‑Band Mixer (PTBM) that removes recurrent unrolling within blocks. PTBM combines intra‑band temporal mixing with per‑frame cross‑band attention in a fully parallel architecture, while a learned Observation‑Adding (LOA) module suppresses ASR‑sensitive artifacts without development‑set tuning. Experiments on DNS Challenge and CHiME‑4 using frozen Whisper back‑ends show that this lightweight front‑end (0.96 M parameters, 0.58 GMAC/s) consistently lowers word error rate compared to recurrent band‑split baselines.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 22

Closing the Speech-Text Gap with Limited Audio for Effective Domain Adaptation in LLM-Based ASR

arXiv:2604.06487v2 Announce Type: replace Abstract: Conventional end-to-end automatic speech recognition (ASR) systems rely on paired speech-text data for domain adaptation. Recent LLM-based ASR arch...

By Thibault Ba\~neras-Roux, Sergio Burdisso, Esa\'u Villatoro-Tello, Dairazalia S\'anchez-Cort\'es, Shiran Liu, Severin Baroudi, Shashi Kumar, Hasindri Watawana, Manjunath K E, Kadri Hacioglu, Petr Motlicek, Andreas Stolcke