arXiv Computer Vision

Pose Adaptive Dynamic FiLM Modulation for Visual Speech Recognition

The paper introduces Pose Adaptive Dynamic FiLM Modulation for Visual Speech Recognition, addressing head‑pose variation that causes appearance changes in VSR. It proposes a Dynamic Residual FiLM (DR‑FiLM) modulator that predicts input‑dependent weights to control pose‑conditioned modulation strength. Experiments on LRS2 and LRS3 show that DR‑FiLM reduces phoneme error rates compared to unweighted multi‑pathway modulation and that deeper FiLM pathways receive higher weights as head‑pose variation increases.

arXiv AI
Sep 17

GrainSpeech: Less Context, More Detail for Compact Speech Synthesis

GrainSpeech is a compact speech synthesis model that uses a fixed‑receptive‑field convolutional encoder to reduce pitch, energy, and duration prediction errors by 36.0%, 17.3%, and 3.4% respectively. It introduces a Mel‑specific gradient‑variance supervision that improves fine‑scale variation while avoiding quality degradation. With only 264.8K parameters, GrainSpeech achieves 17.9× real‑time Mel generation on a microcontroller and attains UTMOS scores comparable to much larger models, using less than 1.5% of their parameters.

By Zitao Liang, Chang Gao
arXiv AI
Jun 30

How to Leverage Synthetic Speech for LLM-Based ASR Systems?

arXiv:2606. 29031v1 Announce Type: cross Abstract: In regulated domains such as banking and healthcare, where privacy constraints make real speech costly to collect and retain, synthetic speech from modern text-to-speech (TTS) is an appealing alternative for training automatic speech recognition (ASR) without exposing sensitive customer recordings.

By Yanis Labrak, Dairazalia Sanchez-Cortes, Sergio Burdisso, S\'everin Baroudi, Shashi Kumar, Esa\'u Villatoro-Tello, Srikanth Madikeri, Manjunath K E, Old\v{r}ich Plchot, Kadri Hacio\u{g}lu, Petr Motlicek, Andreas Stolcke
arXiv Computation and Language
Sep 22

Closing the Speech-Text Gap with Limited Audio for Effective Domain Adaptation in LLM-Based ASR

arXiv:2604.06487v2 Announce Type: replace Abstract: Conventional end-to-end automatic speech recognition (ASR) systems rely on paired speech-text data for domain adaptation. Recent LLM-based ASR arch...

By Thibault Ba\~neras-Roux, Sergio Burdisso, Esa\'u Villatoro-Tello, Dairazalia S\'anchez-Cort\'es, Shiran Liu, Severin Baroudi, Shashi Kumar, Hasindri Watawana, Manjunath K E, Kadri Hacioglu, Petr Motlicek, Andreas Stolcke