arXiv Computer Vision By Matthew Kit Khinn Teng, Haibo Zhang, Takeshi Saitoh

Pose Adaptive Dynamic FiLM Modulation for Visual Speech Recognition

Read the original on arXiv Computer Vision →

The paper introduces Pose Adaptive Dynamic FiLM Modulation for Visual Speech Recognition, addressing head‑pose variation that causes appearance changes in VSR. It proposes a Dynamic Residual FiLM (DR‑FiLM) modulator that predicts input‑dependent weights to control pose‑conditioned modulation strength. Experiments on LRS2 and LRS3 show that DR‑FiLM reduces phoneme error rates compared to unweighted multi‑pathway modulation and that deeper FiLM pathways receive higher weights as head‑pose variation increases.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Sep 17

GrainSpeech: Less Context, More Detail for Compact Speech Synthesis

GrainSpeech is a compact speech synthesis model that uses a fixed‑receptive‑field convolutional encoder to reduce pitch, energy, and duration prediction errors by 36.0%, 17.3%, and 3.4% respectively. It introduces a Mel‑specific gradient‑variance supervision that improves fine‑scale variation while avoiding quality degradation. With only 264.8K parameters, GrainSpeech achieves 17.9× real‑time Mel generation on a microcontroller and attains UTMOS scores comparable to much larger models, using less than 1.5% of their parameters.

By Zitao Liang, Chang Gao