arXiv:2503. 00340v2 Announce Type: cross Abstract: Lightweight models are essential for real-time speech enhancement applications.
By Xiaobin Rong, Leyan Yang, Dahan Wang, Yuxiang Hu, Changbao Zhu, Kai Chen, Jing Lu
The paper presents a redesign of the LiSenNet speech‑enhancement model for deployment on the STM32N6570‑DK Neural‑ART microcontroller accelerator. By replacing the recurrent bottleneck with convolutional mixers, converting unsupported operations to static int8 primitives, and using bounded decoder activations, the authors achieve an NPU‑compatible model that matches or surpasses the original LiSenNet in quality (PESQ 3.08 vs 3.01 FP32) while running each 16 ms input hop in 4.83 ms (real‑time factor 0.30). The study demonstrates that co‑designing parameter count, operator compatibility, quantization range, and streaming state is essential for efficient real‑time speech enhancement on constrained NPUs.
By Cl\'ement Laroche, Rasmus Kongsgaard Olsson
The paper introduces a sequence‑parallel band‑split enhancement front‑end called Parallel Time‑Band Mixer (PTBM) that removes recurrent unrolling within blocks. PTBM combines intra‑band temporal mixing with per‑frame cross‑band attention in a fully parallel architecture, while a learned Observation‑Adding (LOA) module suppresses ASR‑sensitive artifacts without development‑set tuning. Experiments on DNS Challenge and CHiME‑4 using frozen Whisper back‑ends show that this lightweight front‑end (0.96 M parameters, 0.58 GMAC/s) consistently lowers word error rate compared to recurrent band‑split baselines.
By Xingyu Shen, Runze Wang, Wei-Ping Zhu, Benoit Champagne
The paper investigates whether per‑frame early exit can improve compute‑matched performance for on‑device speech enhancement. By supervising every intermediate depth of a causal model and fine‑tuning output heads, the authors produce a family of static models that are more Pareto‑efficient than those trained from scratch, achieving up to 0.11 higher PESQ for equivalent compute and matching the best PESQ at 30% less compute. After int8 quantization, the dynamic enhancer performs on the same latency‑quality frontier as static models on an STM32N6 microcontroller, with the policy execution adding only 26 µs per frame and a 2.2% latency overhead from graph splitting.
By Cl\'ement Laroche, Riccardo Miccini
arXiv:2609.18009v2 Announce Type: replace-cross
Abstract: Lightweight audio-visual speech enhancement (AVSE) models face a critical trade-off between computational efficiency and cross-modal alignmen...
By Guo-Ruei Tseng, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen
arXiv:2609.18009v1 Announce Type: cross
Abstract: Lightweight audio-visual speech enhancement (AVSE) models face a critical trade-off between computational efficiency and cross-modal alignment accura...
By Guo-Ruei Tseng, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen