arXiv AI

BASENet: Band-Adapted Speech Enhancement Network with Cross-Band Attention

arXiv:2606. 12662v1 Announce Type: cross Abstract: Speech enhancement models typically apply uniform capacity across all frequencies, disregarding the non-uniform spectral resolution of human hearing.

arXiv Machine Learning
Sep 25

Beyond Model Size: Redesigning LiSenNet for embedded speech enhancement

The paper presents a redesign of the LiSenNet speech‑enhancement model for deployment on the STM32N6570‑DK Neural‑ART microcontroller accelerator. By replacing the recurrent bottleneck with convolutional mixers, converting unsupported operations to static int8 primitives, and using bounded decoder activations, the authors achieve an NPU‑compatible model that matches or surpasses the original LiSenNet in quality (PESQ 3.08 vs 3.01 FP32) while running each 16 ms input hop in 4.83 ms (real‑time factor 0.30). The study demonstrates that co‑designing parameter count, operator compatibility, quantization range, and streaming state is essential for efficient real‑time speech enhancement on constrained NPUs.

By Cl\'ement Laroche, Rasmus Kongsgaard Olsson
arXiv AI
Sep 1

Parallel Time-Band Mixing with Learned Observation-Adding for Robust ASR Front-Ends

The paper introduces a sequence‑parallel band‑split enhancement front‑end called Parallel Time‑Band Mixer (PTBM) that removes recurrent unrolling within blocks. PTBM combines intra‑band temporal mixing with per‑frame cross‑band attention in a fully parallel architecture, while a learned Observation‑Adding (LOA) module suppresses ASR‑sensitive artifacts without development‑set tuning. Experiments on DNS Challenge and CHiME‑4 using frozen Whisper back‑ends show that this lightweight front‑end (0.96 M parameters, 0.58 GMAC/s) consistently lowers word error rate compared to recurrent band‑split baselines.

By Xingyu Shen, Runze Wang, Wei-Ping Zhu, Benoit Champagne
arXiv Machine Learning
Sep 25

Does per-frame early exit pay? A compute-matched study of dynamic depth for on-device speech enhancement

The paper investigates whether per‑frame early exit can improve compute‑matched performance for on‑device speech enhancement. By supervising every intermediate depth of a causal model and fine‑tuning output heads, the authors produce a family of static models that are more Pareto‑efficient than those trained from scratch, achieving up to 0.11 higher PESQ for equivalent compute and matching the best PESQ at 30% less compute. After int8 quantization, the dynamic enhancer performs on the same latency‑quality frontier as static models on an STM32N6 microcontroller, with the policy execution adding only 26 µs per frame and a 2.2% latency overhead from graph splitting.

By Cl\'ement Laroche, Riccardo Miccini
arXiv AI
Jul 14

Breaking the Quality--Intelligibility Trade-off in Streaming Target Speaker Extraction via Deep-Feature-Anchored Preference Optimization

arXiv:2607. 10191v1 Announce Type: cross Abstract: Generative streaming models for Target Speaker Extraction (TSE) commonly exhibit a quality--intelligibility trade-off, wherein naive optimization for perceptual audio quality tends to degrade speech intelligibility, and conversely.

By Shuhai Peng, Jinjiang Liu, Hui Lu, Liyang Chen, Guiping Zhong, Jiakui Li, Shiyin Kang, Zhiyong Wu