UL-UNAS: Ultra-Lightweight U-Nets for Real-Time Speech Enhancement via Network Architecture Search
arXiv:2503. 00340v2 Announce Type: cross Abstract: Lightweight models are essential for real-time speech enhancement applications.
arXiv:2606. 12662v1 Announce Type: cross Abstract: Speech enhancement models typically apply uniform capacity across all frequencies, disregarding the non-uniform spectral resolution of human hearing.
arXiv:2503. 00340v2 Announce Type: cross Abstract: Lightweight models are essential for real-time speech enhancement applications.
The paper presents a redesign of the LiSenNet speech‑enhancement model for deployment on the STM32N6570‑DK Neural‑ART microcontroller accelerator. By replacing the recurrent bottleneck with convolutional mixers, converting unsupported operations to static int8 primitives, and using bounded decoder activations, the authors achieve an NPU‑compatible model that matches or surpasses the original LiSenNet in quality (PESQ 3.08 vs 3.01 FP32) while running each 16 ms input hop in 4.83 ms (real‑time factor 0.30). The study demonstrates that co‑designing parameter count, operator compatibility, quantization range, and streaming state is essential for efficient real‑time speech enhancement on constrained NPUs.
The paper introduces a sequence‑parallel band‑split enhancement front‑end called Parallel Time‑Band Mixer (PTBM) that removes recurrent unrolling within blocks. PTBM combines intra‑band temporal mixing with per‑frame cross‑band attention in a fully parallel architecture, while a learned Observation‑Adding (LOA) module suppresses ASR‑sensitive artifacts without development‑set tuning. Experiments on DNS Challenge and CHiME‑4 using frozen Whisper back‑ends show that this lightweight front‑end (0.96 M parameters, 0.58 GMAC/s) consistently lowers word error rate compared to recurrent band‑split baselines.
The paper investigates whether per‑frame early exit can improve compute‑matched performance for on‑device speech enhancement. By supervising every intermediate depth of a causal model and fine‑tuning output heads, the authors produce a family of static models that are more Pareto‑efficient than those trained from scratch, achieving up to 0.11 higher PESQ for equivalent compute and matching the best PESQ at 30% less compute. After int8 quantization, the dynamic enhancer performs on the same latency‑quality frontier as static models on an STM32N6 microcontroller, with the policy execution adding only 26 µs per frame and a 2.2% latency overhead from graph splitting.
arXiv:2609.18009v2 Announce Type: replace-cross Abstract: Lightweight audio-visual speech enhancement (AVSE) models face a critical trade-off between computational efficiency and cross-modal alignmen...
arXiv:2609.18009v1 Announce Type: cross Abstract: Lightweight audio-visual speech enhancement (AVSE) models face a critical trade-off between computational efficiency and cross-modal alignment accura...
arXiv:2606. 14120v1 Announce Type: cross Abstract: Auditory attention decoding (AAD) aims to infer the attended speaker from neural responses in multi-speaker acoustic environments and is a key problem for neuro-steered hearing systems.
arXiv:2604. 01832v1 Announce Type: cross Abstract: We introduce GAP-URGENet, a generative-predictive fusion framework developed for Track 1 of the ICASSP 2026 URGENT Challenge.
arXiv:2604. 14606v2 Announce Type: cross Abstract: Universal speech enhancement (USE) aims to restore speech signals from diverse distortions across multiple sampling rates.
arXiv:2607. 10191v1 Announce Type: cross Abstract: Generative streaming models for Target Speaker Extraction (TSE) commonly exhibit a quality--intelligibility trade-off, wherein naive optimization for perceptual audio quality tends to degrade speech intelligibility, and conversely.
arXiv:2609.36577v1 Announce Type: cross Abstract: Audio large language models (ALLMs) can reason about the content of audio recordings to perform complex tasks. However, these capabilities usually co...
arXiv:2606. 10046v1 Announce Type: cross Abstract: Flow-matching transformers achieve strong audio separation, yet their attention dynamics are opaque.