arXiv:2606. 12662v1 Announce Type: cross Abstract: Speech enhancement models typically apply uniform capacity across all frequencies, disregarding the non-uniform spectral resolution of human hearing.
By Damien Martins Gomes, Fran\c{c}ois Capman
The paper presents a redesign of the LiSenNet speech‑enhancement model for deployment on the STM32N6570‑DK Neural‑ART microcontroller accelerator. By replacing the recurrent bottleneck with convolutional mixers, converting unsupported operations to static int8 primitives, and using bounded decoder activations, the authors achieve an NPU‑compatible model that matches or surpasses the original LiSenNet in quality (PESQ 3.08 vs 3.01 FP32) while running each 16 ms input hop in 4.83 ms (real‑time factor 0.30). The study demonstrates that co‑designing parameter count, operator compatibility, quantization range, and streaming state is essential for efficient real‑time speech enhancement on constrained NPUs.
By Cl\'ement Laroche, Rasmus Kongsgaard Olsson
The paper introduces a reparameterization technique that injects feature noise to jointly optimize speech model performance and computational complexity during training. Unlike traditional pruning, this method dynamically adjusts model size for a desired performance‑complexity trade‑off without heuristic weight removal. The authors validate their approach with a synthetic example and two real‑world applications—voice activity detection and audio anti‑spoofing—providing publicly available code for further research.
By Esteban G\'omez, Tom B\"ackstr\"om
The paper investigates whether per‑frame early exit can improve compute‑matched performance for on‑device speech enhancement. By supervising every intermediate depth of a causal model and fine‑tuning output heads, the authors produce a family of static models that are more Pareto‑efficient than those trained from scratch, achieving up to 0.11 higher PESQ for equivalent compute and matching the best PESQ at 30% less compute. After int8 quantization, the dynamic enhancer performs on the same latency‑quality frontier as static models on an STM32N6 microcontroller, with the policy execution adding only 26 µs per frame and a 2.2% latency overhead from graph splitting.
By Cl\'ement Laroche, Riccardo Miccini
arXiv:2601. 06199v3 Announce Type: replace-cross Abstract: Scaling Multimodal Large Language Models (MLLMs) to long-form speech is bottlenecked by the explosive growth of input tokens.
By Junseok Lee, Sangyong Lee, Chang-Jae Chun
arXiv:2604. 01832v1 Announce Type: cross Abstract: We introduce GAP-URGENet, a generative-predictive fusion framework developed for Track 1 of the ICASSP 2026 URGENT Challenge.
By Xiaobin Rong, Yushi Wang, Zheng Wang, Jing Lu