The paper presents a redesign of the LiSenNet speech‑enhancement model for deployment on the STM32N6570‑DK Neural‑ART microcontroller accelerator. By replacing the recurrent bottleneck with convolutional mixers, converting unsupported operations to static int8 primitives, and using bounded decoder activations, the authors achieve an NPU‑compatible model that matches or surpasses the original LiSenNet in quality (PESQ 3.08 vs 3.01 FP32) while running each 16 ms input hop in 4.83 ms (real‑time factor 0.30). The study demonstrates that co‑designing parameter count, operator compatibility, quantization range, and streaming state is essential for efficient real‑time speech enhancement on constrained NPUs.
By Cl\'ement Laroche, Rasmus Kongsgaard Olsson
The paper presents a billing‑aware neural text‑to‑speech system for serverless CPUs, focusing on minimizing CPU‑seconds and GB‑seconds rather than just throughput or latency. By using request‑sized concurrent inference and a reclaimable instance lifecycle, the system limits per‑request CPU parallelism and releases idle memory while keeping the server process alive. On the Kokoro‑82M benchmark, it achieves 2.71 audio‑seconds per CPU‑second versus 0.89 with ONNX Runtime, cuts cost per audio‑hour from $0.0631 to $0.0153, and reduces idle billed memory from 8.7 GB to 1.33 GB, with faster restoration times.
By Pakorn Nathong, Kunat Pipatanakul
arXiv:2606. 22790v2 Announce Type: replace-cross Abstract: In this paper, we investigate the tradeoffs between compute allocation and model performance for two speech processing tasks: Automatic Speech Recognition (ASR) and Speech Emotion Recognition (SER).
By Vyom Agarwal, Mokshda Gangrade, Siddharth Pal, Jerry Wu
GEPARD is a streaming text‑to‑speech model that uses a standard large language model backbone to generate speech autoregressively, decoding audio with an FSQ‑based neural codec. It streams audio chunk‑by‑chunk as text arrives, achieving a real‑time factor of about 0.067 and an aggregate speedup of roughly 204× on a single GPU with 256 concurrent streams. The design keeps all complex auxiliary mechanisms outside the decode loop, enabling deployment with a standard LLM engine (vLLM) without kernel modifications.
By Denis Pavlov, Ulanbek Abdurazakov, Nursultan Bakashov
arXiv:2606. 12662v1 Announce Type: cross Abstract: Speech enhancement models typically apply uniform capacity across all frequencies, disregarding the non-uniform spectral resolution of human hearing.
By Damien Martins Gomes, Fran\c{c}ois Capman
arXiv:2503. 00340v2 Announce Type: cross Abstract: Lightweight models are essential for real-time speech enhancement applications.
By Xiaobin Rong, Leyan Yang, Dahan Wang, Yuxiang Hu, Changbao Zhu, Kai Chen, Jing Lu