arXiv Machine Learning By Cl\'ement Laroche, Riccardo Miccini

Does per-frame early exit pay? A compute-matched study of dynamic depth for on-device speech enhancement

Read the original on arXiv Machine Learning →

The paper investigates whether per‑frame early exit can improve compute‑matched performance for on‑device speech enhancement. By supervising every intermediate depth of a causal model and fine‑tuning output heads, the authors produce a family of static models that are more Pareto‑efficient than those trained from scratch, achieving up to 0.11 higher PESQ for equivalent compute and matching the best PESQ at 30% less compute. After int8 quantization, the dynamic enhancer performs on the same latency‑quality frontier as static models on an STM32N6 microcontroller, with the policy execution adding only 26 µs per frame and a 2.2% latency overhead from graph splitting.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 25

Beyond Model Size: Redesigning LiSenNet for embedded speech enhancement

The paper presents a redesign of the LiSenNet speech‑enhancement model for deployment on the STM32N6570‑DK Neural‑ART microcontroller accelerator. By replacing the recurrent bottleneck with convolutional mixers, converting unsupported operations to static int8 primitives, and using bounded decoder activations, the authors achieve an NPU‑compatible model that matches or surpasses the original LiSenNet in quality (PESQ 3.08 vs 3.01 FP32) while running each 16 ms input hop in 4.83 ms (real‑time factor 0.30). The study demonstrates that co‑designing parameter count, operator compatibility, quantization range, and streaming state is essential for efficient real‑time speech enhancement on constrained NPUs.

By Cl\'ement Laroche, Rasmus Kongsgaard Olsson
arXiv AI
2d ago

Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing

The paper presents a billing‑aware neural text‑to‑speech system for serverless CPUs, focusing on minimizing CPU‑seconds and GB‑seconds rather than just throughput or latency. By using request‑sized concurrent inference and a reclaimable instance lifecycle, the system limits per‑request CPU parallelism and releases idle memory while keeping the server process alive. On the Kokoro‑82M benchmark, it achieves 2.71 audio‑seconds per CPU‑second versus 0.89 with ONNX Runtime, cuts cost per audio‑hour from $0.0631 to $0.0153, and reduces idle billed memory from 8.7 GB to 1.33 GB, with faster restoration times.

By Pakorn Nathong, Kunat Pipatanakul
arXiv Machine Learning
Sep 7

GEPARD - Generative, Prosody-aware, Autoregressive text-to-speech model for Realtime Dialogue

GEPARD is a streaming text‑to‑speech model that uses a standard large language model backbone to generate speech autoregressively, decoding audio with an FSQ‑based neural codec. It streams audio chunk‑by‑chunk as text arrives, achieving a real‑time factor of about 0.067 and an aggregate speedup of roughly 204× on a single GPU with 256 concurrent streams. The design keeps all complex auxiliary mechanisms outside the decode loop, enabling deployment with a standard LLM engine (vLLM) without kernel modifications.

By Denis Pavlov, Ulanbek Abdurazakov, Nursultan Bakashov