arXiv Computation and Language

RT-SEMamba: Real-Time Speech Enhancement Mamba via Progressive Knowledge Distillation

RT-SEMamba is a fully causal speech enhancement model that uses causal time‑frequency Mamba blocks instead of Transformer‑based architectures, allowing efficient long‑form inference with a fixed‑size recurrent state. The authors introduce a progressive knowledge distillation strategy that compresses an 8‑layer teacher into a single‑layer student by jointly distilling spectral outputs and intermediate representations. On the Voicebank‑DEMAND benchmark, the 8‑layer model achieves 3.32 PESQ under a 25 ms latency constraint, while the distilled 1‑layer student improves from 3.06 to 3.18 PESQ, maintains the same steady‑state real‑time factor, and runs 2.64× faster than the teacher.

arXiv AI
6d ago

Acoustic-to-Text KV Compression for Full-Duplex Speech Models

The paper introduces acoustic-to-text KV compression for full‑duplex speech models, converting acoustic key‑value states into compact textual memory during listening‑time slack. When the KV cache exceeds a target budget, older acoustic states are evicted while transcripts and recent acoustic context are retained. Experiments on ten‑minute LongSpeech sessions show a 64.6% reduction in peak streaming KV‑cache size and improved transcription, temporal question answering, and summarization, with comparable pause‑handling, turn‑taking, and interruption performance in Full‑Duplex‑Bench.

By Yejin Lee, Seungbeom Kim, Yongha Lee, Kyuhong Shim
arXiv AI
2d ago

REALM: Retrospective Encoder Alignment for LFP Modeling

REALM is a retrospective knowledge distillation framework that enables causal decoding of behavior from local field potentials (LFPs). It trains a bidirectional Mamba‑2 teacher on multi‑session data using continuous masked autoencoding, then distills its representations into a compact causal student model. The resulting LFP‑only decoder achieves the highest mean accuracy among compared methods, surpassing state‑of‑the‑art baselines while using fewer parameters and less pretraining time.

By Peicheng Wu, Zhenyu Bu, Runze Ma, Lin Du
arXiv AI
Jun 9

Liberating LLM Capabilities in Full-Duplex Speech Models

arXiv:2606. 07547v1 Announce Type: cross Abstract: Speech-based large language models are typically constrained to spoken replies, which limits their user-facing outputs to what can be verbalized and suppresses text-native capabilities such as code generation, structured analysis, and multi-step reasoning in realtime interaction, for tasks that require persistent, structured, and inspectable intermediate outputs.

By Luoyuan Zhang, Bokai Xu, Junbo Cui, Weiyue Sun, Yingjing Xu, Hanyu Liu, Yuan Yao
arXiv Machine Learning
Sep 25

Does per-frame early exit pay? A compute-matched study of dynamic depth for on-device speech enhancement

The paper investigates whether per‑frame early exit can improve compute‑matched performance for on‑device speech enhancement. By supervising every intermediate depth of a causal model and fine‑tuning output heads, the authors produce a family of static models that are more Pareto‑efficient than those trained from scratch, achieving up to 0.11 higher PESQ for equivalent compute and matching the best PESQ at 30% less compute. After int8 quantization, the dynamic enhancer performs on the same latency‑quality frontier as static models on an STM32N6 microcontroller, with the policy execution adding only 26 µs per frame and a 2.2% latency overhead from graph splitting.

By Cl\'ement Laroche, Riccardo Miccini