arXiv AI

TF-MoE: Time-Frequency Mixture-of-Experts for Efficient Speech Separation

arXiv:2606. 29575v1 Announce Type: cross Abstract: Recent advances in speech separation (SS) have led to compact front-end models with small parameter sizes, yet their high computational cost remains a major barrier for deployment on edge devices.

Hugging Face Trending Papers
Jun 28

TF-MoE: Time-Frequency Mixture-of-Experts for Efficient Speech Separation

Recent advances in speech separation (SS) have led to compact front-end models with small parameter sizes, yet their high computational cost remains a major barrier for deployment on edge devices. To address this, we propose TF-MoE, a sparse Mixture-of-Experts (MoE) framework that enhances model capacity with almost no increase in inference cost.

arXiv AI
5d ago

HybridSB-MoE: Dual-Domain Schr\"odinger Bridges with Scene-Adaptive Expert Routing for Speech Enhancement

arXiv:2608. 12715v1 Announce Type: cross Abstract: Generative speech enhancement faces three gaps: spectral models capture harmonic structure but often disrupt phase, waveform models preserve phase but miss harmonics, and Schr\"odinger Bridges (SB) shorten transport from noise to clean speech but leave inference cost only loosely tied to training.

By Zhengyi Lu, Aswini Sivakumar, Jie Hu, Yao Qiang
arXiv AI
2d ago

DeaMoE: Efficient MoE Structure for Fast Small-Batch Decoding

arXiv:2608. 14385v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models have been widely adopted in real-time interactive applications such as coding assistants, real-time audio-video interaction systems.

By Zewen Jin, Shen Fu, Zeping Duan, Shannon Wang, Weihao Wu, Chengjie Tang, Congkun Ai, Ping Gong, Zijian Dai, Youhui Bai, Cheng Li
arXiv Machine Learning
Aug 10

Cloud-Boosted Low-Compute Multi-Channel Speech Enhancement

arXiv:2608. 07423v1 Announce Type: cross Abstract: Low-latency, low-compute speech enhancement is essential for wearable devices with real-time communication requirements, but strict computational constraints significantly limit on-device performance.

By Xulin Fan, Juan Azcarreta, Ashutosh Pandey, Jesus Alvarez, Ke Tan, Jacob Donley, Ritwik Giri, Buye Xu
arXiv Machine Learning
Jul 28

Music-Source-Separation-Training (MSST): A Unified Framework for Training and Evaluating Music Demixing Models

arXiv:2607. 23395v1 Announce Type: cross Abstract: Music Source Separation (MSS), the task of recovering individual sound components (stems) from a polyphonic mixture, is central to applications ranging from karaoke and remixing to audio restoration and content production.

By Roman Solovyev, Ilya Kiselev, Alexander Stempkovskiy, Tatiana Gabruseva