arXiv:2604. 01832v1 Announce Type: cross Abstract: We introduce GAP-URGENet, a generative-predictive fusion framework developed for Track 1 of the ICASSP 2026 URGENT Challenge.
By Xiaobin Rong, Yushi Wang, Zheng Wang, Jing Lu
arXiv:2607. 08717v1 Announce Type: new Abstract: Narrowband interference (NBI) severely degrades orthogonal frequency-division multiplexing (OFDM) systems by corrupting subcarriers and rendering classical soft demodulation ineffective.
By Emmanouil Kavvousanos, Francky Catthoor, Vassilis Paliouras
The paper introduces SNAP, a speaker‑nulling framework designed to improve deepfake speech detection. By estimating a speaker subspace and orthogonally projecting out speaker‑dependent components, SNAP isolates synthesis artifacts in the residual features. This reduction of speaker entanglement enables detectors to focus on artifact‑related cues, achieving state‑of‑the‑art performance.
By Kyudan Jung, Jihwan Kim, Minwoo Lee, Soyoon Kim, Jeonghoon Kim, Jaegul Choo, Cheonbok Park
arXiv:2609.13045v1 Announce Type: new
Abstract: Speech-to-speech translation (S2ST) has advanced significantly with speech LLMs, offering the potential for joint optimization and preserving non-lingu...
By Hayato Futami, Hassan Shahmohammadi, Tushar Dhyani, Alkis Koudounas, Rapha\"el Lafargue, Yosuke Kashiwagi, Quentin Jodelet, Emiru Tsunoo
arXiv:2601. 09239v5 Announce Type: replace-cross Abstract: Speech tokenizers are a key building block of fully discrete Speech LLMs.
By Hanlin Zhang, Daxin Tan, Dehua Tao, Xiao Chen, Haochen Tan, Yunhe Li, Yuchen Cao, Linqi Song
The study examines how the realism of synthetic room impulse response (RIR) datasets influences the training of DeepFilterNet3 for single‑channel speech enhancement. By comparing a DNS4 image‑source‑method RIR set with a higher‑fidelity hybrid wave‑based and geometrical acoustics RIR set, the authors find that the more realistic dataset consistently improves objective speech enhancement metrics and significantly reduces ASR word error rates on unseen measured RIRs. The results suggest that overall realism in synthetic acoustic training data enhances DeepFilterNet3’s generalization to new environments.
By Alessia Milo, Georg G\"otz, Steinar Gu{\dh}j\'onsson, Daniel Gert Nielsen, Jesper Pedersen, Finnur Pind
arXiv:2606. 09677v1 Announce Type: cross Abstract: While discriminative models for multi-channel speech separation excel in reference-based metrics, they often exhibit suboptimal human listening quality.
By Dohwan Kim, Jung-Woo Choi
BRIDLE is a self‑supervised encoder pretraining framework that extends bidirectional training to audio, image, and video by incorporating residual quantization (RQ) with multiple hierarchical codebooks. This approach allows fine‑grained discretization of latent representations and interleaves training between the encoder and tokenizer. Experiments show that BRIDLE achieves state‑of‑the‑art results on audio classification benchmarks and competitive performance on image and video classification tasks, outperforming traditional vector‑quantization methods.
By Hoang M. Nguyen, Satya N. Shukla, Qiang Zhang, Hanchao Yu, Sreya D. Roy, Dipesh Tamboli, Taipeng Tian, Lingjiong Zhu, Yuchen Liu
Neural audio codecs are a key component of speech processing pipelines, compressing audio into discrete tokens for downstream modeling. However, existing codecs struggle to balance reconstruction quality with token efficiency, often encoding perceptually irrelevant information such as background noise and recording artifacts at the expense of linguistically and acoustically meaningful content.
arXiv:2607. 10191v1 Announce Type: cross Abstract: Generative streaming models for Target Speaker Extraction (TSE) commonly exhibit a quality--intelligibility trade-off, wherein naive optimization for perceptual audio quality tends to degrade speech intelligibility, and conversely.
By Shuhai Peng, Jinjiang Liu, Hui Lu, Liyang Chen, Guiping Zhong, Jiakui Li, Shiyin Kang, Zhiyong Wu
arXiv:2608. 09288v1 Announce Type: cross Abstract: Audio-visual speech enhancement under real-world conditions remains challenging due to unreliable visual inputs and the lack of large-scale training data with realistic acoustic conditions.
By Wei Zhou, Wanyi Ning, Yinshang Guo, Qianxiao Fang, Haitao Qian, Yingpeng Li
arXiv:2604. 24199v4 Announce Type: replace-cross Abstract: We propose Speech Enhancement based on Drifting Models (DriftSE), a novel generative framework that formulates denoising as an equilibrium problem.
By Liang Xu, Diego Caviedes-Nozal, W. Bastiaan Kleijn, Longfei Felix Yan, Rasmus Kongsgaard Olsson