arXiv Machine Learning By Rahul Jain, Pierre Trepagnier, Rick Gentile, Joey Botero, Alexia Schulz

Clearing the Underbrush: AI-Enhanced RF Interference Suppression

Read the original on arXiv Machine Learning →

The paper presents an AI‑enhanced method for radio frequency interference suppression that builds on autoregressive transformer models by adding a Finite Scalar Quantization tokenizer layer. This addition improves interference rejection while maintaining low latency, and the authors also test other inference optimizations to speed up processing with minimal accuracy loss. Experiments using a digitally modulated RF signal as the signal of interest and a digital television OFDM signal as interference show that the approach outperforms traditional techniques and prior AI methods, with benefits demonstrated through audio quality metrics like PESQ and potential operational applications.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 7

SNAP: Speaker Nulling for Artifact Projection in Speech Deepfake Detection

The paper introduces SNAP, a speaker‑nulling framework designed to improve deepfake speech detection. By estimating a speaker subspace and orthogonally projecting out speaker‑dependent components, SNAP isolates synthesis artifacts in the residual features. This reduction of speaker entanglement enables detectors to focus on artifact‑related cues, achieving state‑of‑the‑art performance.

By Kyudan Jung, Jihwan Kim, Minwoo Lee, Soyoon Kim, Jeonghoon Kim, Jaegul Choo, Cheonbok Park
arXiv Computation and Language
Sep 14

Kraken: LLM-based Speech-to-Speech Translation via Low-bitrate VQ and Dual-path Source Conditioning

arXiv:2609.13045v1 Announce Type: new Abstract: Speech-to-speech translation (S2ST) has advanced significantly with speech LLMs, offering the potential for joint optimization and preserving non-lingu...

By Hayato Futami, Hassan Shahmohammadi, Tushar Dhyani, Alkis Koudounas, Rapha\"el Lafargue, Yosuke Kashiwagi, Quentin Jodelet, Emiru Tsunoo
arXiv Machine Learning
Aug 24

Training DeepFilterNet with Accurate Room Acoustic Simulations Improves Single-Channel Speech Enhancement

The study examines how the realism of synthetic room impulse response (RIR) datasets influences the training of DeepFilterNet3 for single‑channel speech enhancement. By comparing a DNS4 image‑source‑method RIR set with a higher‑fidelity hybrid wave‑based and geometrical acoustics RIR set, the authors find that the more realistic dataset consistently improves objective speech enhancement metrics and significantly reduces ASR word error rates on unseen measured RIRs. The results suggest that overall realism in synthetic acoustic training data enhances DeepFilterNet3’s generalization to new environments.

By Alessia Milo, Georg G\"otz, Steinar Gu{\dh}j\'onsson, Daniel Gert Nielsen, Jesper Pedersen, Finnur Pind