arXiv AI

Deep Learning Based Relative Transfer Matrix Estimation for Multiple Sources and Multiple Microphones

arXiv:2608. 11627v1 Announce Type: cross Abstract: The Relative Transfer Matrix (ReTM), recently introduced as a generalization of the relative transfer function for multiple receivers and sources, shows promising performance when applied to speech enhancement in noisy environments.

arXiv AI
Sep 4

Test-time adaptation for speech enhancement with an autoregressive speech prior

The paper proposes a single‑utterance test‑time adaptation (TTA) method for speech enhancement that uses an autoregressive prior trained on clean speech latent representations from a neural audio codec. The adaptation regularizes a pretrained enhancement model by minimizing the Kullback‑Leibler divergence between the enhanced speech distribution and the clean speech prior. Experiments on multiple noisy speech datasets demonstrate consistent improvements in speech quality, especially when training and testing noise conditions differ.

By Sofiene Kammoun, Simon Leglaive, Xavier Alameda-Pineda, Timo Gerkmann
arXiv AI
Aug 18

Adding Voice Cloning to Text-to-Audio-Video Models with a Single Zero-Initialised Layer

arXiv:2608. 15690v1 Announce Type: cross Abstract: Text-to-audio-video (T2AV) generation models produce a video and its soundtrack from a textual description, but offer no control over whose voice speaks in the output.

By Ivan Mikheev, Viacheslav Vasilev, Anna Dmitrienko, Alexey Letunovskiy, Ivan Kirillov, Kirill Chernyshev, Denis Dimitrov
arXiv Machine Learning
Sep 11

Single Microphone Own Voice Detection based on Simulated Transfer Functions for Hearing Aids

The paper introduces a simulation-based method for detecting a user's own voice in hearing aids using only a single microphone. It employs a data augmentation strategy with simulated acoustic transfer functions to train a transformer classifier, achieving over 90% accuracy on both simulated and real-world recordings. The approach reduces hardware complexity and power consumption while maintaining robust performance across varied spatial conditions.

By Mathuranathan Mayuravaani, W. Bastiaan Kleijn, Andrew Lensen, Charlotte S{\o}rensen
arXiv AI
4d ago

Estimation of Room Impulse Responses from Handclaps

arXiv:2609.35839v1 Announce Type: cross Abstract: Handclaps provide an equipment-free excitation for room acoustics, but their unknown and variable source waveform makes room impulse response (RIR) e...

By Shih-Yu Lai, Kyung Yun Lee, Nils Meyer-Kahlen, Eloi Moliner, Bing-Yu Chen, Vesa V\"alim\"aki
arXiv AI
Jun 9

Speech Enhancement Based on Drifting Models

arXiv:2604. 24199v4 Announce Type: replace-cross Abstract: We propose Speech Enhancement based on Drifting Models (DriftSE), a novel generative framework that formulates denoising as an equilibrium problem.

By Liang Xu, Diego Caviedes-Nozal, W. Bastiaan Kleijn, Longfei Felix Yan, Rasmus Kongsgaard Olsson
arXiv Computation and Language
Sep 23

Challenges of Multi-Speaker Extraction for Real Conversational Speech Enhancement

The paper addresses challenges in extracting target and multiple speakers from real conversational speech, noting that real conversations contain more silence and enrolment samples that differ from the target speech. It introduces a new loss function that reduces the impact of excess silence during training, yielding improvements in STOI (from 0.55 to 0.60) and frequency‑weighted segmental SNR (from 4.35 to 5.12). The study also investigates how mismatches between enrolment and target speech affect performance.

By Robert Sutherland, Stefan Goetze, Jon Barker