arXiv AI By Kwok-Ho Ng, Tingting Song, Bingwen Feng, Zhihua Xia

WST-Graph: Topology-Preserving Wavelet Scattering Front-End for Speech Deepfake Detection

Read the original on arXiv AI →

The paper introduces WST-Graph, a topology-preserving wavelet scattering front-end designed for speech deepfake detection. It reconstructs wavelet scattering paths into a sparse modulation‑carrier grid that feeds an AASIST graph backend, employing modulation‑level normalization and adaptive local attention pooling to produce fixed relative‑time representations while keeping acoustic axes intact. The resulting waveform‑to‑graph interface uses about 60% fewer trainable parameters than AASIST yet remains competitive, showing clear improvements on selected out‑of‑domain benchmarks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 24

Multilevel Graph Wavelet Compressed Sensing with Scale-Aware Neural Recovery

arXiv:2607. 20857v1 Announce Type: cross Abstract: Scientific machine learning methods such as neural operators and physics-informed neural networks have advanced engineering applications and inverse problems, but their training typically requires large volumes of simulated data.

By Amirhossein Nouranizadeh, Sarang Rajendra Patil, Alan John Varghese, Varsha Narayanan, Amit Chakraborty, Mengjia Xu
arXiv AI
Sep 10

AudioFuse: Unified Spectral-Temporal Learning via a Hybrid ViT-1D CNN Architecture for Robust Phonocardiogram Classification

AudioFuse is a hybrid architecture that jointly learns from spectrograms and raw waveforms to classify phonocardiograms. It combines a wide-and-shallow Vision Transformer for spectral features with a shallow 1D CNN for temporal waveforms, reducing overfitting while capturing complementary information. On the PhysioNet 2016 dataset, AudioFuse achieves a state‑of‑the‑art ROC‑AUC of 0.8608 and shows superior robustness to domain shift on the PASCAL dataset, outperforming both spectrogram‑only and waveform‑only baselines.

By Md. Saiful Bari Siddiqui, Utsab Saha
arXiv Machine Learning
2d ago

Harmonizing Spectral Evolution in Conditional Flow Matching for TTS

Conditional Flow Matching models for text‑to‑speech often produce incoherent frequency evolution during inference. The authors propose a training‑free, frequency‑selective boosting strategy that uses the Discrete Wavelet Transform to dynamically modulate mel‑spectrogram sub‑bands during ODE integration, penalizing aggressive low‑frequency growth while boosting lagging high‑frequency details. Across multiple architectures, this method reduces the number of function evaluations from 32 to 26 and improves Frechet Audio Distance by up to 61% without harming mean opinion scores, speaker similarity, or intelligibility.

By Isha Pandey, Varad Deshpande, Abhijat Bharadwaj, Ganesh Ramakrishnan