arXiv Computer Vision

Complementary rPPG-Derived and Lip-Region Frequency Cues for Talking-Face Deepfake Detection

arXiv AI
Sep 10

Noise Adaptive Streaming Audio-Visual Speech Token Enhancement for Robust Full-Duplex Spoken Dialogue Models

The paper introduces AV-STE, a modular streaming audio‑visual front‑end that enhances corrupted semantic speech tokens using noisy audio and lip video before they reach a frozen speech LLM. By preserving the downstream dialogue model’s pretrained conversational abilities, AV‑STE improves response coherence from 1.42 to 1.91 in same‑dataset speaker interference scenarios while maintaining turn‑taking behavior. These gains also transfer to out‑of‑domain Seamless Interaction.

By Bella Godiva, Yeonju Kim, Yong Man Ro
arXiv AI
Sep 10

AudioFuse: Unified Spectral-Temporal Learning via a Hybrid ViT-1D CNN Architecture for Robust Phonocardiogram Classification

AudioFuse is a hybrid architecture that jointly learns from spectrograms and raw waveforms to classify phonocardiograms. It combines a wide-and-shallow Vision Transformer for spectral features with a shallow 1D CNN for temporal waveforms, reducing overfitting while capturing complementary information. On the PhysioNet 2016 dataset, AudioFuse achieves a state‑of‑the‑art ROC‑AUC of 0.8608 and shows superior robustness to domain shift on the PASCAL dataset, outperforming both spectrogram‑only and waveform‑only baselines.

By Md. Saiful Bari Siddiqui, Utsab Saha
arXiv AI
Jun 10

RAT: Reference-Augmented Training for ASV Anti-Spoofing

arXiv:2606. 10908v1 Announce Type: cross Abstract: We introduce a spoofing countermeasure architecture conditioned on speaker-reference recordings, but observe that it converges to a solution that effectively ignores the reference during inference.

By Vojt\v{e}ch Stan\v{e}k, Anton Firc, Jakub Re\v{s}, Kamil Malinka