arXiv Computation and Language By Avishai Weizman, Yehuda Ben-Shimol, Itshak Lapidot

Spooftral: Can Voxtral Audio-Language Model Detect Speech Spoofing?

Read the original on arXiv Computation and Language →

The paper investigates whether the Voxtral audio‑language model can detect speech spoofing. It shows that without task‑specific adaptation, the model’s language‑model layers prioritize semantic content, making spoof‑discriminative acoustic cues less separable. By applying lightweight weight‑decomposed low‑rank adaptation (DoRA), the authors create Spooftral, which achieves an equal error rate of 4.25% on the ASVspoof5 evaluation set.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jun 10

Linguistically Augmented Audio Speech Data (LinguAS)

arXiv:2606. 10246v1 Announce Type: cross Abstract: Maliciously-created fake speech, including deepfaked and spoofed audio, is proliferating at an alarming rate, and detection models are racing to stay ahead of the curve.

By Ashley R. Keaton, Zahra Khanjani, Christine Mallinson, Vandana P. Janeja