arXiv AI By Mahtab Masoudi Nezhad, Nima Karimian

Robust Spoofed Speech Detection via Temporal Pyramid Modeling

Read the original on arXiv AI →

arXiv:2606. 16837v1 Announce Type: cross Abstract: Spoofed speech detection is increasingly challenged by realistic synthesis, voice conversion, and replay attacks, with cross-dataset generalization remaining a major limitation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 25

Spooftral: Can Voxtral Audio-Language Model Detect Speech Spoofing?

The paper investigates whether the Voxtral audio‑language model can detect speech spoofing. It shows that without task‑specific adaptation, the model’s language‑model layers prioritize semantic content, making spoof‑discriminative acoustic cues less separable. By applying lightweight weight‑decomposed low‑rank adaptation (DoRA), the authors create Spooftral, which achieves an equal error rate of 4.25% on the ASVspoof5 evaluation set.

By Avishai Weizman, Yehuda Ben-Shimol, Itshak Lapidot
arXiv AI
Jun 10

Linguistically Augmented Audio Speech Data (LinguAS)

arXiv:2606. 10246v1 Announce Type: cross Abstract: Maliciously-created fake speech, including deepfaked and spoofed audio, is proliferating at an alarming rate, and detection models are racing to stay ahead of the curve.

By Ashley R. Keaton, Zahra Khanjani, Christine Mallinson, Vandana P. Janeja