arXiv:2607. 26472v1 Announce Type: cross Abstract: Audio deepfake detectors often degrade when generators, corpora, or recording conditions change.
By Haotian Mo, Jie Liu, Siqi Shen, Songzhu Mei, Xinhai Chen, Xiangyang Wang, Yigui Feng, Shuai Li, Gencheng Liu, Keqi Yang, Qinglin Wang
arXiv:2511. 21325v2 Announce Type: replace-cross Abstract: Deepfake (DF) audio detectors still struggle to generalize to out of distribution inputs.
By Ido Nitzan Hidekel, Gal lifshitz, Khen Cohen, Dan Raviv
arXiv:2609.08899v2 Announce Type: replace-cross
Abstract: Speech deepfakes can mimic a speaker's voice convincingly enough to deceive listeners and automated systems. This has driven strong progress...
By Mengzhe Geng, Yujia Lu, Patrick Littell, Manuela Kunz, Xie Chen
arXiv:2606. 10223v1 Announce Type: cross Abstract: Attributing a synthetic utterance to its originating system remains an open challenge: closed-set models fail to reject unseen synthesizers and produce overconfident predictions.
By Awais Khan, Kutub Uddin, Khalid Malik
arXiv:2607. 28351v2 Announce Type: replace-cross Abstract: Speech deepfake detection has expanded in scope with increasingly heterogeneous spoofing mechanisms, including speech synthesis, voice conversion, vocoder reconstruction, and neural-codec resynthesis.
By Wan Lin, Li Wang, Jindong Wang, Kunyu Feng, Zhizheng Wu
arXiv:2609.13842v1 Announce Type: cross
Abstract: Recent advances in speech synthesis and voice conversion have made deepfake speech increasingly realistic, making generalization to unseen spoofing a...
By Minh-Xuan Phan, Khalid Zaman, Candy Olivia Mawalim, Masashi Unoki
The paper introduces CAFNet, a lightweight cross‑attentive neural network that fuses MFCC, LFCC, and Chroma‑STFT features to detect and localise partially manipulated (half‑truth) speech. CAFNet achieves high ternary accuracy (97.55%) and low boundary mean absolute error (0.037 s) on the MLADDC benchmark, while demonstrating that cross‑corpus transfer depends on both capability and corpus characteristics. Ablation studies show that cross‑attention fusion is the most critical component, and removing a deeply supervised auxiliary head improves in‑domain performance and reduces variance.
By S. Sutharya, Remya K. Sasi
The paper introduces CUES, a lightweight heuristic for selecting encoder combinations in large audio‑language models by estimating complementarity through Pearson correlations of single‑encoder performance profiles. Using a frozen SmolLM2‑135M backbone, CUES consistently identifies optimal encoder sets for each track on the XARES‑LLM benchmark without requiring fusion training or test data. On broad audio tasks, CUES selects a diverse trio of encoders, improving performance by 4.3% over Whisper‑medium, while on text generation it opts for a focused speech‑only pair, outperforming mHuBERT‑147 by 6.3%. The results illustrate how correlation signals guide a diversity–interference trade‑off across different task families.
By Pei-Jun Liao, Hung-Shin Lee, Wenze Ren, Kuo-Hsuan Hung, Hung-yi Lee, Hsin-Min Wang
The paper introduces modality‑gated deep adapters, a parameter‑efficient method for adding new modalities to a frozen multimodal embedding language model without altering its existing outputs. These adapters are bottleneck modules attached to each decoder layer, grouped into modality‑specific packs that activate only during encoding of their own modality, ensuring exact preservation of the base model’s computation graph. Experiments on a 2B base model show significant gains in audio‑to‑text and thermal‑to‑text retrieval metrics, and the authors release the audio and thermal packs along with training and evaluation code.
By Abdul Basit Tonmoy, Kazi Fardinul Hoque, Md. Shahrier Islam Arham, Arman Luthra
arXiv:2609.23830v1 Announce Type: new
Abstract: Comparing audio-visual deepfake detectors requires coordinating dataset adaptation, temporal input representation, model interfaces and experimental co...
By Jan Rybarczyk, Mateusz Roszkowski, Jacek Komorowski
The study evaluates audio provenance attribution systems, showing that high clean‑benchmark accuracy does not translate to robustness after codec compression. Using a prospectively registered protocol, the authors measured closed‑set attribution performance on two corpora after single‑stage codec transport, finding significant degradation—up to 70.3 Macro‑F1 points for WavLM‑Base+ and 61.0 for W2V2‑BERT 2.0—depending on codec settings and representation. The results demonstrate that clean accuracy alone cannot guarantee deployment robustness across different codecs and representations.
By Gang Shi (Independent Researcher)
arXiv:2508. 07048v2 Announce Type: replace-cross Abstract: Autoregressive (AR) encoder-decoder models dominate high-quality multilingual ASR, but their left-to-right decoders make inference latency scale with transcript length.
By Taeyoun Kwon, Junhyuk Ahn, Taegeun Yun, Heeju Jwa, Yoonchae Choi, Siwon Park, Jongchan Kim, Hyungon Ryu, Hyuk-Jae Lee, Nam-Joon Kim