arXiv AI

Audio-Anchored Fusion of Multi-Ratio DiT Reconstruction Residuals for Cross-Domain Audio Deepfake Detection

arXiv:2607. 26472v1 Announce Type: cross Abstract: Audio deepfake detectors often degrade when generators, corpora, or recording conditions change.

arXiv Machine Learning
1d ago

NEUROTOKEN: Joint Source and Directional AAD with Envelope Decoding via Conditional Flow Matching

The paper introduces NEUROTOKEN, a unified neural network for auditory attention decoding (AAD) that jointly predicts the attended speaker’s direction and source by modeling the conditional likelihood of the attended envelope given EEG. It employs a conditional flow‑matching head (ATTUNEFLOW) and two inference‑time ensembles (QUADTRACK and ENV‑FLOW) to improve source‑AAD accuracy and reduce variance across subjects. Experiments on KU Leuven, DTU, and NJU datasets show significant gains over existing baselines and reveal that prior direction‑AAD results overestimate performance under stricter protocols.

By Ali Alavi, Donald S. Williamson
arXiv Machine Learning
Sep 3

Half-Truth Audio Detection and Localisation: A Lightweight Cross-Attentive Architecture and a Cross-Corpus Diagnostic Study

The paper introduces CAFNet, a lightweight cross‑attentive neural network that fuses MFCC, LFCC, and Chroma‑STFT features to detect and localise partially manipulated (half‑truth) speech. CAFNet achieves high ternary accuracy (97.55%) and low boundary mean absolute error (0.037 s) on the MLADDC benchmark, while demonstrating that cross‑corpus transfer depends on both capability and corpus characteristics. Ablation studies show that cross‑attention fusion is the most critical component, and removing a deeply supervised auxiliary head improves in‑domain performance and reduces variance.

By S. Sutharya, Remya K. Sasi
arXiv AI
Aug 17

Teffic-Audio: Tell Fact from Fiction

arXiv:2607. 28351v2 Announce Type: replace-cross Abstract: Speech deepfake detection has expanded in scope with increasingly heterogeneous spoofing mechanisms, including speech synthesis, voice conversion, vocoder reconstruction, and neural-codec resynthesis.

By Wan Lin, Li Wang, Jindong Wang, Kunyu Feng, Zhizheng Wu
arXiv AI
Jun 10

Whisfusion: Parallel ASR Decoding with Masked Diffusion

arXiv:2508. 07048v2 Announce Type: replace-cross Abstract: Autoregressive (AR) encoder-decoder models dominate high-quality multilingual ASR, but their left-to-right decoders make inference latency scale with transcript length.

By Taeyoun Kwon, Junhyuk Ahn, Taegeun Yun, Heeju Jwa, Yoonchae Choi, Siwon Park, Jongchan Kim, Hyungon Ryu, Hyuk-Jae Lee, Nam-Joon Kim
arXiv Computation and Language
Sep 17

Correlation-Guided Encoder Selection for Multi-Encoder Large Audio-Language Models

The paper introduces CUES, a lightweight heuristic for selecting encoder combinations in large audio‑language models by estimating complementarity through Pearson correlations of single‑encoder performance profiles. Using a frozen SmolLM2‑135M backbone, CUES consistently identifies optimal encoder sets for each track on the XARES‑LLM benchmark without requiring fusion training or test data. On broad audio tasks, CUES selects a diverse trio of encoders, improving performance by 4.3% over Whisper‑medium, while on text generation it opts for a focused speech‑only pair, outperforming mHuBERT‑147 by 6.3%. The results illustrate how correlation signals guide a diversity–interference trade‑off across different task families.

By Pei-Jun Liao, Hung-Shin Lee, Wenze Ren, Kuo-Hsuan Hung, Hung-yi Lee, Hsin-Min Wang
arXiv Machine Learning
Sep 10

Clean Accuracy Does Not Guarantee Provenance Robustness: A Prospective Codec-Stress Evaluation of Audio Attribution

The study evaluates audio provenance attribution systems, showing that high clean‑benchmark accuracy does not translate to robustness after codec compression. Using a prospectively registered protocol, the authors measured closed‑set attribution performance on two corpora after single‑stage codec transport, finding significant degradation—up to 70.3 Macro‑F1 points for WavLM‑Base+ and 61.0 for W2V2‑BERT 2.0—depending on codec settings and representation. The results demonstrate that clean accuracy alone cannot guarantee deployment robustness across different codecs and representations.

By Gang Shi (Independent Researcher)