arXiv AI

Dual-Branch Gated Fusion for Open-Set Audio Deepfake Source Tracing

arXiv:2606. 10223v1 Announce Type: cross Abstract: Attributing a synthetic utterance to its originating system remains an open challenge: closed-set models fail to reject unseen synthesizers and produce overconfident predictions.

arXiv AI
3d ago

UniAE-MoE: A Unified Audio Encoder via Mixture of Experts

UniAE-MoE is a unified audio encoder that uses a Mixture-of-Experts architecture to model cross‑domain audio representations. It integrates encoder components from Qwen2‑Audio and Audio‑Flamingo 3, enhances them with SwiGLU and shared experts, and applies a two‑stage instruction‑tuning strategy along with task‑specific data scaling. The model achieves state‑of‑the‑art results on the XARES‑LLM benchmark (0.802) and tops the Interspeech 2026 Audio Encoder Capability Challenge, demonstrating strong generalization across speech, music, and general audio tasks.

By Shengbo Cai, Zhisheng Zhang, Zichao Nie, Jing Peng, Jingran Xie, Zhiyong Wu
arXiv Machine Learning
Sep 3

Half-Truth Audio Detection and Localisation: A Lightweight Cross-Attentive Architecture and a Cross-Corpus Diagnostic Study

The paper introduces CAFNet, a lightweight cross‑attentive neural network that fuses MFCC, LFCC, and Chroma‑STFT features to detect and localise partially manipulated (half‑truth) speech. CAFNet achieves high ternary accuracy (97.55%) and low boundary mean absolute error (0.037 s) on the MLADDC benchmark, while demonstrating that cross‑corpus transfer depends on both capability and corpus characteristics. Ablation studies show that cross‑attention fusion is the most critical component, and removing a deeply supervised auxiliary head improves in‑domain performance and reduces variance.

By S. Sutharya, Remya K. Sasi