arXiv Computation and Language

FRAUDSkill: Structured Frozen-Weight Skill Optimization for Audio Anti-Fraud Detection

FRAUDSkill is a structured frozen‑weight adaptation framework for audio anti‑fraud detection that keeps the underlying audio‑language model unchanged while optimizing external skill programs, route‑specific policies, and decision rules. It combines structured output control with validation‑guided multi‑path inference to produce protocol‑compliant predictions. On the TeleAntiFraud benchmark, FRAUDSkill achieves a 73.50% Macro‑F1 score, outperforming the shared frozen‑model baseline by 31.96% and reducing invalid outputs to 1.94%.

arXiv Computation and Language
1d ago

Before the Warning Comes Too Late: Incremental Phone-Scam Detection from Speech

The paper introduces StreamFraudNet, a weakly supervised model that detects phone scams from raw telephone audio in an incremental fashion. It processes audio through overlapping windows with a frozen self‑supervised encoder, uses recurrent temporal modeling, and aggregates window scores to update predictions every two seconds. On an English benchmark, the model achieves a ROC‑AUC of 0.9953, outperforming baselines while producing its first prediction after 10 seconds and running faster than real time.

By Khang Nhat Hoang Vo, Anh Trac Duc Dinh, Tai Tien Ta, Tho Quan
arXiv Computation and Language
2d ago

TeleAntiFraud 2.0: A Refreshable, Profile-Grounded, and Audio-Based Benchmark for Telecom Fraud Detection

TeleAntiFraud 2.0 is a monthly‑frozen, audio‑based benchmark for telecom fraud detection that incorporates newly observed scam patterns while preserving earlier test sets. It uses a Mixed‑Tree Anti‑Fraud Generation Pipeline to create profile‑grounded scenarios, expands them into mixed‑tree dialogues, and renders validated speech for 900 Chinese calls (600 fraud, 300 near‑domain non‑fraud) each month. Experiments show that classifiers perform well against unrelated negatives but drop significantly against near‑domain negatives, highlighting the need for near‑domain construction and collapse‑aware reporting in realistic evaluation settings.

By Huiyuan Liu, Zhiming Ma, Yanxing Liu, Shun Zhang, Qifan Wang, Di Liu, Yifan Wang, Yuyang Deng, Haoyang Meng, Yijin Zhou, Yuxi Zhao, Chengxian Hu, Peidong Wang, Peng Chen
arXiv AI
Jun 10

Linguistically Augmented Audio Speech Data (LinguAS)

arXiv:2606. 10246v1 Announce Type: cross Abstract: Maliciously-created fake speech, including deepfaked and spoofed audio, is proliferating at an alarming rate, and detection models are racing to stay ahead of the curve.

By Ashley R. Keaton, Zahra Khanjani, Christine Mallinson, Vandana P. Janeja
arXiv AI
Jul 28

Traceable LLM Reasoning for Fake-Order Fraud Detection

arXiv:2607. 23075v1 Announce Type: cross Abstract: Detecting fake-order fraud at scale remains a critical challenge for large online-to-offline (O2O) service platforms, as existing approaches often rely on expert-designed features, produce black-box decisions, and provide limited interpretability.

By Siqi You, Bingsong Xu, Zhixian Zheng, Xinjian Peng, Yang Xie, Ying Wang, Jiarong Xu
arXiv AI
Sep 12

Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems

The paper introduces the Spectral Masking and Interpolation Attack (SMIA), a black‑box adversarial technique that subtly alters inaudible frequency regions of AI‑generated audio to fool voice authentication systems and their countermeasures. Experiments show SMIA achieves at least 82% success against combined verification and countermeasure systems, 97.5% against standalone speaker verification, and 100% against countermeasures, revealing a critical security gap. The authors argue that current static defenses are inadequate and call for dynamic, context‑aware defenses that can adapt to evolving threats.

By Kamel Kamel, Hridoy Sankar Dutta, Keshav Sood, Sunil Aryal