arXiv AI

Poster: Exploring the Limits of Audio-Based Detection of Turkish Phone Call Scams

arXiv:2606. 24523v1 Announce Type: cross Abstract: Scam phone calls exploit vulnerable communities worldwide, yet research on detection has focused almost exclusively on English and other high-resource languages.

arXiv Machine Learning
Sep 25

A Corpus of Real Scam- and Spam-Call Conversations from an Active Voice-Agent Honeypot

The paper introduces a dataset of 10,015 real scam and spam phone calls collected over 53 days using an active voice‑agent honeypot. Each call is recorded, transcribed, and automatically labeled, yielding 328,869 turn‑level transcripts and 895 hours of audio from 5,665 distinct numbers. The corpus distinguishes between predatory‑but‑legal lead generation and outright scams, with labels validated by human review and technical checks on realism.

By Ethan Traister, Dennis Tsang Ng, Siyu Zhang, Huaiyu Guo, Tommy Duong, Tyler Wu, Yuchen Zhou, Xingyu Shen, Jiaqi Wu, Simiao Ren
arXiv Computation and Language
Sep 17

TeleAntiFraud 2.0: A Refreshable, Profile-Grounded, and Audio-Based Benchmark for Telecom Fraud Detection

TeleAntiFraud 2.0 is a monthly‑frozen, audio‑based benchmark for telecom fraud detection that incorporates newly observed scam patterns while preserving earlier test sets. It uses a Mixed‑Tree Anti‑Fraud Generation Pipeline to create profile‑grounded scenarios, expands them into mixed‑tree dialogues, and renders validated speech for 900 Chinese calls (600 fraud, 300 near‑domain non‑fraud) each month. Experiments show that classifiers perform well against unrelated negatives but drop significantly against near‑domain negatives, highlighting the need for near‑domain construction and collapse‑aware reporting in realistic evaluation settings.

By Huiyuan Liu, Zhiming Ma, Yanxing Liu, Shun Zhang, Qifan Wang, Di Liu, Yifan Wang, Yuyang Deng, Haoyang Meng, Yijin Zhou, Yuxi Zhao, Chengxian Hu, Peidong Wang, Peng Chen
arXiv Computation and Language
Sep 18

Before the Warning Comes Too Late: Incremental Phone-Scam Detection from Speech

The paper introduces StreamFraudNet, a weakly supervised model that detects phone scams from raw telephone audio in an incremental fashion. It processes audio through overlapping windows with a frozen self‑supervised encoder, uses recurrent temporal modeling, and aggregates window scores to update predictions every two seconds. On an English benchmark, the model achieves a ROC‑AUC of 0.9953, outperforming baselines while producing its first prediction after 10 seconds and running faster than real time.

By Khang Nhat Hoang Vo, Anh Trac Duc Dinh, Tai Tien Ta, Tho Quan
arXiv AI
Jun 10

Linguistically Augmented Audio Speech Data (LinguAS)

arXiv:2606. 10246v1 Announce Type: cross Abstract: Maliciously-created fake speech, including deepfaked and spoofed audio, is proliferating at an alarming rate, and detection models are racing to stay ahead of the curve.

By Ashley R. Keaton, Zahra Khanjani, Christine Mallinson, Vandana P. Janeja
arXiv Computation and Language
Sep 17

FRAUDSkill: Structured Frozen-Weight Skill Optimization for Audio Anti-Fraud Detection

FRAUDSkill is a structured frozen‑weight adaptation framework for audio anti‑fraud detection that keeps the underlying audio‑language model unchanged while optimizing external skill programs, route‑specific policies, and decision rules. It combines structured output control with validation‑guided multi‑path inference to produce protocol‑compliant predictions. On the TeleAntiFraud benchmark, FRAUDSkill achieves a 73.50% Macro‑F1 score, outperforming the shared frozen‑model baseline by 31.96% and reducing invalid outputs to 1.94%.

By Chengxian Hu, Zhiming Ma, Mingjun Pan, Yifan Wang, Shun Zhang, Qifan Wang, Zhilei Zhao, Yijin Zhou, Yuxi Zhao, Huiyuan Liu, Peidong Wang, Peng Chen
arXiv Computation and Language
Sep 7

Auditing Bias and Safety in Voice AI Customer Care

The paper introduces a validation‑gated audit framework for voice AI customer‑care systems, treating them as stateful, multi‑turn, tool‑mediated interactions where bias and safety can manifest as added burdens before a final decision. The framework distinguishes between native speech‑to‑speech, cascaded ASR‑to‑LM‑to‑TTS, and hybrid architectures, and applies matched service facts across controlled caller presentation conditions to validate fact invariance, presentation cues, artifacts, and acoustic measurements. It outlines seven validation gates, a six‑family metric set, and demonstrates the approach with a synthetic refund‑dispute audit example, while noting that production results are withheld until the protocol is satisfied.

By Vignesh Ethiraj, Ashwath David
arXiv AI
Sep 12

A Survey of Threats Against Voice Authentication and Anti-Spoofing Systems

The paper reviews how voice authentication has evolved from handcrafted acoustic features to deep learning speaker embeddings, expanding its use in finance, smart devices, and law enforcement. It surveys modern threats—including data poisoning, adversarial, deepfake, and adversarial spoofing attacks—tracing their development alongside technological advances. For each attack type, the authors summarize methods, datasets, performance, and limitations, and organize the literature using accepted taxonomies to highlight emerging risks and open challenges.

By Kamel Kamel, Keshav Sood, Hridoy Sankar Dutta, Sunil Aryal
arXiv AI
Sep 25

To Trust or Not to Trust: Retrieval-Augmented Fact Checking in Speech

The paper introduces VeriSpeak, a benchmark of 3,879 spoken claims for evaluating fact verification in Large Audio Language Models (LALMs). It shows a clear modality gap: models that verify written claims well often fail on spoken versions, and retrieval alone offers limited improvement. Combining retrieval with explicit reasoning yields the best performance, reaching 86.1% accuracy and demonstrating the need for grounded reasoning over retrieved evidence in speech misinformation detection.

By Debajyoti Mazumder, Mamta, Abhirama Subramanyam Penamakuri