TeleAntiFraud 2.0 is a monthly‑frozen, audio‑based benchmark for telecom fraud detection that incorporates newly observed scam patterns while preserving earlier test sets. It uses a Mixed‑Tree Anti‑Fraud Generation Pipeline to create profile‑grounded scenarios, expands them into mixed‑tree dialogues, and renders validated speech for 900 Chinese calls (600 fraud, 300 near‑domain non‑fraud) each month. Experiments show that classifiers perform well against unrelated negatives but drop significantly against near‑domain negatives, highlighting the need for near‑domain construction and collapse‑aware reporting in realistic evaluation settings.
By Huiyuan Liu, Zhiming Ma, Yanxing Liu, Shun Zhang, Qifan Wang, Di Liu, Yifan Wang, Yuyang Deng, Haoyang Meng, Yijin Zhou, Yuxi Zhao, Chengxian Hu, Peidong Wang, Peng Chen
arXiv:2606. 28002v1 Announce Type: cross Abstract: Insurance fraud imposes substantial financial losses and operational inefficiencies, raising premiums and impacting trust among legitimate policyholders.
By Muhammad Shakeel Akram, Amal Htait, Abdul Hamid Sadka, Emma Meisingseth, Karishma Jaitly
arXiv:2606. 24523v1 Announce Type: cross Abstract: Scam phone calls exploit vulnerable communities worldwide, yet research on detection has focused almost exclusively on English and other high-resource languages.
By Arda Eren, Micheal Cheung, Youqian Zhang, Grace Ngai, Eugene Yujun Fu
The paper introduces a dataset of 10,015 real scam and spam phone calls collected over 53 days using an active voice‑agent honeypot. Each call is recorded, transcribed, and automatically labeled, yielding 328,869 turn‑level transcripts and 895 hours of audio from 5,665 distinct numbers. The corpus distinguishes between predatory‑but‑legal lead generation and outright scams, with labels validated by human review and technical checks on realism.
By Ethan Traister, Dennis Tsang Ng, Siyu Zhang, Huaiyu Guo, Tommy Duong, Tyler Wu, Yuchen Zhou, Xingyu Shen, Jiaqi Wu, Simiao Ren
The paper introduces StreamFraudNet, a weakly supervised model that detects phone scams from raw telephone audio in an incremental fashion. It processes audio through overlapping windows with a frozen self‑supervised encoder, uses recurrent temporal modeling, and aggregates window scores to update predictions every two seconds. On an English benchmark, the model achieves a ROC‑AUC of 0.9953, outperforming baselines while producing its first prediction after 10 seconds and running faster than real time.
By Khang Nhat Hoang Vo, Anh Trac Duc Dinh, Tai Tien Ta, Tho Quan
The paper "Beyond EER: Multi-Dimensional Evaluation of Information Leakage in Speaker De-Identification" argues that evaluating speaker de-identification systems solely by Equal Error Rate (EER) is insufficient. It proposes a holistic framework using five complementary metrics—EER, soft biometric leakage score, cumulative match characteristic re-identification analysis, canonical correlation analysis with Procrustes embedding alignment, and intelligibility via word error rate and semantic similarity—to capture independent dimensions of information leakage. Experiments on five IARPA ARTS SDID systems show that these metrics reveal leakage that a single metric would miss.
By Seungmin Seo, Oleg Aulov, P. Jonathon Phillips, Kevin Mangold, Jonathan Eskin
The paper introduces a validation‑gated audit framework for voice AI customer‑care systems, treating them as stateful, multi‑turn, tool‑mediated interactions where bias and safety can manifest as added burdens before a final decision. The framework distinguishes between native speech‑to‑speech, cascaded ASR‑to‑LM‑to‑TTS, and hybrid architectures, and applies matched service facts across controlled caller presentation conditions to validate fact invariance, presentation cues, artifacts, and acoustic measurements. It outlines seven validation gates, a six‑family metric set, and demonstrates the approach with a synthetic refund‑dispute audit example, while noting that production results are withheld until the protocol is satisfied.
By Vignesh Ethiraj, Ashwath David
The paper introduces VeriSpeak, a benchmark of 3,879 spoken claims for evaluating fact verification in Large Audio Language Models (LALMs). It shows a clear modality gap: models that verify written claims well often fail on spoken versions, and retrieval alone offers limited improvement. Combining retrieval with explicit reasoning yields the best performance, reaching 86.1% accuracy and demonstrating the need for grounded reasoning over retrieved evidence in speech misinformation detection.
By Debajyoti Mazumder, Mamta, Abhirama Subramanyam Penamakuri
arXiv:2609.39344v1 Announce Type: cross
Abstract: Speech transcripts used as long-term memory must preserve both words and stable speaker identities. Existing meeting-transcription metrics either ign...
By Shantanu Vispute, Aditya Mishra, Siddhartha Saxena
The article critiques the notion that a person's voice is a stable, unique biometric trace—termed a voiceprint—by reviewing historical, forensic, and technological evidence. It argues that voices are highly dynamic and context-dependent, and that the voiceprint metaphor misrepresents probabilistic speaker information as a fixed identity marker. The authors emphasize that speaker recognition should account for within-speaker variability, domain mismatch, and synthetic manipulation rather than rely on an assumed stable voiceprint.
By Tianle Yang, Cuiling Zhang, Chengzhe Sun, Siwei Lyu, Phil Rose
arXiv:2606. 06740v1 Announce Type: cross Abstract: Discrete speech units obtained via k-means clustering of self supervised embeddings entangle phonetic, speaker, and language information, causing speaker mixing and cross-lingual interference in multilingual multi-speaker speech generation.
By Naman Kothari, Arjun Gangwar, Adarsh Arigala, S Umesh
arXiv:2609.39162v1 Announce Type: cross
Abstract: Speaker diarization systems based on speaker embeddings and neural diarization exploit complementary forms of speaker information, but their intermed...
By Yehoshua Dissen, Joseph Keshet, Eduard Golshtein