The paper introduces a dataset of 10,015 real scam and spam phone calls collected over 53 days using an active voice‑agent honeypot. Each call is recorded, transcribed, and automatically labeled, yielding 328,869 turn‑level transcripts and 895 hours of audio from 5,665 distinct numbers. The corpus distinguishes between predatory‑but‑legal lead generation and outright scams, with labels validated by human review and technical checks on realism.
By Ethan Traister, Dennis Tsang Ng, Siyu Zhang, Huaiyu Guo, Tommy Duong, Tyler Wu, Yuchen Zhou, Xingyu Shen, Jiaqi Wu, Simiao Ren
The study analyzes 10,211 real scam and spam calls collected by an AI voice‑agent honeypot, revealing that scammers operate on a templated, office‑hour schedule and use disposable numbers to recycle scripts. Callers predominantly seek identity anchors such as home addresses and dates of birth, and the amount of conversation increases with the target’s age, though the requested information remains unchanged. Early detection is feasible, with escalation predictability reaching 0.87 ROC‑AUC by the eighth line using simple bag‑of‑words models.
By Ethan Traister, Ankit Raj, Jiaqi Gan, Xingyu Shen, Tyler Wu, Yuchen Zhou, Tommy Duong, Kidus Zewde, Siying Chen, Simiao Ren
The paper presents a cumulative turn‑based risk assessment framework for detecting financial scams targeting older adults, which aggregates conversational turns and updates risk estimates at each step. A multi‑turn dialogue dataset covering investment, charity, and tech support scams is created, with annotations for risk level, score, rationale, and safety recommendation at every cumulative stage. Four small language models (Phi‑4, LLaMA‑3.2, DeepSeek‑R1, Qwen3) are fine‑tuned; Phi‑4 and LLaMA‑3.2 outperform others in turn‑aware risk estimation, demonstrating that compact models can effectively support incremental scam detection in resource‑constrained, privacy‑aware deployments.
By Parviz Ghafariasl, Weimin Fu, Xiaolong Guo, Shing I. Chang
Telephone fraud is pervasive and costly, but its inner workings are rarely observed at scale. We analyze a complete corpus of 10,211 inbound scam and spam calls -- 913 hours of audio and 330,956 trans...
arXiv:2406. 13049v3 Announce Type: replace-cross Abstract: Personalized phishing is difficult to defend against because messages can be tailored to a target's work, interests, and social context.
By Jerson Francia, Derek Hansen, Benjamin Schooley, Matthew Taylor, Shydra Valynn Murray, Rebekah Cornelius, Greg Snow
arXiv:2602. 05056v2 Announce Type: replace-cross Abstract: Online scams increasingly leverage fluent and context-aware social engineering strategies, creating growing demand for AI systems that explain why a message may be risky.
By Heajun An, Connor Ng, Sandesh Sharma Dulal, Junghwan Kim, Jin-Hee Cho
arXiv:2606. 24523v1 Announce Type: cross Abstract: Scam phone calls exploit vulnerable communities worldwide, yet research on detection has focused almost exclusively on English and other high-resource languages.
By Arda Eren, Micheal Cheung, Youqian Zhang, Grace Ngai, Eugene Yujun Fu
arXiv:2607. 18429v1 Announce Type: cross Abstract: Phishing emails remain one of the most persistent cybersecurity threats, and machine-learning classifiers are widely used to detect them.
By Tanveer Ahmed, Seyedali Pourmoafil
arXiv:2507.15393v2 Announce Type: replace-cross
Abstract: Phishing email is a critical step in the cybercrime kill chain due to the high reachability of victims' email accounts and the low cost of la...
By Ruofan Liu, Yun Lin, Yuxin Wang, Xiwen Teoh, Zhenkai Liang, Gongshen Liu, Haojin Zhu, Jin Song Dong
arXiv:2512. 10104v2 Announce Type: cross Abstract: Email phishing is one of the most prevalent and globally consequential vectors of cyber intrusion.
By Najmul Hasan, Prashanth BusiReddyGari, Haitao Zhao, Yihao Ren, Jinsheng Xu, Shaohu Zhang
arXiv:2608. 15893v1 Announce Type: new Abstract: The rise of social media bots poses a persistent threat, enabling misinformation, opinion manipulation, and the erosion of trust in online platforms.
By Nof Orenstein, Yoni Birman
arXiv:2511. 12085v3 Announce Type: replace-cross Abstract: Phishing and related cyber threats are becoming increasingly sophisticated, with email-based phishing remaining the most persistent attack vector.
By Sajad U P