Telephone fraud is pervasive and costly, but its inner workings are rarely observed at scale. We analyze a complete corpus of 10,211 inbound scam and spam calls -- 913 hours of audio and 330,956 trans...
arXiv:2607. 11707v1 Announce Type: cross Abstract: Following the rapid progress of generative Artificial Intelligence, there is a growing threat posed by conversational scams.
By Ahmed Omar Salim Adnan, Yogananda Manjunath, Shivanjali Khare
TeleAntiFraud 2.0 is a monthly‑frozen, audio‑based benchmark for telecom fraud detection that incorporates newly observed scam patterns while preserving earlier test sets. It uses a Mixed‑Tree Anti‑Fraud Generation Pipeline to create profile‑grounded scenarios, expands them into mixed‑tree dialogues, and renders validated speech for 900 Chinese calls (600 fraud, 300 near‑domain non‑fraud) each month. Experiments show that classifiers perform well against unrelated negatives but drop significantly against near‑domain negatives, highlighting the need for near‑domain construction and collapse‑aware reporting in realistic evaluation settings.
By Huiyuan Liu, Zhiming Ma, Yanxing Liu, Shun Zhang, Qifan Wang, Di Liu, Yifan Wang, Yuyang Deng, Haoyang Meng, Yijin Zhou, Yuxi Zhao, Chengxian Hu, Peidong Wang, Peng Chen
arXiv:2606. 27936v1 Announce Type: cross Abstract: The widespread collection of fine-grained location data by commercial data brokers creates a re-identification risk that is not widely recognised by the public.
By Oscar Thees, Roman M\"uller, Matthias Templ
The paper presents a cumulative turn‑based risk assessment framework for detecting financial scams targeting older adults, which aggregates conversational turns and updates risk estimates at each step. A multi‑turn dialogue dataset covering investment, charity, and tech support scams is created, with annotations for risk level, score, rationale, and safety recommendation at every cumulative stage. Four small language models (Phi‑4, LLaMA‑3.2, DeepSeek‑R1, Qwen3) are fine‑tuned; Phi‑4 and LLaMA‑3.2 outperform others in turn‑aware risk estimation, demonstrating that compact models can effectively support incremental scam detection in resource‑constrained, privacy‑aware deployments.
By Parviz Ghafariasl, Weimin Fu, Xiaolong Guo, Shing I. Chang
arXiv:2606. 27944v1 Announce Type: cross Abstract: Phone-use Agents can execute complex tasks end to end across real mobile applications.
By Yiming Sun, Chen Chen, Zifan Zhou, Mi Zhang
arXiv:2510. 14207v3 Announce Type: replace Abstract: Large Language Model (LLM) agents are powering a growing share of interactive web applications, yet remain vulnerable to misuse and harm.
By Trilok Padhi, Pinxian Lu, Abdulkadir Erol, Tanmay Sutar, Gauri Sharma, Mina Sonmez, Munmun De Choudhury, Ugur Kursuncu
arXiv:2507.15393v2 Announce Type: replace-cross
Abstract: Phishing email is a critical step in the cybercrime kill chain due to the high reachability of victims' email accounts and the low cost of la...
By Ruofan Liu, Yun Lin, Yuxin Wang, Xiwen Teoh, Zhenkai Liang, Gongshen Liu, Haojin Zhu, Jin Song Dong
arXiv:2606. 17555v1 Announce Type: cross Abstract: Banks simultaneously face signature-based fraud (card-not-present attacks, account takeover, ATM cloning) and behavioural financial crime (structuring, layering, mule networks, business email compromise) -- two threat families with fundamentally different detection requirements.
By Joseph Walusimbi, Joshua Benjamin Ssentongo
arXiv:2607. 10712v1 Announce Type: cross Abstract: Scientific fraud is the instrument of doubt that malicious entities can use to establish controversy in science.
By B\'alint Gyevn\'ar, Atoosa Kasirzadeh, Nihar B. Shah
arXiv:2607. 13078v1 Announce Type: cross Abstract: LLMs are now proposed for fraud detection, scam investigation, content moderation, and other trust-and-safety workflows.
By Keyur Gabani
arXiv:2607. 10252v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly consumed through opaque serving chains - API aggregators, resellers, and inference providers - in which the client has no technical means to confirm that the model answering is the model advertised, and recent audits show that a substantial fraction of commercial endpoints deviate from the vendor's reference weights.
By Tomas Bruckner