Anatomy of a Scam Call: What 10,000 real scam and spam calls reveal about how phone scammers operate
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The study analyzes 10,211 real scam and spam calls collected by an AI voice‑agent honeypot, revealing that scammers operate on a templated, office‑hour schedule and use disposable numbers to recycle scripts. Callers predominantly seek identity anchors such as home addresses and dates of birth, and the amount of conversation increases with the target’s age, though the requested information remains unchanged. Early detection is feasible, with escalation predictability reaching 0.87 ROC‑AUC by the eighth line using simple bag‑of‑words models.
arXiv:2607. 11707v1 Announce Type: cross Abstract: Following the rapid progress of generative Artificial Intelligence, there is a growing threat posed by conversational scams.
TeleAntiFraud 2.0 is a monthly‑frozen, audio‑based benchmark for telecom fraud detection that incorporates newly observed scam patterns while preserving earlier test sets. It uses a Mixed‑Tree Anti‑Fraud Generation Pipeline to create profile‑grounded scenarios, expands them into mixed‑tree dialogues, and renders validated speech for 900 Chinese calls (600 fraud, 300 near‑domain non‑fraud) each month. Experiments show that classifiers perform well against unrelated negatives but drop significantly against near‑domain negatives, highlighting the need for near‑domain construction and collapse‑aware reporting in realistic evaluation settings.
arXiv:2606. 27944v1 Announce Type: cross Abstract: Phone-use Agents can execute complex tasks end to end across real mobile applications.
The paper presents a cumulative turn‑based risk assessment framework for detecting financial scams targeting older adults, which aggregates conversational turns and updates risk estimates at each step. A multi‑turn dialogue dataset covering investment, charity, and tech support scams is created, with annotations for risk level, score, rationale, and safety recommendation at every cumulative stage. Four small language models (Phi‑4, LLaMA‑3.2, DeepSeek‑R1, Qwen3) are fine‑tuned; Phi‑4 and LLaMA‑3.2 outperform others in turn‑aware risk estimation, demonstrating that compact models can effectively support incremental scam detection in resource‑constrained, privacy‑aware deployments.
arXiv:2606. 27936v1 Announce Type: cross Abstract: The widespread collection of fine-grained location data by commercial data brokers creates a re-identification risk that is not widely recognised by the public.