arXiv Machine Learning

AURA: Adaptive Uncertainty-Routed Analysis for Email Threat Detection

AURA: Adaptive Uncertainty-Routed Analysis for Email Threat Detection is a multimodal system that evaluates both email content and embedded URLs to detect spam and phishing. It uses a two-layer approach: first, a URL classifier estimates prediction uncertainty, and only messages with high uncertainty are passed to a fine-tuned transformer encoder for deeper semantic analysis. Evaluated on eight diverse training corpora and two real-world datasets covering a decade of attacks, AURA achieves a macro F1-score of 0.9858 in-distribution and maintains scores above 0.94 on the NazPhish-Eval and GuenterTrap-Eval datasets, demonstrating strong generalization to new attack scenarios.

arXiv Computation and Language
6d ago

Prompt Injection Detection for Email Agents Through Attack Chain Modeling

The paper introduces a prompt‑injection detection framework for email assistants that models attacks as a chain of stages. It combines a text detector, stage‑specific verifiers, rule‑based risk signals, user intent consistency checks, and a logistic decision policy. Experiments on five benchmarks show the framework outperforms pretrained detectors, achieving a mean F1 of 0.406 versus 0.216, and demonstrate that training on benign emails resembling attacks reduces false alarms.

By Ahmad Hashmi, Dhyey Patel, Yunting Yin
arXiv AI
Sep 24

WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents

WAInjectBench introduces the first comprehensive benchmark for detecting prompt injection attacks against web agents, offering a fine‑grained categorization of threats and datasets that include malicious and benign text and image samples. The study systematically evaluates both text‑based and image‑based detection methods across multiple scenarios, revealing that detectors perform well on attacks with explicit instructions or visible perturbations but struggle with subtle or instruction‑free attacks. The authors release the datasets and code to facilitate further research in this area.

By Yinuo Liu, Xilong Wang, Ruohan Xu, Yuqi Jia, Neil Zhenqiang Gong
arXiv AI
Sep 10

CogniDir: Combating Cognitive Malicious Comments via Adaptive Distributional Learning for Robust Fake News Detection

CogniDir is an adaptive distributional learning framework designed to improve fake news detection against new psychologically grounded malicious comments generated by Large Language Models. It reframes robust detection as a dynamic data mixture optimization problem, using cognitive psychology to formalize adversarial paradigms and an information‑theoretic score to guide adaptive sampling of training data. Experiments on three benchmarks show that CogniDir achieves state‑of‑the‑art robustness, boosting F1 scores by up to 17.9% over existing baselines under heterogeneous AI‑generated attacks.

By Zhao Tong, Chunlin Gong, Yimeng Gu, Haichao Shi, Qiang Liu, Shu Wu, Xingcheng Xu, Xiao-Yu Zhang
arXiv Machine Learning
1d ago

Evasion Attacks: How Adversarial Noise Bypasses ML Classifiers

The paper reports a reproducible study of evasion attacks on image and text classifiers. A compact convolutional network on MNIST achieved 98.63% clean accuracy but dropped to 60.20% under FGSM with ε=0.15 and 1.72% with ε=0.30, while PGD reduced accuracy to 32.47% and 0.41%; a bit‑depth‑reduction defense only partially restored performance. In contrast, a DistilBERT model fine‑tuned on the SMS Spam Collection reached 98.75% accuracy and 94.96% F1‑score, yet a sequence of predefined perturbations produced only modest probability shifts and did not flip spam to ham predictions.

By Parker Hummel (Minot State University), Ryne Skabo (Minot State University), Muhammad Abusaqer (Minot State University)