← Back to all news
arXiv AI September 15, 2026 By Yusuf Khalid Shire, Sang-Chul Kim

PIDS-Bench: Evaluating Prompt-Injection Detectors Under Over-Defense, Obfuscation, and Distribution Shift

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

  • fine-tuning
  • benchmarks
  • safety

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Aug 21

MaliciousSkillBench: A Comprehensive Benchmark for Malicious Agent Skill Detection

arXiv:2608. 19901v1 Announce Type: cross Abstract: Agent Skills extend LLM agents with reusable instruction packages that may also include scripts, resources, and service configuration.

By Yue Wang, Yi Liu, Gelei Deng, Ying Zhang, Yuekang Li, Zhenyu Chen, Leo Zhang
llmsagentsbenchmarks
More like this →
arXiv Machine Learning
Jul 21

When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift

arXiv:2602. 14161v2 Announce Type: replace Abstract: Detecting prompt injection, jailbreak attacks, and harmful requests is critical for deploying LLM-based agents safely, yet current evaluation practices in this literature overestimate generalization.

By Max Fomin
llmsagentsbenchmarkssafety
More like this →
arXiv Machine Learning
Jun 18

Evaluating Prompting-Based Defenses Against Domain-Camouflaged Injection Attacks

arXiv:2606. 18530v1 Announce Type: cross Abstract: Domain-camouflaged injection attacks embed malicious instructions in retrieved content using domain-appropriate vocabulary, evading standard detectors that rely on syntactic injection markers.

By Aaditya Pai
llmsagentsbenchmarks
More like this →
arXiv Machine Learning
Aug 4

Crushing the Evidence: A Dual-Penalty Evasion Framework for Fooling White-Box Explainable AI Auditors

arXiv:2608. 00566v1 Announce Type: new Abstract: Post-hoc model explainers such as LIME, SHAP, and Integrated Gradients are widely deployed to audit models in high-stakes sensitive domains, including finance, healthcare, and social welfare.

By Niraj Kumar, Harsh Kasyap
ragbenchmarkssafety
More like this →
arXiv Machine Learning
Aug 11

BASIS: Breach-Aware Selective Prompt Injection Shielding with Prefill Attention Probes

arXiv:2608. 08027v1 Announce Type: cross Abstract: Prompt injection is a critical security threat in large language model (LLM) applications, where attackers hijack model behavior by embedding malicious instructions in user or external data.

By Laiqiao Qin, Tianqing Zhu, Longxiang Gao, Wanlei Zhou
llmsrag
More like this →
arXiv AI
Jun 19

Calibration Without Comprehension: Diagnosing the Limits of Fine-Tuning LLMs for Vulnerability Detection in Systems Software

arXiv:2606. 20502v1 Announce Type: cross Abstract: Whether LLMs scoring well on vulnerability benchmarks genuinely reason about security or merely pattern-match on contaminated data remains unresolved.

By Arastoo Zibaeirad, Marco Vieira
llmsfine-tuningbenchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea