arXiv:2512. 18542v3 Announce Type: replace-cross Abstract: AI coding assistants produce vulnerable code in 45\% of security-relevant scenarios~\cite{veracode2025}, yet no public training dataset teaches both traditional web security and AI/ML-specific defenses in a format suitable for instruction tuning.
By Scott Thornton
arXiv:2512. 10485v2 Announce Type: replace-cross Abstract: Vulnerability detection methods based on deep learning (DL) have shown strong performance on benchmark datasets, yet their real-world effectiveness remains underexplored.
By Chaomeng Lu, Bert Lagaisse
arXiv:2510. 13817v2 Announce Type: replace Abstract: The growth of IoT devices in shared environments has outpaced our ability to identify them, posing urgent risks to privacy, safety, and accountability.
By Rameen Mahmood, Tousif Ahmed, Sai Teja Peddinti, Danny Yuxing Huang
arXiv:2606. 20502v1 Announce Type: cross Abstract: Whether LLMs scoring well on vulnerability benchmarks genuinely reason about security or merely pattern-match on contaminated data remains unresolved.
By Arastoo Zibaeirad, Marco Vieira
CESBench is a new benchmark for evaluating large language models on cryptographic engineering security for IoT devices, comprising 380 expert‑written items across six sub‑domains such as side‑channel, fault injection, and implementation. The benchmark includes four task types—multiple‑choice, judgment, scenario, and code—each designed to test different competences, with automatic scoring for the first two and LLM‑based judging for the latter two. Evaluation of 11 open‑weight and proprietary LLMs shows strong performance on multiple‑choice and code tasks but weaker results on judgment and scenario tasks, highlighting gaps in justifying security verdicts.
By Wenquan Zhou, An Wang, Jing Liang, Peien Feng, Jingqi Zhang, Yaoling Ding, Liehuang Zhu
arXiv:2606. 03453v1 Announce Type: cross Abstract: Vulnerability disclosure volumes now far exceed organizational assessment capacity, yet three adjacent research communities (proof-of-concept generation, vulnerability prioritization, and detection rule engineering) operate largely in isolation.
By Farooq Shaikh