arXiv AI By Wenquan Zhou, An Wang, Jing Liang, Peien Feng, Jingqi Zhang, Yaoling Ding, Liehuang Zhu

CESBench: Benchmarking Large Language Models on Cryptographic Engineering Security for IoT Devices

Read the original on arXiv AI →

CESBench is a new benchmark for evaluating large language models on cryptographic engineering security for IoT devices, comprising 380 expert‑written items across six sub‑domains such as side‑channel, fault injection, and implementation. The benchmark includes four task types—multiple‑choice, judgment, scenario, and code—each designed to test different competences, with automatic scoring for the first two and LLM‑based judging for the latter two. Evaluation of 11 open‑weight and proprietary LLMs shows strong performance on multiple‑choice and code tasks but weaker results on judgment and scenario tasks, highlighting gaps in justifying security verdicts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 2

Toward Cybersecurity-Expert Small Language Models

arXiv:2510. 14113v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are transforming everyday applications, yet deployment in cybersecurity lags due to a lack of high-quality, domain-specific models and training datasets.

By Matan Levi, Daniel Ohayon, Ariel Blobstein, Ravid Sagi, Ian Molloy, Yair Allouche