CESBench: Benchmarking Large Language Models on Cryptographic Engineering Security for IoT Devices
Read the original on arXiv AI →CESBench is a new benchmark for evaluating large language models on cryptographic engineering security for IoT devices, comprising 380 expert‑written items across six sub‑domains such as side‑channel, fault injection, and implementation. The benchmark includes four task types—multiple‑choice, judgment, scenario, and code—each designed to test different competences, with automatic scoring for the first two and LLM‑based judging for the latter two. Evaluation of 11 open‑weight and proprietary LLMs shows strong performance on multiple‑choice and code tasks but weaker results on judgment and scenario tasks, highlighting gaps in justifying security verdicts.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.