arXiv Machine Learning By Sadib Hassan Rumman, Md. Shariful Islam, Md. Rayhanur Rahman

Cross-Corpus Evaluation of Generalizable Vulnerability Detection in IoT Firmware

Read the original on arXiv Machine Learning →

arXiv:2608. 11492v1 Announce Type: cross Abstract: IoT firmware vulnerability detection remains challenging due to heterogeneous firmware ecosystems, resource-constrained platforms, and limitations in existing benchmarks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 21

CESBench: Benchmarking Large Language Models on Cryptographic Engineering Security for IoT Devices

CESBench is a new benchmark for evaluating large language models on cryptographic engineering security for IoT devices, comprising 380 expert‑written items across six sub‑domains such as side‑channel, fault injection, and implementation. The benchmark includes four task types—multiple‑choice, judgment, scenario, and code—each designed to test different competences, with automatic scoring for the first two and LLM‑based judging for the latter two. Evaluation of 11 open‑weight and proprietary LLMs shows strong performance on multiple‑choice and code tasks but weaker results on judgment and scenario tasks, highlighting gaps in justifying security verdicts.

By Wenquan Zhou, An Wang, Jing Liang, Peien Feng, Jingqi Zhang, Yaoling Ding, Liehuang Zhu