arXiv Machine Learning

Ruby: Unmasking Unsafe Rust in Stripped Binaries via Machine Learning

arXiv:2211. 00111v3 Announce Type: replace-cross Abstract: Rust, as an emerging system programming language, introduces $\texttt{unsafe}$ to allow developers to bypass safety checks during compilation.

arXiv AI
Jul 7

RustMizan: A Compilable, Contamination-Aware Benchmarking Framework for Rust Vulnerabilities

arXiv:2607. 04729v1 Announce Type: cross Abstract: LLM agents are increasingly applied to vulnerability analysis, but existing benchmarks have not kept pace.

By Tarek Elsayed, Shiping Yang, Eunsong Koh, Sanika Goyal, Vincent Huang, Paul Ngo, Nathan Young, Mohammad Omidvar Tehrani, Alvyn Kang, Arnell Kang, Zeyu Chen, Ang\'elica Moreira, Xuan Feng, Angel X. Chang, Nick Sumner, Steven Y. Ko
arXiv AI
Sep 7

When LLM Decompilers Recompile More and Preserve Less

The paper examines how large‑language‑model (LLM) decompilers, which produce clean, idiomatic C code, are currently evaluated mainly on recompilability and passing shipped tests. It shows that these metrics can mask significant behavioral differences: a decompiled function may recompile and pass all tests yet diverge on other inputs or lose disclosed vulnerabilities. To address this, the authors propose Decompile‑Diverge, a behavioral oracle that synthesizes drivers, fuzzes inputs, and compares the decompiled code’s behavior to the original, revealing divergences in up to 13% of cases and exposing gaps in current evaluation suites.

By Chang Liu, Edward Raff, Kristopher Micinski
arXiv Machine Learning
5d ago

What Do They Fix? LLM-Aided Categorization of Security Patches for Critical Memory Bugs

The paper introduces DUALLM, a dual-method pipeline that uses a Large Language Model and a fine‑tuned small language model to classify Linux kernel security patches with high precision. By analyzing commit titles, messages, diffs, and code context, DUALLM achieves 87.4% accuracy and an F1‑score of 0.875, outperforming existing methods. It successfully identified 111 recent patches addressing out‑of‑bounds or use‑after‑free vulnerabilities, with 90 confirmed true positives and proof‑of‑concept exploits demonstrating the validity of the classifications.

By Xingyu Li (UC Riverside), Juefei Pu (UC Riverside), Yifan Wu (UC Riverside), Xiaochen Zou (UC Riverside), Shitong Zhu (UC Riverside), Qiushi Wu (UC Riverside), Zheng Zhang (UC Riverside), Joshua Hsu (UC Riverside), Yue Dong (UC Riverside), Zhiyun Qian (UC Riverside), Kangjie Lu (UC Riverside), Trent Jaeger (UC Riverside), Michael De Lucia (UC Riverside), Srikanth V. Krishnamurthy (UC Riverside)
arXiv AI
Sep 11

CS-Guard: Benchmarking LLM Guardrails for Code Generation Security

CS-Guard is a new benchmark that systematically evaluates guardrails for code generation security, covering 1,000 malware-generation prompts, 7 jailbreak attacks, and a novel fictional scenario attack (FSA) for text-to-code generation, as well as 331 code prompts for code-to-code generation. The study empirically tests nine guardrails across seven large language models, finding that many guardrails fail to prevent malicious code generation, with attack success rates reaching about 50% for text-to-code and up to nearly 100% for code-to-code and FSA scenarios. CS-Guard introduces a modular three-layer guardrail taxonomy and releases its benchmark and data to support future research.

By Jinyang Li, Mingyu Guo, Hung X. Nguyen
arXiv AI
Sep 7

The History Is the Detector: Executing CVE Patch History, End-to-End

The paper introduces BUGSTONE‑E2E, a framework that converts vulnerability history into executable detection rules and validates them. It mines reusable rules from fixing commits, organizes them by CWE and language, and applies a funnel‑shaped pipeline that starts with lightweight analysis and culminates in LLM‑guided inspection, runtime verification, and patch generation. Using 19,325 high‑severity CVEs, the system identified 2,710 fixing commits, created 1,033 detection rules across 56 CWE families, and produced runtime evidence for 644 findings in 14 programs.

By Qiushi Wu, Kevin Eykholt, Youngja Park, Xiaokui Shu, Dhilung Kirat, Douglas Lee Schales, Ian Molloy