arXiv:2608. 02084v1 Announce Type: cross Abstract: Binary function embedding models are trained to encode the semantics of binary code in such a way that they can be generalized to a variety of reverse engineering tasks, such as binary code search, vulnerability detection, or malware classification.
By Samuel Valenzuela, Johannes Kinder
arXiv:2608. 11766v1 Announce Type: cross Abstract: Binary code representation learning is a fundamental problem in software security and reverse engineering.
By Huaijin Wang, Shuai Wang
The paper introduces LSem2Vec, a two‑stage method that first uses a large language model to extract source code semantics and then applies a sentence embedding model to produce vector representations. This approach removes the need for task‑specific training or fine‑tuning, addressing errors in LLM outputs. Experiments on three datasets across multiple programming languages show that LSem2Vec outperforms five state‑of‑the‑art unsupervised methods.
By Zixiang Xian, Chenhui Cui, Rubing Huang, Chunrong Fang, Zhenyu Chen
arXiv:2605. 16046v2 Announce Type: replace-cross Abstract: Semantic code search has been widely adopted in both academia and industry.
By Yiming Liu, Ruofan Liu, Yun Lin, Zicong Zhang, Weiyu Kong, Pengnian Qi, Xiao Cheng, Weinan Zhang, Qianxiang Wang, Linpeng Huang
The paper surveys 87 influential studies on machine‑learning‑based automated vulnerability detection (ML4AVD), categorizing them by problem formulation, input and detection granularity, target languages, evaluation metrics, datasets, and detection approaches. It identifies twelve self‑reinforcing pain points—such as overreliance on binary classification of C/C++ function‑level vulnerabilities, limited language coverage, and intertwined datasets, baselines, and metrics—that trap the field in a narrow, artificial problem space. The authors propose concrete recommendations to break these feedback loops and evaluate a recent high‑profile effort, AIxCC, against these guidelines, reflecting on ML4AVD’s relevance amid the rise of agentic AI.
By Dan Ristea, Shae McFadden, Ezzeldin Shereen, Madeleine Dwyer, Sanyam Vyas, Chris Hicks, Vasilios Mavroudis
arXiv:2608. 07894v1 Announce Type: new Abstract: Recent advances have highlighted the potential of machine learning, particularly Large Language Models (LLMs), for analyzing and optimizing programs.
By Calvin Higgins, Marco Alvarez
arXiv:2505. 03818v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) can achieve strong performance on everyday coding tasks, but they can fail on complex tasks that require non-trivial reasoning about program semantics.
By Antonio Valerio Miceli-Barone, Vaishak Belle, Ali Payani
The paper investigates whether code large language models (CodeLLMs) inadvertently reproduce proprietary or sensitive code by evaluating seven state‑of‑the‑art training data detection (TDD) methods on eight CodeLLMs. It introduces CodeSnitch, a benchmark of 9,000 function‑level code samples across three languages, each labeled as included or excluded from training data, and applies mutation strategies based on the Type‑1 to Type‑4 code clone taxonomy to test TDD robustness. The study offers a systematic assessment of current TDD techniques for code and suggests directions for developing more effective detection methods.
By Tianlin Li, Yunxiang Wei, Zhiming Li, Aishan Liu, Qing Guo, Xianglong Liu, Dongning Sun, Yang Liu
arXiv:2606. 02834v1 Announce Type: cross Abstract: Malware analysis starts with the raw bytes of an executable program, and tools to "lift" these to higher-level representations, such as assembly, are expensive and subject to error.
By Florian St\"ortz, Catalin-Andrei Stan, Alexandru Dinu, Sandra Servia-Rodr\'iguez, Mihaela Gaman, Calin Miron, Edward Raff
arXiv:2606. 16244v1 Announce Type: cross Abstract: Large language models routinely generate code with exploitable security flaws.
By Xiaoyun Xu, Lichao Wu, Jona te Lintelo, Siyu Zhang, Stjepan Picek
SpIDER is a dense retrieval method that combines LLM reasoning with graph-based exploration of codebases to locate relevant functions, classes, or files for user queries. It introduces a graph-structured benchmark, SpIDER-Bench, covering multiple programming languages and demonstrates significant recall improvements over traditional BM25 and dense approaches. The method’s graph-based candidate expansion provides auditable structural reasons for each retrieved item while keeping the retrieval budget fixed.
By Shravan Chaudhari, Rahul Thomas Jacob, Jiajun Cao, Shihab Rashid, Mononito Goswami, Christian Bock
arXiv:2509. 14335v2 Announce Type: replace-cross Abstract: Automated malware classifiers achieve strong detection performance, but auditing requires more than flagging a sample: analysts must explain malicious behaviors and justify them with code evidence.
By Xinran Zheng, Xingzhi Qian, Yiling He, Shuo Yang, Lorenzo Cavallaro