arXiv AI

Instruction Alignment for Binary Code Representation Learning

arXiv:2608. 11766v1 Announce Type: cross Abstract: Binary code representation learning is a fundamental problem in software security and reverse engineering.

arXiv AI
Jun 3

Large Byte Model: Teaching Language Models About Compiled Code

arXiv:2606. 02834v1 Announce Type: cross Abstract: Malware analysis starts with the raw bytes of an executable program, and tools to "lift" these to higher-level representations, such as assembly, are expensive and subject to error.

By Florian St\"ortz, Catalin-Andrei Stan, Alexandru Dinu, Sandra Servia-Rodr\'iguez, Mihaela Gaman, Calin Miron, Edward Raff
arXiv AI
Sep 25

Detecting Data Poisoning in Code Generation LLMs via Black-Box, Vulnerability-Oriented Scanning

The paper introduces CodeScan, a black-box, vulnerability-oriented scanning framework designed to detect data poisoning and backdoor attacks in code generation large language models (LLMs). CodeScan operates by analyzing structural similarities across multiple code generations, normalizing them with abstract syntax tree (AST) techniques, and then applying LLM-based vulnerability analysis to identify recurring insecure patterns. Evaluations on 117 models across three architectures and multiple sizes show over 97% detection accuracy with fewer false positives compared to prior methods.

By Shenao Yan, Shan Jin, Shimaa Ahmed, Sunpreet Singh Arora, Yiwei Cai, Yizhen Wang, Yuan Hong
arXiv AI
2d ago

Are AI Coders Snitches? An Empirical Study of Pretraining Data Detection on Code Large Language Models

The paper investigates whether code large language models (CodeLLMs) inadvertently reproduce proprietary or sensitive code by evaluating seven state‑of‑the‑art training data detection (TDD) methods on eight CodeLLMs. It introduces CodeSnitch, a benchmark of 9,000 function‑level code samples across three languages, each labeled as included or excluded from training data, and applies mutation strategies based on the Type‑1 to Type‑4 code clone taxonomy to test TDD robustness. The study offers a systematic assessment of current TDD techniques for code and suggests directions for developing more effective detection methods.

By Tianlin Li, Yunxiang Wei, Zhiming Li, Aishan Liu, Qing Guo, Xianglong Liu, Dongning Sun, Yang Liu
arXiv AI
Sep 17

Echo: Learning-based Matching Decompilation using Trusted Back Translation

Echo is a matching decompilation system that uses trusted back‑translation to guide an iterative search for source code whose recompiled assembly exactly matches a target binary. It generates candidate programs and compilation configurations with a domain‑specific model, compiles them, measures assembly similarity, and refines mismatches through rule‑based, neural, and reasoning‑based techniques. In evaluations on function‑level benchmarks and the Mirai malware binary, Echo achieves 2.43× more exact matches than the strongest baseline and outperforms GPT‑5.6 and Codex by 2.75× and 7.4× on Mirai, respectively.

By Jun Bi, Xiangxin Fang, Aarsh Chaube, Jos\'e Wesley De Souza Magalh\~aes, Rodrigo C. O. Rocha, Michael O'Boyle
arXiv Computation and Language
Sep 15

Toward Secure Code Generation: Bridging Correctness and Security via Task-Adaptive Vulnerability Modeling and Execution-Based Benchmarking

arXiv:2407.02395v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used for program synthesis, yet they often generate code that is functionally plausible but ins...

By Jiexin Wang, Liuwen Cao, Xitong Luo, Yang Cao, Zhenghao Li, Yunyi Xiao, Mengchen Zhao, Adam Jatowt, Yi Cai