arXiv AI

Practical Source Code Recovery from Binary Functions Using Anchor-Based Retrieval and LLM Reasoning

arXiv:2607. 09452v1 Announce Type: cross Abstract: We present a practical pipeline for recovering source code from stripped binary functions by combining reverse engineering, anchor-based source code retrieval, and large language model reasoning.

arXiv AI
Sep 17

Echo: Learning-based Matching Decompilation using Trusted Back Translation

Echo is a matching decompilation system that uses trusted back‑translation to guide an iterative search for source code whose recompiled assembly exactly matches a target binary. It generates candidate programs and compilation configurations with a domain‑specific model, compiles them, measures assembly similarity, and refines mismatches through rule‑based, neural, and reasoning‑based techniques. In evaluations on function‑level benchmarks and the Mirai malware binary, Echo achieves 2.43× more exact matches than the strongest baseline and outperforms GPT‑5.6 and Codex by 2.75× and 7.4× on Mirai, respectively.

By Jun Bi, Xiangxin Fang, Aarsh Chaube, Jos\'e Wesley De Souza Magalh\~aes, Rodrigo C. O. Rocha, Michael O'Boyle
arXiv AI
Jun 9

Efficient and Scalable Provenance Tracking for LLM-Generated Code Snippets

arXiv:2605. 28510v2 Announce Type: replace-cross Abstract: Large language models (LLMs) for code completion and generation are increasingly used in software development, yet they may reproduce training examples verbatim and without authorship attribution, raising legal and ethical concerns around plagiarism and license compliance.

By Andrea Gurioli, Davide D'Ascenzo, Federico Pennino, Maurizio Gabbrielli, Stefano Zacchiroli
arXiv AI
Jul 3

BLAgent: Agentic RAG for File-Level Bug Localization

arXiv:2605. 17965v2 Announce Type: replace-cross Abstract: Bug localization remains a key bottleneck for large language model (LLM)-based software maintenance, where accurately identifying faulty code is essential for debugging, root cause analysis, triage, and automated program repair (APR).

By Md Afif Al Mamun, Gias Uddin
arXiv AI
Aug 10

Beyond Text Matching: Towards Reference-Free Evaluation for Human-Oriented Binary Reverse Engineering

arXiv:2608. 07038v1 Announce Type: cross Abstract: Human-Oriented Binary Reverse Engineering (HOBRE) aims to transform decompiled pseudocode into a more human-friendly representation, thereby reducing the cognitive burden of reverse analysis and improving efficiency.

By Xiuwei Shang, Li Hu, Xiao Jiang, Jieke Shi, Junda He, Zhou Yang, Shaoyin Cheng, Guoqiang Chen, Weiming Zhang, David Lo