arXiv AI By Nicolas Koller, Andreas u. Schmidt

REFORGE: A Method for Benchmarking LLMs' Reverse Engineering Capabilities in Decompiled Binary Function Naming

Read the original on arXiv AI →

arXiv:2607. 07738v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly applied to reverse-engineering tasks, and recent threat-intelligence reporting shows them operating inside live offensive-security workflows.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 15

Introspective Uncertainty Estimation for LLM-Based Code Generation

The thesis explores Introspective Uncertainty Estimation (IUE) for large language models (LLMs) in code generation, aiming to determine whether hidden-state representations can indicate functional correctness at both response and line levels. Using LiveCodeBench and BigCodeBench, the study finds that hidden states provide a strong signal for overall correctness, with static single-token probes performing best, while dynamic strategies offer no consistent advantage. Although line-level fault localization is more challenging, a conditional Top‑K ranking approach remains effective, suggesting a two‑stage workflow that first screens responses for risk and then prioritizes line‑level checks.

By Thomas Klassert
arXiv AI
Sep 17

Echo: Learning-based Matching Decompilation using Trusted Back Translation

Echo is a matching decompilation system that uses trusted back‑translation to guide an iterative search for source code whose recompiled assembly exactly matches a target binary. It generates candidate programs and compilation configurations with a domain‑specific model, compiles them, measures assembly similarity, and refines mismatches through rule‑based, neural, and reasoning‑based techniques. In evaluations on function‑level benchmarks and the Mirai malware binary, Echo achieves 2.43× more exact matches than the strongest baseline and outperforms GPT‑5.6 and Codex by 2.75× and 7.4× on Mirai, respectively.

By Jun Bi, Xiangxin Fang, Aarsh Chaube, Jos\'e Wesley De Souza Magalh\~aes, Rodrigo C. O. Rocha, Michael O'Boyle
arXiv AI
Aug 10

Beyond Text Matching: Towards Reference-Free Evaluation for Human-Oriented Binary Reverse Engineering

arXiv:2608. 07038v1 Announce Type: cross Abstract: Human-Oriented Binary Reverse Engineering (HOBRE) aims to transform decompiled pseudocode into a more human-friendly representation, thereby reducing the cognitive burden of reverse analysis and improving efficiency.

By Xiuwei Shang, Li Hu, Xiao Jiang, Jieke Shi, Junda He, Zhou Yang, Shaoyin Cheng, Guoqiang Chen, Weiming Zhang, David Lo