arXiv Machine Learning By Samuel Valenzuela, Johannes Kinder

Pretraining on Call Graphs: When Binary Analysis Tasks Profit From Context

Read the original on arXiv Machine Learning →

arXiv:2608. 02084v1 Announce Type: cross Abstract: Binary function embedding models are trained to encode the semantics of binary code in such a way that they can be generalized to a variety of reverse engineering tasks, such as binary code search, vulnerability detection, or malware classification.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 19

LSem2Vec: A Simple yet Effective Two-Stage Approach for Source Code Embedding

The paper introduces LSem2Vec, a two‑stage method that first uses a large language model to extract source code semantics and then applies a sentence embedding model to produce vector representations. This approach removes the need for task‑specific training or fine‑tuning, addressing errors in LLM outputs. Experiments on three datasets across multiple programming languages show that LSem2Vec outperforms five state‑of‑the‑art unsupervised methods.

By Zixiang Xian, Chenhui Cui, Rubing Huang, Chunrong Fang, Zhenyu Chen
arXiv Machine Learning
Sep 18

Evaluating Out-of-Distribution Robustness in Graph-Based Android Malware Classification: A New Principled Benchmark

The paper introduces a new benchmark for assessing out-of-distribution robustness in graph-based Android malware classifiers, highlighting that current models drop up to 45% accuracy on unseen malware variants. It presents two scenarios—MalNet-Tiny-Common for covariate shift and MalNet-Tiny-Distinct for domain shift—and identifies a limitation in existing benchmarks that rely solely on structure-only function call graphs. To address this, the authors propose a semantic enrichment framework that augments graph topology with function-level attributes and LLM-based code embeddings, demonstrating that this data-centric approach improves robustness under distribution shift and complements model-based methods.

By Ngoc N. Tran, Anwar Said, Waseem Abbas, Tyler Derr, Xenofon D. Koutsoukos