arXiv:2606. 09957v1 Announce Type: cross Abstract: Semantic faults specific to the use of machine learning models are a common problem for machine learning developers, causing suboptimal predictions, high computational cost, or incorrect outputs.
By Willem Meijer, Kristian Sandahl, D\'aniel Varr\'o
arXiv:2607. 10411v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used for code smell detection tasks due to their ability to interpret program semantics.
By Istiaq Ahmed Fahad, Kamruzzaman Asif, Md. Nurul Ahad Tawhid
The paper introduces SWE-Flux, a repository‑level benchmark designed to test large language models’ ability to reason about runtime behavior. It contains 480 execution‑grounded instances from 12 real Python repositories, with gold answers automatically harvested from instrumented test executions. Evaluation of five LLMs shows the task remains difficult, with the best model achieving only 37% accuracy, and the benchmark can generate challenging variants through input perturbation.
By Hamed Taherkhani, Mohammad Abdollahi, Melika Sepidband, Hridya Dhulipala, Tien N. Nguyen, Hadi Hemmati
The paper introduces CodeScan, a black-box, vulnerability-oriented scanning framework designed to detect data poisoning and backdoor attacks in code generation large language models (LLMs). CodeScan operates by analyzing structural similarities across multiple code generations, normalizing them with abstract syntax tree (AST) techniques, and then applying LLM-based vulnerability analysis to identify recurring insecure patterns. Evaluations on 117 models across three architectures and multiple sizes show over 97% detection accuracy with fewer false positives compared to prior methods.
By Shenao Yan, Shan Jin, Shimaa Ahmed, Sunpreet Singh Arora, Yiwei Cai, Yizhen Wang, Yuan Hong
arXiv:2506. 02791v4 Announce Type: replace-cross Abstract: In recent years, code intelligence has gained increasing importance in the field of automated software engineering.
By Zhen Yang, Hongyi Lin, Yifan He, Junqi Wang, Zeyu Sun, Shuo Liu, Jie Xu, Pengpeng Wang, Zhongxing Yu, Qingyuan Liang
arXiv:2606. 19149v1 Announce Type: cross Abstract: Automated vulnerability discovery in large codebases remains challenging: traditional static analysis produces high false-positive rates, while dynamic approaches such as fuzzing require substantial infrastructure and often target narrow classes of bugs.
By Nahum Korda, Gadi Evron
arXiv:2509. 14335v2 Announce Type: replace-cross Abstract: Automated malware classifiers achieve strong detection performance, but auditing requires more than flagging a sample: analysts must explain malicious behaviors and justify them with code evidence.
By Xinran Zheng, Xingzhi Qian, Yiling He, Shuo Yang, Lorenzo Cavallaro
arXiv:2606. 17283v1 Announce Type: cross Abstract: Achieving reproducibility, quantity, and diversity in vulnerability datasets has long been viewed as an inherent three-way trade-off, where improving one dimension often comes at the cost of the others.
By Xiang Mei, Jordi Del Castillo, Pulkit Singh Singaria, Haoran Xi, Abdelouahab Benchikh, Tiffany Bao, Ruoyu Wang, Yan Shoshitaishvili, Adam Doup\'e, Hammond Pearce, Brendan Dolan-Gavitt
arXiv:2512. 22827v2 Announce Type: replace-cross Abstract: Code often suffers from performance bugs.
By Yue Wu, Minghao Han, Ruiyin Li, Peng Liang, Amjed Tahir, Zengyang Li, Qiong Feng, Mojtaba Shahin
arXiv:2507. 22580v2 Announce Type: replace-cross Abstract: Automated Program Repair (APR) seeks to automatically correct software bugs without requiring human intervention.
By Marcos Fuster-Pena, David de-Fitero-Dominguez, Antonio Garcia-Cabot, Eva Garcia-Lopez
arXiv:2606. 16244v1 Announce Type: cross Abstract: Large language models routinely generate code with exploitable security flaws.
By Xiaoyun Xu, Lichao Wu, Jona te Lintelo, Siyu Zhang, Stjepan Picek
The paper surveys 87 influential studies on machine‑learning‑based automated vulnerability detection (ML4AVD), categorizing them by problem formulation, input and detection granularity, target languages, evaluation metrics, datasets, and detection approaches. It identifies twelve self‑reinforcing pain points—such as overreliance on binary classification of C/C++ function‑level vulnerabilities, limited language coverage, and intertwined datasets, baselines, and metrics—that trap the field in a narrow, artificial problem space. The authors propose concrete recommendations to break these feedback loops and evaluate a recent high‑profile effort, AIxCC, against these guidelines, reflecting on ML4AVD’s relevance amid the rise of agentic AI.
By Dan Ristea, Shae McFadden, Ezzeldin Shereen, Madeleine Dwyer, Sanyam Vyas, Chris Hicks, Vasilios Mavroudis