The paper investigates whether code large language models (CodeLLMs) inadvertently reproduce proprietary or sensitive code by evaluating seven state‑of‑the‑art training data detection (TDD) methods on eight CodeLLMs. It introduces CodeSnitch, a benchmark of 9,000 function‑level code samples across three languages, each labeled as included or excluded from training data, and applies mutation strategies based on the Type‑1 to Type‑4 code clone taxonomy to test TDD robustness. The study offers a systematic assessment of current TDD techniques for code and suggests directions for developing more effective detection methods.
By Tianlin Li, Yunxiang Wei, Zhiming Li, Aishan Liu, Qing Guo, Xianglong Liu, Dongning Sun, Yang Liu
SpecMine is a large-scale corpus that documents Spec-Driven Development (SDD) artifacts in public GitHub repositories. It includes a broad census of 470,795 spec files from 73,030 repositories linked to 17 tools, a focused census of 98,574 Kiro layout files from 12,910 repositories, and a sweep of 5,992 pull requests across 581 repositories that modify specs. The dataset provides enriched metadata, full commit histories, parsed document structures, and over 2.4 million typed references connecting specs to code, sibling documents, PRs, branches, and issues.
By Shyam Agarwal, Anmol Singhal, Travis Breaux, Bogdan Vasilescu
arXiv:2605. 13138v2 Announce Type: replace-cross Abstract: Automated detection of vulnerability-fixing commits (\vfcs) is critical for timely security patch deployment, as advisory databases lag patch releases by a median of 25 days and many fixes never receive advisories.
By Nils Loose, Joseph Bienh\"uls, Kristoffer Hempel, Felix M\"achtle, Thomas Eisenbarth
arXiv:2607. 01867v1 Announce Type: cross Abstract: The use of LLMs in software development has become increasingly widespread on tasks such as code generation and summarization.
By Yongyi Ji, Jiaji Wang, Yi Zhou, Fuxiang Chen, Hongji Yang
arXiv:2507. 11059v3 Announce Type: replace-cross Abstract: The rapid advancement of Large Language Models (LLMs) in software engineering has revealed critical limitations in existing benchmarks, particularly the widely used SWE-bench dataset.
By Pavel Adamenko, Mikhail Ivanov, Aidar Valeev, Rodion Levichev, Pavel Zadorozhny, Ivan Lopatin, Dmitry Babaev, Alena Fenogenova, Valentin Malykh
arXiv:2606. 17283v1 Announce Type: cross Abstract: Achieving reproducibility, quantity, and diversity in vulnerability datasets has long been viewed as an inherent three-way trade-off, where improving one dimension often comes at the cost of the others.
By Xiang Mei, Jordi Del Castillo, Pulkit Singh Singaria, Haoran Xi, Abdelouahab Benchikh, Tiffany Bao, Ruoyu Wang, Yan Shoshitaishvili, Adam Doup\'e, Hammond Pearce, Brendan Dolan-Gavitt
The use of LLMs in software development has become increasingly widespread on tasks such as code generation and summarization. Reports from large technology companies showed that around 20% to 30% of their code are generated by LLMs.
arXiv:2506. 11066v3 Announce Type: replace-cross Abstract: Code retrieval is essential in modern software development, as it boosts code reuse and accelerates debugging.
By Jiahui Geng, Fengyu Cai, Shaobo Cui, Qing Li, Liangwei Chen, Chenyang Lyu, Haonan Li, Derui Zhu, Walter Pretschner, Heinz Koeppl, Fakhri Karray
The paper introduces ADFD‑Migrate, a method that extracts a latent declarative representation of code—an annotated data‑flow diagram (ADFD)—to aid large‑scale repository migration. By using an LLM to infer the source ADFD from repository context and guiding target‑language generation with dependency‑aware chunking, the approach improves porting soundness and completeness. Evaluated on 50 Fortran repositories, the system achieves high behavioral agreement and a superior migration outcome index compared to baseline translation methods.
By Shraddha Surana, Ashwin Srinivasan, Michael Bain
arXiv:2608. 06640v1 Announce Type: cross Abstract: The widespread integration of AI coding assistants offers undeniable boosts to engineering velocity.
By Michael Tran, Fred Lewis, Kun Yang, Saksham Thakur, Aditya Kini, Aditya Patil, Milad Hashemi, Parthasarathy Ranganathan
arXiv:2606. 12620v1 Announce Type: cross Abstract: Thanks to the rapid adoption of AI code assistants powered by large language models (LLMs), industry codebases are, increasingly, a hybrid of AI- and human-authored code.
By Luke Patterson, Li Wang, Adam Faulkner
arXiv:2608. 14742v1 Announce Type: cross Abstract: Pandas has emerged as the de facto library for data processing and machine learning, widely used for tasks, such as data loading, transformation, and analysis.
By Syrym Abdikhan, Mazhar Hameed