arXiv AI

To See is Not to Master: Teaching LLMs to Use Private Libraries for Code Generation

The paper introduces PriCoder, a method for teaching large language models (LLMs) to effectively use private library APIs for code generation. PriCoder synthesizes training data by constructing a graph and applying two operators—Progressive Graph Evolution to increase diversity and Multidimensional Graph Pruning to enhance quality. Experiments on three mainstream LLMs demonstrate that PriCoder boosts private‑library code generation by over 20% in pass@1, while leaving general code generation largely unchanged.

arXiv AI
2d ago

Are AI Coders Snitches? An Empirical Study of Pretraining Data Detection on Code Large Language Models

The paper investigates whether code large language models (CodeLLMs) inadvertently reproduce proprietary or sensitive code by evaluating seven state‑of‑the‑art training data detection (TDD) methods on eight CodeLLMs. It introduces CodeSnitch, a benchmark of 9,000 function‑level code samples across three languages, each labeled as included or excluded from training data, and applies mutation strategies based on the Type‑1 to Type‑4 code clone taxonomy to test TDD robustness. The study offers a systematic assessment of current TDD techniques for code and suggests directions for developing more effective detection methods.

By Tianlin Li, Yunxiang Wei, Zhiming Li, Aishan Liu, Qing Guo, Xianglong Liu, Dongning Sun, Yang Liu