arXiv AI

Learn from Your Mistakes: Tree-like Self-Play for Secure Code LLMs

arXiv:2606. 03489v1 Announce Type: cross Abstract: While Large Language Models (LLMs) excel in code generation, they remain prone to replicating subtle yet critical vulnerabilities endemic to their training data.

Hugging Face Trending Papers
Jun 2

Learn from Your Mistakes: Tree-like Self-Play for Secure Code LLMs

While Large Language Models (LLMs) excel in code generation, they remain prone to replicating subtle yet critical vulnerabilities endemic to their training data. Current alignment techniques, such as Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL), typically apply coarse-grained optimization at the sequence level.

arXiv Computation and Language
Sep 15

Toward Secure Code Generation: Bridging Correctness and Security via Task-Adaptive Vulnerability Modeling and Execution-Based Benchmarking

arXiv:2407.02395v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used for program synthesis, yet they often generate code that is functionally plausible but ins...

By Jiexin Wang, Liuwen Cao, Xitong Luo, Yang Cao, Zhenghao Li, Yunyi Xiao, Mengchen Zhao, Adam Jatowt, Yi Cai
arXiv AI
Aug 25

Evaluating Inference-Time Defenses Against Package Hallucination in LLM-Generated Code

The paper investigates how large language models (LLMs) hallucinate nonexistent software packages during code generation and evaluates methods to mitigate this issue. It finds that current evaluation practices overestimate hallucination rates, especially for Python, and that Retrieval-Augmented Generation (RAG) and Self-Refine reduce hallucinations across multiple models and languages. The study also introduces Package Utility (PU) to measure whether defenses preserve useful recommendations and shows that Greedy decoding offers the best trade‑off between mitigation and utility, while adversarial prompts significantly increase hallucination rates, particularly in Ruby.

By Alberick Euraste Djire, Iyiola E. Olatunji, Melissa Tessa, Earl T. Barr, Jacques Klein, Tegawend\'e F. Bissyand\'e
arXiv Computation and Language
3d ago

SecureVibe: Making Vibe Coding More Secure

SecureVibe is a training recipe designed to enhance the security of vibe coding by explicitly targeting planning and testing for code security. It combines supervised fine‑tuning on a security suite with post‑training methods (SECUREVIBE_rl and SECUREVIBE_hg) that use verifiable execution feedback and hint‑based self‑supervision. The approach outperforms baselines on multiple security coding benchmarks, improving security pass@1 by up to 6.9 points on BaxBench and 11.5 points on SusVibes, while also boosting functionality pass@1 on both security and generic coding tasks.

By Danqing Wang, Baolin Peng, Zhepei Wei, Isadora White, Wenlin Yao, Hao Cheng, Qianhui Wu, Minseon Kim, Xingdi Yuan, Lei Li, Jianfeng Gao
arXiv AI
Jun 16

DualGauge: Automated Joint Security-Functionality Benchmarking of Specification-Only Code Generation by LLMs and Coding Agents

arXiv:2511. 20709v2 Announce Type: replace-cross Abstract: Large language models (LLMs) and LLM-based coding agents are now used to generate code from natural-language specifications, yet ensuring such code is both functionally correct and secure remains a challenge.

By Rupam Patir, Keyan Guo, Suvadra Barua, Abhijeet Pathak, Dinesh Gudimetla, Jiawei Guo, Hongxin Hu, Haipeng Cai