arXiv Machine Learning
Sep 14

Performance, Efficiency and Collapse -- Advantages and Challenges in Offline Post-training of Code LLMs

The paper investigates whether reinforcement‑learning post‑training of code‑generating large language models can be done entirely offline using existing datasets, avoiding costly online code generation and GPU‑CPU communication. Experiments show that a few hours of offline RL can substantially boost zero‑shot code generation performance across models from 0.5 B to 7 B parameters, though the magnitude of improvement differs by model family.

By Abhinav Anand, Sanjana Reddy Pachika, Shweta Verma, Mira Mezini