Training CodeParrot 🦜 from Scratch
Related stories
Text and code embeddings by contrastive pre-training
StarCoder: A State-of-the-Art LLM for Code
Creating a Coding Assistant with StarCoder
Using Local Coding Agents
Using Open-Weight Models in Local Coding Harnesses as an Alternative to Claude Code and Codex Subscriptions
Invariant Pretraining for Robust Code Representations
arXiv:2608. 15412v1 Announce Type: cross Abstract: Encoder-based code representation models remain widely deployed for discriminative tasks such as clone detection and code classification, where their small size and low inference cost are decisive.
Evaluating large language models trained on code
OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling
Context lengths of language models (LMs) have dramatically increased, driven by the demands for in-context learning, self-improvement, and long-horizon agentic workflows. Existing long-context corpora, however, are dominated by books, academic articles, and code repositories, which are finite resources and often scarce in long-distance dependencies.
The Age of Machine Learning As Code Has Arrived
Code Llama: Llama 2 learns to code
CodeAlchemy: Synthetic Code Rewriting at Scale
arXiv:2606. 10087v1 Announce Type: cross Abstract: Pre-training on raw code teaches syntax but provides sparse signal for diverse real-world task formats.
OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling
arXiv:2608. 05141v1 Announce Type: new Abstract: Context lengths of language models (LMs) have dramatically increased, driven by the demands for in-context learning, self-improvement, and long-horizon agentic workflows.
