StarCoder2-Instruct: Fully Transparent and Permissive Self-Alignment for Code Generation
Related stories
Creating a Coding Assistant with StarCoder
Exploring the Potential of Diffusion Large Language Models in Code Generation
arXiv:2509.11252v3 Announce Type: replace-cross Abstract: LLMs have become the mainstream approaches to code generation. Existing LLMs mainly employ autoregressive generation, i.e. generating code to...
Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation
arXiv:2606. 28998v1 Announce Type: cross Abstract: Large Language Model (LLM) alignment trains an LLM using preference data to produce outputs that better meet established quality standards.
CodeTD: Topology of Attention Detects Hallucinations in Code LLMs
The paper introduces CodeTD, a novel method that uses topological data analysis of attention maps from code language models to pre‑execution assess code correctness and detect hallucinations. It quantifies prompt‑generation mismatch through topological patterns and is evaluated on multiple benchmarks (HumanEval, MBPP, BigCodeBench, MultiPL‑E) across five programming languages and ten Code LLMs up to 34B parameters. Results show CodeTD outperforms recent baselines and transfers well between coding benchmarks.
CodeEvo: Interaction-Driven Synthesis of Code-centric Data through Hybrid and Iterative Feedback
arXiv:2507. 22080v2 Announce Type: replace-cross Abstract: Acquiring high-quality instruction-code pairs is essential for training Large Language Models for code generation.
Towards Functional Correctness of Large Code Models with Selective Generation
arXiv:2505. 13553v3 Announce Type: replace-cross Abstract: The hallucination of code generation models hinders their applicability to systems requiring higher safety standards.
Instruction Alignment for Binary Code Representation Learning
arXiv:2608. 11766v1 Announce Type: cross Abstract: Binary code representation learning is a fundamental problem in software security and reverse engineering.
Self-Spec Verifiable Code Generation
arXiv:2609.39568v1 Announce Type: cross Abstract: Large language models (LLMs) may generate unreliable code on corner cases missed by testing, while formal verification can provide machine-checkable...
StarCoder2 and The Stack v2
ExeCRE: Execution-Consistency Guided Reliability Estimation for Self-Correcting Code Generation
arXiv:2608. 04439v1 Announce Type: cross Abstract: Large language models (LLMs) have made notable progress in code generation, but they still struggle on challenging tasks that require sophisticated algorithms or complex implementations.
Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning
arXiv:2604.06079v2 Announce Type: replace Abstract: Graphics Program Synthesis is pivotal for interpreting and editing visual data, effectively facilitating the reverse-engineering of static visuals...