StarCoder2-Instruct: Fully Transparent and Permissive Self-Alignment for Code Generation
Read the original on Hugging Face Blog →The Flow has not summarised this story yet — read it at Hugging Face Blog.
The Flow has not summarised this story yet — read it at Hugging Face Blog.
arXiv:2509.11252v3 Announce Type: replace-cross Abstract: LLMs have become the mainstream approaches to code generation. Existing LLMs mainly employ autoregressive generation, i.e. generating code to...
arXiv:2606. 28998v1 Announce Type: cross Abstract: Large Language Model (LLM) alignment trains an LLM using preference data to produce outputs that better meet established quality standards.
The paper introduces CodeTD, a novel method that uses topological data analysis of attention maps from code language models to pre‑execution assess code correctness and detect hallucinations. It quantifies prompt‑generation mismatch through topological patterns and is evaluated on multiple benchmarks (HumanEval, MBPP, BigCodeBench, MultiPL‑E) across five programming languages and ten Code LLMs up to 34B parameters. Results show CodeTD outperforms recent baselines and transfers well between coding benchmarks.
arXiv:2507. 22080v2 Announce Type: replace-cross Abstract: Acquiring high-quality instruction-code pairs is essential for training Large Language Models for code generation.