Hugging Face Blog

StarCoder2-Instruct: Fully Transparent and Permissive Self-Alignment for Code Generation

arXiv AI
Sep 10

CodeTD: Topology of Attention Detects Hallucinations in Code LLMs

The paper introduces CodeTD, a novel method that uses topological data analysis of attention maps from code language models to pre‑execution assess code correctness and detect hallucinations. It quantifies prompt‑generation mismatch through topological patterns and is evaluated on multiple benchmarks (HumanEval, MBPP, BigCodeBench, MultiPL‑E) across five programming languages and ten Code LLMs up to 34B parameters. Results show CodeTD outperforms recent baselines and transfers well between coding benchmarks.

By Daria Voronkova, Ilya Trofimov, Anton Dmitriev, Eduard Tulchinskii, Evgeny Burnaev, Serguei Barannikov
arXiv AI
5d ago

Self-Spec Verifiable Code Generation

arXiv:2609.39568v1 Announce Type: cross Abstract: Large language models (LLMs) may generate unreliable code on corner cases missed by testing, while formal verification can provide machine-checkable...

By Jiaru Qian, Yihong Dong, Yongmin Li, Hao Zhu, Bin Gu, Ge Li
arXiv Computer Vision
Aug 31

Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning

arXiv:2604.06079v2 Announce Type: replace Abstract: Graphics Program Synthesis is pivotal for interpreting and editing visual data, effectively facilitating the reverse-engineering of static visuals...

By Juekai Lin, Yun Zhu, Honglin Lin, Sijing Li, Tianwei Lin, Zheng Liu, Xiaoyang Wang, Wenqiao Zhang, Lijun Wu