Hugging Face Blog

StarCoder2 and The Stack v2

arXiv AI
Sep 10

CodeTD: Topology of Attention Detects Hallucinations in Code LLMs

The paper introduces CodeTD, a novel method that uses topological data analysis of attention maps from code language models to pre‑execution assess code correctness and detect hallucinations. It quantifies prompt‑generation mismatch through topological patterns and is evaluated on multiple benchmarks (HumanEval, MBPP, BigCodeBench, MultiPL‑E) across five programming languages and ten Code LLMs up to 34B parameters. Results show CodeTD outperforms recent baselines and transfers well between coding benchmarks.

By Daria Voronkova, Ilya Trofimov, Anton Dmitriev, Eduard Tulchinskii, Evgeny Burnaev, Serguei Barannikov