← Back to all news
Hugging Face Blog October 7, 2025

BigCodeArena: Judging code generations end to end with code executions

Read the original on Hugging Face Blog →

The Flow has not summarised this story yet — read it at Hugging Face Blog.

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Jun 16

Towards Functional Correctness of Large Code Models with Selective Generation

arXiv:2505. 13553v3 Announce Type: replace-cross Abstract: The hallucination of code generation models hinders their applicability to systems requiring higher safety standards.

By Jaewoo Jeong, Taesoo Kim, Sangdon Park
safety
More like this →
arXiv AI
Aug 6

ExeCRE: Execution-Consistency Guided Reliability Estimation for Self-Correcting Code Generation

arXiv:2608. 04439v1 Announce Type: cross Abstract: Large language models (LLMs) have made notable progress in code generation, but they still struggle on challenging tasks that require sophisticated algorithms or complex implementations.

By Yiru Dong, Richong Zhang, Fanshuang Kong, Si Chen
llms
More like this →
arXiv AI
Jun 2

Measuring and Mitigating Bias in Code Generated by Large Language Models

arXiv:2606. 00049v1 Announce Type: cross Abstract: Large language models (LLMs) are widely recognised for their applications in natural language generation and are increasingly used for code generation tasks.

By Yuxi Chen, Yutian Tang, Timothy Storer
llmsagentssafety
More like this →
arXiv AI
Jun 3

FLARE: Fine-Grained Diagnostic Feedback for LLM Code Refinement

arXiv:2606. 03852v1 Announce Type: cross Abstract: Large language models often generate code with bugs.

By Yinsheng Yao, Hongxiang Zhang, Weixi Tong, Tianyi Zhang
llms
More like this →
OpenAI Blog
Mar 3, 2022

A research agenda for assessing the economic impacts of code generation models

More like this →
arXiv AI
Jun 9

Beyond Pass Rate: A Multilingual, Execution-Grounded Evaluation of Open Code LLMs

arXiv:2606. 08840v1 Announce Type: new Abstract: Code generation models are typically compared using compact execution benchmarks and aggregate pass rates, but such summaries obscure how performance varies across programming languages, problem families, and failure modes.

By Sayed Erfan Arefin
llmsbenchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.0.0 · bb4ee0e