BigCodeBench: The Next Generation of HumanEval
Related stories
Introducing the LiveCodeBench Leaderboard - Holistic and Contamination-Free Evaluation of Code LLMs
SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios
arXiv:2509. 22097v5 Announce Type: replace-cross Abstract: Large language model-powered code agents are rapidly transforming software engineering, yet the security risks of their generated code have become a critical concern.
FormulaCode: Evaluating Agentic Optimization on Large Codebases
arXiv:2603. 16011v3 Announce Type: replace-cross Abstract: Large language model (LLM) coding agents increasingly operate at the repository level, motivating benchmarks that evaluate their ability to optimize entire codebases under realistic constraints.
SkillVetBench: LLM-as-Judge for Multi-Dimensional Security Risk Evaluation in Open-Source LLM Agent Skills
arXiv:2606. 15899v1 Announce Type: cross Abstract: Open-source LLM agent ecosystems are growing rapidly, yet the security of community-contributed skills - modular tool definitions that extend agent capabilities - remains largely unvetted.
Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution
arXiv:2608. 08311v1 Announce Type: cross Abstract: We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation improve through reviewed commits that become the runtime for later work.
VICBench: A Multi-Language Benchmark for Code Vulnerability Detection
Evaluating security vulnerability detection tools requires benchmark datasets with vulnerability-inducing commits (VICs) - the commits that first introduce vulnerabilities into codebases. VICs are essential for determining the full range of vulnerable software versions.
An Agentic Evaluation Framework for AI-Generated Scientific Code in PETSc
The paper presents PETSCAgent-Bench, a multidimensional benchmark and agent-based framework designed to evaluate AI-generated scientific code that uses the PETSc library. It combines deterministic checks with LLM-based assessments across five categories—correctness, performance, code quality, algorithmic appropriateness, and library-specific conventions—using a 14-evaluator pipeline. The framework demonstrates that while current large language models produce readable, well-structured code, they often fail on correctness and library conventions in realistic PETSc problems, revealing gaps that simple pass/fail tests miss.
Steerability via constraints: a substrate for scalable oversight of coding agents
arXiv:2607. 02389v1 Announce Type: new Abstract: Coding agents are capable; human oversight is the bottleneck.
SafeCoder vs. Closed-source Code Assistants
ToolUniverse: An open platform for democratizing AI scientists
arXiv:2509. 23426v3 Announce Type: replace Abstract: AI scientists are emerging computational systems that serve as collaborative partners in discovery.
VICBench: A Multi-Language Benchmark for Code Vulnerability Detection
arXiv:2608. 12246v1 Announce Type: cross Abstract: Evaluating security vulnerability detection tools requires benchmark datasets with vulnerability-inducing commits (VICs) - the commits that first introduce vulnerabilities into codebases.