Towards Data Science

Bug Detection Blind Spots in AI Coding Harnesses (GStack and Beyond)

The article "Bug Detection Blind Spots in AI Coding Harnesses (GStack and Beyond)" reports on 28 debugging experiments that show AI coding tools struggle more with missing information than with code complexity. It highlights that these tools exhibit blind spots when key details are absent, affecting their debugging performance.

arXiv Machine Learning
Aug 27

FuzzingBrain-Bench V1: Evaluating Open-Ended Bug Discovery by LLMs

FuzzingBrain‑Bench V1 is a new benchmark that tests large language models (LLMs) on their ability to discover software bugs in open‑source projects. Unlike prior benchmarks that focus on a single target vulnerability, this benchmark gives models a Docker‑based harness and asks them to generate inputs that trigger as many distinct crashes as possible. The first version contains 77 challenges from 43 projects (36 C, 32 C++, 9 Java/JVM) and evaluates Claude Haiku 4.5, Claude Sonnet 4.6, and Claude Opus 4.8, with Claude Opus 4.8 achieving the highest score by triggering crashes in 60 of 77 challenges.

By Ze Sheng, Aleksandar Kezic, Zhicheng Chen, Jeff Huang
Hugging Face Trending Papers
Sep 24

Large Language Models for Programming: Actually Fixing or Reimplementing Incorrect Code?

The paper investigates how Large Language Models (LLMs) handle bug fixing compared to human-written patches by analyzing about 3,000 Codeforces submissions. It finds that LLMs often modify more lines than necessary and sometimes produce entirely new solutions, and that they solve more problems correctly when generating solutions from scratch rather than patching existing code. The study highlights implications for AI‑assisted programming tools, suggesting a shift toward incremental problem‑solving strategies.

arXiv Computation and Language
Sep 25

Large Language Models for Programming: Actually Fixing or Reimplementing Incorrect Code?

The paper investigates how large language models (LLMs) handle bug fixing versus problem solving in competitive programming. Using a dataset of ~3,000 Codeforces submissions and their human fixes, the authors compare LLM-generated patches to human patches and assess whether LLMs prefer to modify buggy code or generate new solutions. Results show that LLMs often alter more lines than necessary and sometimes produce entirely new solutions, performing better when allowed to generate solutions from scratch rather than patching existing code.

By Alexandru Stefan Stoica, Traian Rebedea, Marian Cristian Mihaescu
arXiv AI
4d ago

AgentBug-Smith: Automatically Reproducing Real-World Harness Bugs in Agentic Systems

arXiv:2609.37864v1 Announce Type: cross Abstract: Agent harness bugs exhibit unique characteristics and remain challenging for state-of-the-art software agents to repair. Progress in this area is fur...

By Yiming Cheng (The University of Chicago), Alfin Wijaya Rahardja (Fudan University), Mengshi Zhang (TensorBlock, Inc), Zihao Chen (TensorBlock, Inc), Zhenpeng Chen (Tsinghua University), Yiling Lou (University of Illinois Urbana-Champaign)