arXiv Machine Learning By Masahiro Kato

GTDD: Generative Test-Driven Development for AI Coding Agents with Adversarial Testing

Read the original on arXiv Machine Learning →

The paper introduces Generative Test-Driven Development (GTDD), a method where a separate testing agent continuously generates new inputs after each coding agent’s implementation, guided by a human-specified behavioral contract and prior feedback. A trusted evaluator validates these inputs, provides reduced counterexamples, and stores them for regression testing, ensuring that development repeatedly confronts failures beyond the initial examples. Experimental results on a stateful key-value store show that policies regenerating tests during development achieve lower mean failure rates than a single-generation policy, while providing the tester with the candidate’s source code does not yield additional improvement.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 17

A Study of the Reliability of Agentic AI-Generated Programs

The paper investigates the reliability of software produced by agentic AI by comparing AI-generated versions of ten well-known Linux utilities to their human-written counterparts. Using fuzz testing (both black-box and coverage-guided AFL++), the authors find that AI-generated code is often as reliable or more reliable than the latest human versions, with fewer memory errors but a higher incidence of hangs. The study emphasizes that robust AI-generated software requires careful prompting, skilled human oversight, and that the AI workflow can serve as a cost-effective specification for sustainable code.

By Ayesha Shafique, Barton P. MIller, Elisa R. Heymann
Hugging Face Trending Papers
Aug 17

TDD-Agent: Test-Driven Reasoning for Code Generation

Large Language Models (LLMs) have achieved remarkable progress in code generation, yet ensuring correctness in complex, repository-level tasks remains challenging. Existing approaches often use generated tests as static post-hoc validators, which limits their ability to guide implementation and may introduce misleading feedback when the tests themselves are incomplete or incorrect.

arXiv Machine Learning
Jul 10

TTHE: Test-Time Harness Evolution

arXiv:2607. 08124v1 Announce Type: cross Abstract: The behavior of an LLM agent is determined not only by the underlying model, but also by its harness: the executable program that constructs context, invokes tools, verifies intermediate results, and recovers from failures.

By Jun Nie, Yonggang Zhang, Jun Song, Qianshu Cai, Dahai Yu, Yike Guo, Xinmei Tian, Bo Han