arXiv AI By Sijia Gu, Noor Nashid, Ali Mesbah

SCATE: Learning to Supervise Coding Agents for Cost-Effective Test Generation

Read the original on arXiv AI →

arXiv:2607. 08983v1 Announce Type: cross Abstract: While autonomous coding agents have significantly advanced automated test generation, they remain fundamentally limited by lazy generation, a phenomenon where agents prematurely terminate tasks and systematically avoid complex programmatic logic, resulting in inadequate code coverage.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Aug 17

TDD-Agent: Test-Driven Reasoning for Code Generation

Large Language Models (LLMs) have achieved remarkable progress in code generation, yet ensuring correctness in complex, repository-level tasks remains challenging. Existing approaches often use generated tests as static post-hoc validators, which limits their ability to guide implementation and may introduce misleading feedback when the tests themselves are incomplete or incorrect.

arXiv AI
Jun 4

Can Generalist Agents Automate Data Curation?

arXiv:2606. 04261v1 Announce Type: new Abstract: Curating training data is among the most consequential yet labor-intensive parts of modern AI development: practitioners iteratively propose, implement, evaluate, and revise data policies against noisy benchmark feedback.

By Feiyang Kang, Hanze Li, Adam Nguyen, Mahavir Dabas, Jiaqi W. Ma, Frederic Sala, Dawn Song, Ruoxi Jia
arXiv Computation and Language
Aug 27

Scalable Supervision for Software Agents via Patch Reasoning

The paper introduces R4P, a reasoning‑based supervision method for software agents that eliminates the need for test execution by using a group‑wise training objective to verify multiple patches simultaneously. R4P achieves 72.2% accuracy on the SWE‑bench patch verification task, matching proprietary models, and enables the creation of an execution‑free scaffold called Mini‑SE. Mini‑SE, trained purely with reinforcement learning via R4P, improves Pass@1 from 26.2% to 32.8% over the baseline Qwen3‑32B, demonstrating R4P’s practical utility and scalable performance.

By Junjielong Xu, Boyin Tan, Xiaoyuan Liu, Chao Peng, Pengfei Gao, Pinjia He