arXiv AI By Yanuo Ma, Ben Kereopa-Yorke, Ben Schultz

Building to the Test: Coding Agents Deliver What You Check, Not What You Requested

Read the original on arXiv AI →

arXiv:2606. 28430v1 Announce Type: cross Abstract: Benchmarks are widely used to evaluate task completion by Large Language Models (LLMs), but this approach has accumulated construction-validity problems, and a passing score may not show whether the requested task was delivered.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 4

Counterexamples as Feedback for Agent Self-Correction

The paper introduces A-CEGIS, a lightweight framework that employs counterexamples as feedback to evaluate and improve multi-turn natural-language-to-regex synthesis. In experiments on 30 NL-RX-Turk tasks, counterexample feedback enables agents to solve 90% of tasks within four turns, outperforming zero‑shot generation, generic self‑correction, and error‑only feedback. A full diagnostic run with hardening solves all hidden tasks by the final turn, achieving a mean time‑to‑success of 2.7 turns and robust success of 77% after targeted probing.

By Sidhesh Badrinarayan, Adithya Parthasarathy