arXiv AI

Building to the Test: Coding Agents Deliver What You Check, Not What You Requested

arXiv:2606. 28430v1 Announce Type: cross Abstract: Benchmarks are widely used to evaluate task completion by Large Language Models (LLMs), but this approach has accumulated construction-validity problems, and a passing score may not show whether the requested task was delivered.

arXiv AI
Sep 4

Counterexamples as Feedback for Agent Self-Correction

The paper introduces A-CEGIS, a lightweight framework that employs counterexamples as feedback to evaluate and improve multi-turn natural-language-to-regex synthesis. In experiments on 30 NL-RX-Turk tasks, counterexample feedback enables agents to solve 90% of tasks within four turns, outperforming zero‑shot generation, generic self‑correction, and error‑only feedback. A full diagnostic run with hardening solves all hidden tasks by the final turn, achieving a mean time‑to‑success of 2.7 turns and robust success of 77% after targeted probing.

By Sidhesh Badrinarayan, Adithya Parthasarathy
arXiv AI
Aug 28

BekchiAI: Measuring, Observing, and Controlling LLM Agents in One Click

BekchiAI introduces a benchmark and platform for evaluating large language model agents. The benchmark comprises 13 tool‑using ReAct agents across seven task categories, totaling 2,057 deterministic test tasks with verifier‑checkable gold answers. The platform offers web‑based observability, token and latency telemetry, and remote run termination for live agents.

By Mesut Toruk