arXiv AI By Sergey Rodionov

Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?

Read the original on arXiv AI →

arXiv:2607. 15439v1 Announce Type: new Abstract: Our previous ARC-AGI-3 agent bundled executable world modeling, scheduled simplification, and exact replay verification, leaving unclear which idea accounted for its performance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
2d ago

Kepler: Auditable World Models for ARC-AGI-3

Kepler is an open‑source harness that represents hypotheses as executable world models and validates them through retrospective transition checks and conditional prediction checks. In the ARC‑AGI‑3 benchmark, a frozen Claude Opus 5 configuration achieved a perfect 100.00 RHAE on all 25 public games without per‑game model selection or score‑conditioned reruns, and matched or outperformed median‑human action counts on 181 of 183 levels. The study also identified three evaluation failures and highlighted that public‑set score alone has limited discriminative value, advocating for first‑attempt, cost‑conditioned, and verification‑aware reporting. whyItMatters":"The results demonstrate that a purely score‑based evaluation can be misleading, underscoring the need for more rigorous, cost‑aware, and verification‑aware metrics in AI benchmark assessments."

By Wensen Wu
arXiv AI
Sep 21

GameLogicBench: Evaluating Coding Agents on Runtime Game Logic with Tick-Level State Assertions

arXiv:2609.21562v1 Announce Type: cross Abstract: Coding agents can modify and test code across large software projects. Game development is a domain where agents must implement gameplay rules. A gam...

By Xinyu Che, Yunfei Ge, Shihao Li, Yanchen Liu, Hang Yan, Xinping Lei, Yanghai Wang, Zixuan Dong, Yifan Yao, Qianqian Xie, Letian Zhu, Jiaheng Liu
arXiv AI
Sep 25

Policy as Code: A Coroutine-Bridge Harness for Fast-Reasoning Reliability on CAR-bench

The paper introduces a coroutine-bridge harness that lets a language model emit a Python program to manage tool calls in the CAR-bench evaluation. By decoupling model invocations from tool round-trips, the approach reduces model calls to a median of two per task while maintaining seven agent turns, achieving a median latency of 1.8 s on a Cerebras gpt‑oss‑120b. The harness achieved 60.0 % Pass³ on the official hidden evaluation, outperforming the baseline by 4.5× and matching frontier-model agents on GPT‑5.5, all while keeping the prompt largely cached and minimizing input compute.

By Ivan Matveev
arXiv AI
2d ago

Agents Are Systems, Not Models: Rethinking Agentic Evaluation

The paper argues that evaluating agents as fixed models is insufficient, proposing instead to treat them as configurable systems. Using a new benchmark of four scientific tasks, the authors analyze how five configuration aspects—task information, reasoning, self‑verification, time budget, and backbone model—affect performance, noting that about 54% of outcome variance arises from run‑to‑run differences even with the same settings. The study finds that providing more task information has the strongest impact, while interactions among settings (e.g., extra time only helps with adequate information or model capability) and the choice of verification tools significantly shape agent behavior.

By Luis Wiedmann, Leander Girrbach, Cordelia Schmid, Zeynep Akata
arXiv AI
Sep 17

Compiled Agency: Frontier General-Purpose Coding Agents Build Winning Game Players from Bare Interaction - from Flappy Bird to StarCraft II and Civilization

The paper introduces Gauntlet, a framework that lets large language models autonomously build game-playing agents from a bare contract—just a game description, raw observation/action interface, and an empty policy file. In a single session, the model experiments with the game, compiles a standalone controller, and the resulting program is evaluated on held‑out instances without further model calls. The authors demonstrate that these compiled agents can win full‑scale games such as StarCraft II and Civilization, marking the first time a language‑agent system has achieved standalone victory in such complex titles.

By Joey Xiao, Haonan Huang