Kepler is an open‑source harness that represents hypotheses as executable world models and validates them through retrospective transition checks and conditional prediction checks. In the ARC‑AGI‑3 benchmark, a frozen Claude Opus 5 configuration achieved a perfect 100.00 RHAE on all 25 public games without per‑game model selection or score‑conditioned reruns, and matched or outperformed median‑human action counts on 181 of 183 levels. The study also identified three evaluation failures and highlighted that public‑set score alone has limited discriminative value, advocating for first‑attempt, cost‑conditioned, and verification‑aware reporting.
whyItMatters":"The results demonstrate that a purely score‑based evaluation can be misleading, underscoring the need for more rigorous, cost‑aware, and verification‑aware metrics in AI benchmark assessments."
By Wensen Wu
arXiv:2607. 15439v1 Announce Type: new Abstract: Our previous ARC-AGI-3 agent bundled executable world modeling, scheduled simplification, and exact replay verification, leaving unclear which idea accounted for its performance.
By Sergey Rodionov
The paper introduces Gauntlet, a framework that lets large language models autonomously build game-playing agents from a bare contract—just a game description, raw observation/action interface, and an empty policy file. In a single session, the model experiments with the game, compiles a standalone controller, and the resulting program is evaluated on held‑out instances without further model calls. The authors demonstrate that these compiled agents can win full‑scale games such as StarCraft II and Civilization, marking the first time a language‑agent system has achieved standalone victory in such complex titles.
By Joey Xiao, Haonan Huang
The paper introduces VHD-Play, a pipeline that first samples and solves a mathematical model before generating agentic reinforcement learning environments, ensuring that dynamics and evaluation are aligned from the outset. This approach yields 3,300 diverse environments at a low cost and significantly improves the performance of a large language‑model agent (Qwen3.6‑35B‑A3B) across multiple diagnostic families and external benchmarks. The study demonstrates that stateful interaction is a key factor in learning gains and that scaling the training substrate can further enhance performance.
By Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
arXiv:2608.22533v1 Announce Type: new
Abstract: Autonomous computer-use agents are increasingly applied to long-horizon tasks requiring coordinated application calls, persistent state tracking, and v...
By Zheyuan Deng, Binghang Lu, Hanqi Feng, Shirley Huang, Dianzhuo Wang, Yuanda Xu, Zhiwei Zhang, Yige Sun, Changhong Mou, Runyu Zhang, Yuexing Hao, Barnabas Poczos, Xiaomin Li
arXiv:2609.27532v1 Announce Type: new
Abstract: Long-horizon agentic tasks require an agent to modify an environment through a sequence of tool calls, with success determined by the final state. The...
By Ming Ma, Yi Zhu, Yiran Zhong, Feida Zhu, Chonghan Liu, Pengkun Jiao, Qichao Wang, Yanhao Jia, Tianming Yang, Steven Hoi