arXiv AI

DiG-bench: Discovery in Games

arXiv:2608. 12593v1 Announce Type: new Abstract: Discovery---formulating novel generalizations---is a central part of the scientific process.

arXiv AI
Sep 25

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

ExplorationBench is a new benchmark designed to evaluate AI systems’ ability to conduct scientific exploration in verifiable alien worlds. It comprises two sandbox environments—AlienCode and AlienLogic—each containing discovery targets, tasks, flawed manuals, and tool‑call schemas that force systems to formulate hypotheses, design experiments, and iterate on results. Ten AI systems were tested, revealing that while the best performers can learn and apply unfamiliar rules, their progress varies across exploration trajectories and can even regress with continued exploration.

By Ming Zhang, Zhenghao Xiang, Peizhong Gao, Yujiong Shen, Yuhui Wang, Zhonghan Yue, Shihan Dou, Zhangyue Yin, Junjie Ye, Shichun Liu, Weihuang Zheng, Jiahao Chen, Jiayi Chen, Hongzhang Liu, Jiaqi Shao, Tao Gui, Qi Zhang, Xuanjing Huang, Suncong Zheng, Maxm Pan
Hugging Face Trending Papers
Sep 24

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

ExplorationBench is a benchmark designed to evaluate AI systems’ ability to conduct scientific exploration in verifiable alien worlds, where rules are executable and can be precisely checked. It consists of two sandboxes—AlienCode and AlienLogic—each offering discovery targets, tasks, flawed manuals, environmental feedback, and tool‑call schemas. The benchmark tests whether systems can generate new hypotheses, design experiments, and iterate on results, rather than merely recalling pre‑trained knowledge, and finds that top performers can acquire and apply unfamiliar rules, though performance varies across exploration trajectories.

arXiv AI
Sep 23

Synthesizing Reactive Character Behaviors for Continuous Games via Programmatic Policy Search

The paper introduces a method for creating reactive character behaviors in continuous games as compact, human‑readable programs. It searches over a domain‑specific language that uses reactive geometric decisions and higher‑order constructs to discretize continuous behavior space, while eliminating redundant program forms through synthesis antipatterns. The approach, called agentic sketching, combines bottom‑up symbolic enumeration with top‑down guidance from a coding agent, and outperforms either technique alone on a benchmark of 14 continuous games.

By Maxim Gumin, Hsueh-Ti Derek Liu, Victor Zordan, Daniel Ritchie
arXiv AI
Sep 21

GameLogicBench: Evaluating Coding Agents on Runtime Game Logic with Tick-Level State Assertions

arXiv:2609.21562v1 Announce Type: cross Abstract: Coding agents can modify and test code across large software projects. Game development is a domain where agents must implement gameplay rules. A gam...

By Xinyu Che, Yunfei Ge, Shihao Li, Yanchen Liu, Hang Yan, Xinping Lei, Yanghai Wang, Zixuan Dong, Yifan Yao, Qianqian Xie, Letian Zhu, Jiaheng Liu
arXiv AI
Jul 14

FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights

arXiv:2602. 02905v2 Announce Type: replace Abstract: Autonomous agents powered by large language models (LLMs) promise to accelerate scientific discovery end-to-end, but rigorously evaluating their capacity for verifiable discovery remains a central challenge.

By Zhen Wang, Fan Bai, Zhongyan Luo, Jinyan Su, Kaiser Sun, Xinle Yu, Jieyuan Liu, Kun Zhou, Claire Cardie, Mark Dredze, Zhiting Hu, Eric P. Xing
arXiv AI
4d ago

CheatBench: Measuring Reward Gaming in AI Agents

arXiv:2609.36308v1 Announce Type: new Abstract: Reinforcement learning has helped AI agents solve increasingly difficult tasks, but high rewards do not always reflect the work users intended. In rece...

By Long Phan, Stephen K. Yang, Jason J. Lim, Mantas Mazeika, Wenyu Zhang, Zheyuan Liu, Richard Ren, Jingxiang Meng, Yaoteng Tan, Weiliang Zhao, Addison Wu, Matei Anghel, Dan Hendrycks