arXiv AI By Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig

EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Jun 12

Benchmarking AI Agents for Addressing Scientific Challenges Across Scales

arXiv:2606. 12736v1 Announce Type: new Abstract: AI agents are increasingly being developed to accelerate scientific discovery, yet their practical capabilities in real research settings remain poorly understood.

By Tianyu Liu, Allen Xin Wang, Antonia Panescu, Lisa Xinyi Chen, Wenxin Long, Xinyu Wei, Yueqian Jing, Ziyao Zeng, Jihang Chen, Sihan Jiang, Ziqing Wang, Siyi Gu, Siyu Chen, Xinyang Hu, Haoran Shao, Leqi Xu, Wangjie Zheng, Zhiyuan Cao, Ada Fang, Botao Yu, Kunyang Sun, Rex Ying, Arman Cohan, Qingyu Chen, Lingzhou Xue, Kaize Ding, Yuanqi Du, Wengong Jin, Zhuoran Yang, Marinka Zitnik, James Zou, Hua Xu, Hongyu Zhao
arXiv AI
Sep 18

ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI

The paper introduces ScientistTwo, a fully autonomous multi‑agent framework that takes a scientific problem, establishes baselines, generates hypotheses, and coordinates specialized agents to conduct an end‑to‑end discovery cycle without human intervention. It rigorously tests and refines its methods through automated experiments, ablation studies, and a closed‑loop peer‑review engine. Benchmarking against top conferences (ICLR, ICML, NeurIPS) shows that ScientistTwo produces expert‑level, publishable papers and codebases that outperform human state‑of‑the‑art models and receive higher review ratings under automated AI review.

By Jaehyun Nam, Jinsung Yoon, Yanzhou Pan, Yubo Wang, Rui Meng, Parthasarathy Ranganathan, Tomas Pfister
arXiv AI
Sep 25

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

ExplorationBench is a new benchmark designed to evaluate AI systems’ ability to conduct scientific exploration in verifiable alien worlds. It comprises two sandbox environments—AlienCode and AlienLogic—each containing discovery targets, tasks, flawed manuals, and tool‑call schemas that force systems to formulate hypotheses, design experiments, and iterate on results. Ten AI systems were tested, revealing that while the best performers can learn and apply unfamiliar rules, their progress varies across exploration trajectories and can even regress with continued exploration.

By Ming Zhang, Zhenghao Xiang, Peizhong Gao, Yujiong Shen, Yuhui Wang, Zhonghan Yue, Shihan Dou, Zhangyue Yin, Junjie Ye, Shichun Liu, Weihuang Zheng, Jiahao Chen, Jiayi Chen, Hongzhang Liu, Jiaqi Shao, Tao Gui, Qi Zhang, Xuanjing Huang, Suncong Zheng, Maxm Pan
arXiv AI
Jul 14

FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights

arXiv:2602. 02905v2 Announce Type: replace Abstract: Autonomous agents powered by large language models (LLMs) promise to accelerate scientific discovery end-to-end, but rigorously evaluating their capacity for verifiable discovery remains a central challenge.

By Zhen Wang, Fan Bai, Zhongyan Luo, Jinyan Su, Kaiser Sun, Xinle Yu, Jieyuan Liu, Kun Zhou, Claire Cardie, Mark Dredze, Zhiting Hu, Eric P. Xing
Hugging Face Trending Papers
Sep 24

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

ExplorationBench is a benchmark designed to evaluate AI systems’ ability to conduct scientific exploration in verifiable alien worlds, where rules are executable and can be precisely checked. It consists of two sandboxes—AlienCode and AlienLogic—each offering discovery targets, tasks, flawed manuals, environmental feedback, and tool‑call schemas. The benchmark tests whether systems can generate new hypotheses, design experiments, and iterate on results, rather than merely recalling pre‑trained knowledge, and finds that top performers can acquire and apply unfamiliar rules, though performance varies across exploration trajectories.