arXiv AI

The Case for Model Science: Verify, Explore, Steer, Refine

arXiv:2606. 01189v1 Announce Type: new Abstract: We argue that the AI community is now ready to move beyond benchmarking and consolidate scattered efforts in model analysis into a systematic discipline, a direction we term Model Science.

arXiv AI
Jul 14

FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights

arXiv:2602. 02905v2 Announce Type: replace Abstract: Autonomous agents powered by large language models (LLMs) promise to accelerate scientific discovery end-to-end, but rigorously evaluating their capacity for verifiable discovery remains a central challenge.

By Zhen Wang, Fan Bai, Zhongyan Luo, Jinyan Su, Kaiser Sun, Xinle Yu, Jieyuan Liu, Kun Zhou, Claire Cardie, Mark Dredze, Zhiting Hu, Eric P. Xing
arXiv AI
Sep 25

AI in Science: Early Insights

The paper "AI in Science: Early Insights" analyzes AI’s impact on scientific work using data from 15 million Gemini interactions, 2,600 specialized AI models, and a survey of 600 scientists. It finds widespread AI adoption, complementary use of large language models and specialized tools, significant productivity gains of about seven hours per week, and a shift in research bottlenecks toward hypothesis backlog and verification needs. The study suggests AI can boost scientific productivity but its full effect depends on addressing new downstream challenges.

By Mihai Codreanu, Alex Imas, Juan Mateos-Garcia, Joseph Emmens, Evalyne Muiruri, Arthur Turrell, Julian Jacobs, Atoosa Kasirzadeh, Ana Trisovic, Yiyuan Chen, Tanya Rodchenko, Catherine Pollard, Scott Strand, Daniel Rock, Zanna Iscenko, Fabien Curto Millet, Neil Thompson, James Manyika
arXiv AI
Sep 12

Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents

Sci‑MMR is a new benchmark for multi‑step evidence‑grounded scientific reasoning in multimodal agents, featuring 235 multi‑hop tasks across four disciplines and an average of nine figure panels per task. It evaluates not just final answer accuracy but also the recovery of structured evidence from scientific claims, citations, visual data, and supporting regions. Experiments on eight state‑of‑the‑art models show a gap of over 20 points between answer accuracy and complete evidence recovery, highlighting significant challenges in evidence acquisition and integration.

By Jiaqiang Li, Yajie Yang, Zhiheng Xi, Jiadong Chen, Enyu Zhou, Senjie Jin, Yang Nan, Jiazheng Zhang, Han Wang, Yanxin Li, Dingwei Zhu, Bicheng Deng, Yuhui Wang, Xiang Zheng, Qi Zhang, Lei Bai, Xingjun Ma, Tao Gui
arXiv Machine Learning
Sep 2

Modelpedia: A Catalog of Model Findings for the Meta-Science of AI

Modelpedia is an automated, LLM-assisted framework that extracts and organizes findings about AI models from published papers into a searchable public catalog. It links each finding to the relevant model, dataset, method, and concept, and has already extracted over a thousand findings from ICLR 2024 and 2025 papers. The authors invite the community to explore, contribute to, and build on this open catalog, positioning model findings as a shared foundation for the meta‑science of AI.

By Franciszek Bernat (Centre for Credible AI, Warsaw University of Technology), Dawid P{\l}udowski (Centre for Credible AI, Warsaw University of Technology), Micha{\l} Jan W{\l}odarczyk (Centre for Credible AI, Warsaw University of Technology), Luca Longo (University College Cork), Jianlong Zhou (University of Technology Sydney), Andreas Holzinger (Human-Centered AI Lab), Riccardo Guidotti (University of Pisa, ISTI-CNR), Wojciech Samek (Technical University of Berlin, Berlin Institute for the Foundations of Learning and Data), Przemys{\l}aw Biecek (Centre for Credible AI, University of Warsaw)
arXiv AI
Aug 10

Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery

arXiv:2608. 06931v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly involved in scientific discovery, yet it remains unclear whether they can support complex real laboratory science.

By Taolin Han, Yuchen Zhang, Jinghang Wang, Yun Wu, Wai Yuet Chiu, Zhaohai Li, Yifei Zhang, Jinxin Wang, Yuhao Zhou, Chen Zhao, Jiajia Li, Jiaxin Li, Qile Jin, Kewei Sun, Shuang Wu, Weiqi Zhai, Renquan Lv, Junchao Li, Ruodan Chen, Qingteng Chen, Zhibo Yang, Hu Wei, Lin Qu, Shuai Bai, Bing Zhao
arXiv Computer Vision
Sep 7

From Interpretability Methods to Interpretable Models

The paper argues that explainable AI for computer vision has focused too much on developing interpretability methods rather than assessing how interpretable the models themselves are. It proposes a shift toward model-centric evaluation, using existing tools to compare what different models represent and compute, and emphasizes the need to measure whether humans can truly understand these models. The authors review the current toolbox, survey limited model comparison work, draw parallels to systems neuroscience, and outline a future agenda for model-focused XAI.

By Julien Colin, Nuria Oliver, Thomas Serre
arXiv AI
Aug 13

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

arXiv:2608. 12036v1 Announce Type: new Abstract: AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood.

By Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang, Shumin Deng, Linyi Yang, Zhixiang Cui, Xin Xu, Yunzhi Yao, Buqiang Xu, Fei Shen, Haozhe Luo, Yunxiang Wei, Ningyu Zhang, Julian McAuley, Tat Seng Chua, Huajun Chen
arXiv AI
Sep 18

ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI

The paper introduces ScientistTwo, a fully autonomous multi‑agent framework that takes a scientific problem, establishes baselines, generates hypotheses, and coordinates specialized agents to conduct an end‑to‑end discovery cycle without human intervention. It rigorously tests and refines its methods through automated experiments, ablation studies, and a closed‑loop peer‑review engine. Benchmarking against top conferences (ICLR, ICML, NeurIPS) shows that ScientistTwo produces expert‑level, publishable papers and codebases that outperform human state‑of‑the‑art models and receive higher review ratings under automated AI review.

By Jaehyun Nam, Jinsung Yoon, Yanzhou Pan, Yubo Wang, Rui Meng, Parthasarathy Ranganathan, Tomas Pfister
arXiv AI
Jun 9

A case study of evaluating AI agents on a neuroscience data-to-discovery pipeline

arXiv:2606. 07718v1 Announce Type: new Abstract: Agentic AI tools offer a promising path to automating software development bottlenecks in scientific research pipelines, particularly for stages that take domain experts days to months to build, where scientists care about correctness and robustness, not implementation details.

By Kai A. Horstmann, Ethan Lin, Alice A. Robie, Jennifer J. Sun, Kristin Branson
arXiv AI
Jun 8

Act As a Real Researcher: A Suite of Benchmarks Evaluating Frontier LLMs and Agentic Harnesses in Research Lifecycle

arXiv:2606. 07462v1 Announce Type: new Abstract: As foundation models advance and agent scaffolding becomes increasingly sophisticated, agents have demonstrated remarkable proficiency in complex, long-horizon coding tasks and even autonomous experiment execution.

By Jiayu Wang, Weijiang Lv, Bowen Fu, Jing Fu, Jiayi Song, Lingyu Zhang, Lanxuan Xue, Luodi Chen, Zepeng Xin, Kaiyu Li, Xiangyong Cao