arXiv AI

An Agentic Evaluation Framework for AI-Generated Scientific Code in PETSc

The paper presents PETSCAgent-Bench, a multidimensional benchmark and agent-based framework designed to evaluate AI-generated scientific code that uses the PETSc library. It combines deterministic checks with LLM-based assessments across five categories—correctness, performance, code quality, algorithmic appropriateness, and library-specific conventions—using a 14-evaluator pipeline. The framework demonstrates that while current large language models produce readable, well-structured code, they often fail on correctness and library conventions in realistic PETSc problems, revealing gaps that simple pass/fail tests miss.

arXiv Computation and Language
Sep 17

ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments

arXiv:2609.19134v1 Announce Type: new Abstract: Scientific code repositories encode decades of human knowledge in executable models, methods, and tools. Yet fragmented toolchains, implicit domain con...

By Hejia Geng, Zesen Huang, Haoyang Li, Wenbin Li, Koutian Wu, Zihan Zhou, Yuanbo Pang, Weihao Liu, Zigong Xu, Zhiping Li, Zongzheng Zhang, Chuanfei Dong, Jiankai Sun, Tianzhe Zheng, Fengyu Xie, Yue Ma, Yueheng Shi, Tong Xie, Zonglin Di, Xianrong Liu, Qucheng Gao, Yimin Liu, Jiaming Pan, Sheng Huang, Xiao-Han Ma, Lanqing Yuan, Zhenlin Zhu, Ziang Liu, Ziyang Xu, Junkai Wang, Kangkai Liang, Jiayi Xian, Zehong Zhao, Liuwei Xu, Jingxu Xie, Peijin Zhang, Qiang Gao, Chengyi Xing, Zhe Zhao, Xi Wang, Yaopeng Xing, Xing Meng, Zhenfei Yin, Yingcheng Wu, Ling Yang
arXiv AI
Aug 28

Exploring the Role of LLMs in HPC Programming: A Survey

The survey reviews how Large Language Models (LLMs) are being used in High‑Performance Computing (HPC) programming, covering code generation, parallelization, frameworks, evaluation, and broader challenges. It finds that general‑purpose LLMs perform adequately on serial and OpenMP‑style tasks but struggle with distributed MPI workloads, while domain‑specialized models achieve higher accuracy yet are limited in scope and evaluation. The authors argue that LLMs will not replace HPC experts soon but can act as powerful collaborators, provided richer datasets, integration with performance tools, rigorous evaluation, and governance are developed.

By Strahinja Ljaljevic, Josep Jorba, Sergio Iserte
arXiv AI
Jun 9

A case study of evaluating AI agents on a neuroscience data-to-discovery pipeline

arXiv:2606. 07718v1 Announce Type: new Abstract: Agentic AI tools offer a promising path to automating software development bottlenecks in scientific research pipelines, particularly for stages that take domain experts days to months to build, where scientists care about correctness and robustness, not implementation details.

By Kai A. Horstmann, Ethan Lin, Alice A. Robie, Jennifer J. Sun, Kristin Branson
arXiv AI
Sep 7

Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation

Harbor Adapters is a unified evaluation infrastructure that ports over 80 agentic benchmarks, enabling arbitrary agents to be tested across complex environments. The authors performed a large‑scale evaluation of 8 models on 54 benchmarks, using Terminus‑2 and three native harnesses, revealing detailed agent capabilities and failure modes. They also created Harbor‑Index, a curated set of 82 challenging tasks from 29 benchmarks, designed to be affordable yet comprehensive, with the best model achieving a 28.0% pass rate.

By Lin Shi (Audrey), Haowei Lin (Audrey), Zixuan Zhu (Audrey), Xiaoyue Zhou (Audrey), Xiang Li (Audrey), Xiangning Lin (Audrey), Yaxuan Deng (Audrey), Han Xu (Audrey), Yuangang Li (Audrey), Shanda Li (Audrey), Zizhao Chen (Audrey), Hanwen Xing (Audrey), Harsh Raj (Audrey), Bo Chen (Audrey), Quan Shi (Audrey), Steven Dillmann (Audrey), Yipeng Gao (Audrey), Puneesh Khanna (Audrey), Ruofan Lu (Audrey), Chao Beyond Zhou (Audrey), Michael Yang (Audrey), Robert Zhang (Audrey), Siyuan Chai (Audrey), Jiayu Chang (Audrey), Yizhao Chen (Audrey), Xiaokun Chen (Audrey), Yiwei Dai (Audrey), Wenting Yang (Audrey), Hange Liu (Audrey), Minghao Liu (Audrey), Zihan Wang (Audrey), Adnan El Assadi (Audrey), Benedikt Stroebl (Audrey), E. Kelly Buchanan (Audrey), Han Meng (Audrey), Junwei He (Audrey), Longxuan Yu (Audrey), Radin Shayanfar (Audrey), Yukyung Lee (Audrey), Zhikang Dong (Audrey), Allen G Hart (Audrey), Anjiang Wei (Audrey), Anurag Kashyap (Audrey), Arpandeep Khatua (Audrey), Audrey Jixin Zheng (Audrey), Chengrui Ma (Audrey), David Heineman (Audrey), Dubing Chen (Audrey), Hai-Anh Trinh (Audrey), Haishuo Fang (Audrey), Hefan Zhang (Audrey), Hui Shen (Audrey), Issa Sugiura (Audrey), Jiankai Sun (Audrey), Jiechao Gao (Audrey), Junhong Lin (Audrey), Junnan Li (Audrey), Kai Yang (Audrey), Lei Hsiung (Audrey), Maoyu Wang (Audrey), Mengze Tang (Audrey), Nabil Omi (Audrey), Negin Raoof (Audrey), Nicholas Edwards (Audrey), Octavia Guo (Audrey), Orfeas Menis Mastromichalakis (Audrey), Pengliang Ji (Audrey), Przemys{\l}aw Hejman (Audrey), Qi Qi (Audrey), Qunshu Lin (Audrey), Richard Zhuang (Audrey), Rui Yang (Audrey), Ruichen Zheng (Audrey), Ryan Marten (Audrey), Shaghayegh Fazliani (Audrey), Shizheng Hou (Audrey), Sicong Jiang (Audrey), Sijie Li (Audrey), Song Bian (Audrey), Terry Yue Zhuo (Audrey), Tianqing Wu (Audrey), Tom Tang (Audrey), Wanjia Zhao (Audrey), Weihao Xuan (Audrey), Wenhua Liang (Audrey), Xian Liu (Audrey), Xin Lan (Audrey), Xuan Zhang (Audrey), Xuandong Zhao (Audrey), Yanchuan Tang (Audrey), Yifan Jiang (Audrey), Yijiang Li (Audrey), Yitong Guan (Audrey), Yizhi Li (Audrey), Yonghui Liu (Audrey), Yuheng Tang (Audrey), Yujun (Audrey), Mao, Yunfei Zhao, Yuxin Wang, Yuxuan Tang, Zhenheng Tang, Zhifei Li, Ziruo Wang, Ziyu She, Kaiyuan Liu, Iheb Chaabane, Yuxin Tang, Xiangyi Li, Andy Konwinski, Boxuan Li, Leon Liangyu Chen, Alex Dimakis, Nicholas Carlini, Soroush Vosoughi, Di He, Etash Guha, Benjamin Feuer, Mike Merrill, Ludwig Schmidt, Alex Shaw
arXiv AI
Aug 28

AgentFold: Closed-Loop Agentic Search for Protein Folding Model Design

AgentFold is a multi‑agent framework that treats protein‑folding model design as a closed‑loop search over executable code variants. Starting from the ESMFold codebase, the agents generate hypotheses, modify and debug code, evaluate model variants, and store both successes and failures in structured memory, guided by an MCTS‑style policy that allocates GPU resources. In an engineering‑scale experiment, AgentFold explored about 80 variants using 5,000 GPU‑hours and 170 million LLM tokens, improving the best lDDT score by 7.5% over independent Codex proposals and outperforming a random‑search baseline, while also uncovering empirical design patterns such as the benefits of early, soft, learnable priors.

By Mingquan Liu, Jiangyu Chen, Hanqun Cao, Xujun Zhang, Pengsen Ma, Xiangru Tang, Shuting Jin, Zhuo Yang, Tianfan Fu, Fang Wu, Xiangxiang Zeng
arXiv Computation and Language
Sep 25

An Empirical Study of Automating Agent Evaluation

The paper presents EvalAgent, an AI assistant that automates agent evaluation by encoding domain expertise into evaluation skills such as procedural instructions, reusable code, and dynamic API retrieval. EvalAgent constructs a trace-based pipeline that outputs metrics, executable code, and reports, and is evaluated using a new meta-evaluation framework and AgentEvalBench. Results show that EvalAgent improves the Eval@1 metric from 17.5% to 65% and receives 79.5% human expert preference, while ablation studies confirm the importance of evaluation skills.

By Kang Zhou, Sangmin Woo, Haibo Ding, Kiran Ramnath, Subramanian Chidambaram, Aosong Feng, Vinayak Arannil, Muhyun Kim, Ishan Singh, Darren Wang, Zhichao Xu, Megha Gandhi, Nirmal Prabhu, Soumya Smruti Mishra, Smeet Dhakecha, Vivek Singh, Gouri Pandeshwar, Lin Lee Cheong