arXiv AI By Hong Zhang, Barry Smith, Satish Balay, Le Chen, Murat Keceli, Lois Curfman McInnes, Junchao Zhang

An Agentic Evaluation Framework for AI-Generated Scientific Code in PETSc

Read the original on arXiv AI →

The paper presents PETSCAgent-Bench, a multidimensional benchmark and agent-based framework designed to evaluate AI-generated scientific code that uses the PETSc library. It combines deterministic checks with LLM-based assessments across five categories—correctness, performance, code quality, algorithmic appropriateness, and library-specific conventions—using a 14-evaluator pipeline. The framework demonstrates that while current large language models produce readable, well-structured code, they often fail on correctness and library conventions in realistic PETSc problems, revealing gaps that simple pass/fail tests miss.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 17

ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments

arXiv:2609.19134v1 Announce Type: new Abstract: Scientific code repositories encode decades of human knowledge in executable models, methods, and tools. Yet fragmented toolchains, implicit domain con...

By Hejia Geng, Zesen Huang, Haoyang Li, Wenbin Li, Koutian Wu, Zihan Zhou, Yuanbo Pang, Weihao Liu, Zigong Xu, Zhiping Li, Zongzheng Zhang, Chuanfei Dong, Jiankai Sun, Tianzhe Zheng, Fengyu Xie, Yue Ma, Yueheng Shi, Tong Xie, Zonglin Di, Xianrong Liu, Qucheng Gao, Yimin Liu, Jiaming Pan, Sheng Huang, Xiao-Han Ma, Lanqing Yuan, Zhenlin Zhu, Ziang Liu, Ziyang Xu, Junkai Wang, Kangkai Liang, Jiayi Xian, Zehong Zhao, Liuwei Xu, Jingxu Xie, Peijin Zhang, Qiang Gao, Chengyi Xing, Zhe Zhao, Xi Wang, Yaopeng Xing, Xing Meng, Zhenfei Yin, Yingcheng Wu, Ling Yang
arXiv AI
Aug 28

Exploring the Role of LLMs in HPC Programming: A Survey

The survey reviews how Large Language Models (LLMs) are being used in High‑Performance Computing (HPC) programming, covering code generation, parallelization, frameworks, evaluation, and broader challenges. It finds that general‑purpose LLMs perform adequately on serial and OpenMP‑style tasks but struggle with distributed MPI workloads, while domain‑specialized models achieve higher accuracy yet are limited in scope and evaluation. The authors argue that LLMs will not replace HPC experts soon but can act as powerful collaborators, provided richer datasets, integration with performance tools, rigorous evaluation, and governance are developed.

By Strahinja Ljaljevic, Josep Jorba, Sergio Iserte
arXiv AI
Jun 9

A case study of evaluating AI agents on a neuroscience data-to-discovery pipeline

arXiv:2606. 07718v1 Announce Type: new Abstract: Agentic AI tools offer a promising path to automating software development bottlenecks in scientific research pipelines, particularly for stages that take domain experts days to months to build, where scientists care about correctness and robustness, not implementation details.

By Kai A. Horstmann, Ethan Lin, Alice A. Robie, Jennifer J. Sun, Kristin Branson