arXiv AI

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language

arXiv:2607. 09789v1 Announce Type: new Abstract: We introduce PHITSBench, an execution-scored benchmark for the Monte Carlo Particle and Heavy Ion Transport code System (PHITS).

arXiv AI
Aug 26

Beyond Executable Models: The Pufibara Agent Harness and the Modelica Agent Workflow Benchmark for Physical System Modeling

The paper introduces Pufibara, an agent harness designed to maintain engineering state and evidence across revisions in Modelica-based physical system modeling. It also presents a 232-task Modelica Agent Workflow Benchmark covering model repair, generation, and tuning, evaluated by an external benchmark-owned evaluator. Experiments show Pufibara outperforms Claude Code in task success and resource efficiency across two LLM backends.

By Zizhe Wang
arXiv AI
Aug 14

PhysMaster: Building an Autonomous AI Physicist for Theoretical and Computational Physics Research

arXiv:2512. 19799v2 Announce Type: replace Abstract: Advances in LLM reasoning and tool use have enabled agentic science, yet frontier theoretical and computational physics remains challenging because research requires deep domain expertise, long-horizon reasoning, and reliable numerical computation.

By Tingjia Miao, Wenkai Jin, Jinxin Tan, Muhua Zhang, Xianghe Pang, Zexi Liu, Yuwen Du, Tian Jin, Tu Guo, Zhengliang Zhang, Jingkun Liu, Yuelin Hu, Jiejun Zhang, Yunjie Huang, Yuhan Wang, Wenbo Li, Yinuo Gao, Shuo Chen, Rui Ye, Yuzhi Zhang, Linfeng Zhang, Kun Chen, Wei Wang, Weinan E, Siheng Chen
arXiv AI
Aug 11

SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs

arXiv:2603. 20253v3 Announce Type: replace-cross Abstract: Evaluating LLM agents for scientific tasks has focused on token costs while ignoring tool-use costs like simulation time and experimental resources.

By Yadi Cao, Sicheng Lai, Jiahe Huang, Yang Zhang, Zach Lawrence, Rohan Bhakta, Izzy F. Thomas, Mingyun Cao, Chung-Hao Tsai, Zihao Zhou, Yidong Zhao, Hao Liu, Alessandro Marinoni, Alexey Arefiev, Rose Yu
arXiv AI
6d ago

Agentick: A Unified Benchmark for General Sequential Decision-Making Agents

Agentick is a unified benchmark for sequential decision‑making agents that evaluates RL, LLM, VLM, hybrid, and human agents on 37 procedurally generated tasks across six capability categories, four difficulty levels, and five observation modalities via a single Gymnasium‑compatible interface. It includes a Coding API, oracle reference policies, pre‑built SFT datasets, a composable agent harness, and a live leaderboard. An evaluation of 27 configurations and over 90,000 episodes shows no single approach dominates, with GPT‑5 mini leading overall, PPO excelling in planning and multi‑agent tasks, and the reasoning harness boosting LLM performance by 3‑10×, while ASCII observations outperform natural language.

By Roger Creus Castanyer, Pablo Samuel Castro, Glen Berseth
arXiv AI
Sep 25

MOOSEnger: A Simulation-Aware AI Agent Framework for the MOOSE Ecosystem

MOOSEnger is a simulation‑aware AI agent framework designed for the MOOSE ecosystem, integrating an interchangeable reasoning model with domain knowledge, revised simulation artifacts, MOOSE‑specific validation, and executable solver feedback. Its generate‑check‑repair‑run workflow uses MOOSE knowledge retrieval, HIT‑aware parsing, syntax metadata, diagnostics, and revision‑controlled authoring to bind evidence to each input revision and guide bounded repair before acceptance. Across 200 prompts, MOOSEnger raises executable success from 5% to 89.5% with GPT‑5.2 and from 0% to 76.5% with Gemma 4 31B, and a ten‑case benchmark shows all generated inputs meet semantic alignment, with eight also meeting numerical‑accuracy criteria.

By Mengnan Li, Jason Miller, Zaid Abulawi, Zachary Prince, Matt Kohl, Jack M. Cavaluzzi, Guillaume Giudicelli, Casey T. Icenhour, Alexander Lindsay, Cody Permann
arXiv AI
Jul 1

HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents

arXiv:2606. 31179v1 Announce Type: new Abstract: As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward real-world healthcare applications.

By Qianchu Liu, Sheng Zhang, Guanghui Qin, Jeya Maria Jose Valanarasu, Maximilian Rokuss, Mingyu Lu, Timothy Ossowski, Juan Manuel Zambrano Chaves, Cliff Wong, Peniel Argaw, Yashna Hasija, Mu Wei, Wen-wai Yim, Qin Liu, Zilin Jing, Jason Entenmann, Naoto Usuyama, Tristan Naumann, Hoifung Poon
arXiv AI
Jun 6

AutoDFT: A Closed-Loop Multi-Agent Framework for Autonomous DFT Calculations

arXiv:2605. 26179v2 Announce Type: replace-cross Abstract: Density functional theory (DFT) serves as the basis for computational discovery in materials science and chemistry, yet each calculation demands extensive human effort: adjusting algorithms when convergence stalls, revising plans when unexpected physics emerges, and inserting steps as intermediate results reshape the problem.

By Penghui Yang, Zhonghan Zhang, Yue Li, Xinrun Wang, Yanchen Deng, Yuhao Lu, Bijun Tang, Zheng Liu, Bo An