arXiv Machine Learning

PhysicsBench: A Unified Leaderboard for Generative and Predictive Models in Engineering Design and Simulation

PhysicsBench is a unified benchmark and leaderboard that evaluates both generative and predictive AI models for engineering design and simulation under a single standardized procedure. It covers seven generation and prediction tasks across 1D, 2D, and 3D domains, ranking 66 models on nine industrial‑scale CAD/CFD/FEA datasets and public references, expanded into 28 configurations. The evaluation uses realistic, limited data scales and a common metric suite that captures geometric fidelity, physical‑field accuracy, and engineering‑specific validity, with rankings derived via a PageRank‑based dominance graph and a separate efficiency view.

arXiv Machine Learning
Aug 21

CarBench: A Comprehensive Benchmark for Neural Surrogates on High-Fidelity 3D Car Aerodynamics

arXiv:2512. 07847v2 Announce Type: replace Abstract: Benchmarking has been the cornerstone of progress in computer vision, natural language processing, and the broader deep learning domain, driving algorithmic innovation through standardized datasets and reproducible evaluation protocols.

By Mohamed Elrefaie, Dule Shu, Matt Klenk, Faez Ahmed
arXiv Machine Learning
1d ago

Scientific Discovery under Validation Congestion via Multi-Fidelity Pairwise Rankings

The paper introduces PRISMS, a framework that uses expert pairwise rankings of varying fidelity to curate scientific designs without relying on data-intensive regression models. By escalating queries from lower- to higher-fidelity rankers based on Fisher-information, PRISMS improves discovery recall and reduces the number of screening rounds compared to regression-only and non‑escalated ranking methods. In optimization tasks, PRISMS outperforms Bayesian optimization by achieving higher hypervolume.

By Kevin Tirta Wijaya, Alston Lo, Michael Sun, Wojciech Matusik, Vahid Babaei
arXiv AI
Aug 14

PhysMaster: Building an Autonomous AI Physicist for Theoretical and Computational Physics Research

arXiv:2512. 19799v2 Announce Type: replace Abstract: Advances in LLM reasoning and tool use have enabled agentic science, yet frontier theoretical and computational physics remains challenging because research requires deep domain expertise, long-horizon reasoning, and reliable numerical computation.

By Tingjia Miao, Wenkai Jin, Jinxin Tan, Muhua Zhang, Xianghe Pang, Zexi Liu, Yuwen Du, Tian Jin, Tu Guo, Zhengliang Zhang, Jingkun Liu, Yuelin Hu, Jiejun Zhang, Yunjie Huang, Yuhan Wang, Wenbo Li, Yinuo Gao, Shuo Chen, Rui Ye, Yuzhi Zhang, Linfeng Zhang, Kun Chen, Wei Wang, Weinan E, Siheng Chen
arXiv AI
Jun 2

Towards a Physics Foundation Model

arXiv:2509. 13805v4 Announce Type: replace-cross Abstract: Foundation models have revolutionized natural language processing through a ``train once, deploy anywhere'' paradigm, where a single pre-trained model adapts to countless downstream tasks without retraining.

By Florian Wiesner, Zo\"e J. Gray, Matthias Wessling, Stephen Baek
arXiv Computation and Language
Sep 4

RealCADBench: Benchmarking Parametric CAD Modeling from Industrial Design Intents

RealCADBench is a new benchmark for evaluating intent‑to‑program parametric CAD modeling, featuring 12,632 tasks drawn from 19 factory‑automation categories and covering text, 2D drawings, product photos, and rendered images for both Part and Assembly modeling. The study reports results on a 1,770‑task evaluation slice, using metrics such as executability, Solid IoU, Surface IoU, and a rubric‑based visual‑semantic identity Judge. Across nine standalone and six frontier‑scale large models, no single model dominates all four metrics, highlighting diverse strengths and failure modes like missing fine structures and incorrect assembly placement.

By JoyIndustrial VisCAD Team, Linxin Cai, Qiuhe Hong, Zhichao Huang, Guanlin Li, Zongzhen Li, Hongsen Liu, Yichen Long, Wei Wang, Yuchen Wang, Dongyue Yang, Huimu Yu, Xianwen Zhong
arXiv AI
Aug 17

Don't Claim Benchmark-Oriented Optimization Improves General Coding Capability -- Diverse Evaluation Is Required

arXiv:2608. 13566v1 Announce Type: cross Abstract: Post-training papers, model cards, and blog posts often treat scores on a small set of coding benchmarks (e.

By Egor Shibaev, Vera Kudrevskaia, Timur Galimzyanov, Mikhail Evtikhiev, Ana Terna, Rastislav Rabatin, Timur Kudashev, Timofey Bryksin, Arina Puchkova, Patrik Bartak, Egor Bogomolov, Sergey Titov
arXiv AI
4d ago

MatToolBench: Benchmarking Multimodal Agents in Real-World Materials Science Workflows

MatToolBench is a new benchmark that evaluates multimodal GUI agents on professional materials science software. It contains 204 tasks across 10 tools in three modalities—GUI operation, OriginPro scripting, and code-based database queries—executed inside a Windows 11 VM. The benchmark offers fine-grained, expert-decomposed scoring and a high-performing multimodal judge for aesthetic assessment, revealing that strong general benchmark performance does not transfer to scientific workflows.

By Mei Wu, Rui Xie, Runyu Zhang, Yuqiang Li, Tianfan Fu, Bo Chen, Kai Yu, Xin Chen, Lu Chen