arXiv Machine Learning By Sang Won Lee, Hyogu Jeong, Namwoo Kang

PhysicsBench: A Unified Leaderboard for Generative and Predictive Models in Engineering Design and Simulation

Read the original on arXiv Machine Learning →

PhysicsBench is a unified benchmark and leaderboard that evaluates both generative and predictive AI models for engineering design and simulation under a single standardized procedure. It covers seven generation and prediction tasks across 1D, 2D, and 3D domains, ranking 66 models on nine industrial‑scale CAD/CFD/FEA datasets and public references, expanded into 28 configurations. The evaluation uses realistic, limited data scales and a common metric suite that captures geometric fidelity, physical‑field accuracy, and engineering‑specific validity, with rankings derived via a PageRank‑based dominance graph and a separate efficiency view.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 21

CarBench: A Comprehensive Benchmark for Neural Surrogates on High-Fidelity 3D Car Aerodynamics

arXiv:2512. 07847v2 Announce Type: replace Abstract: Benchmarking has been the cornerstone of progress in computer vision, natural language processing, and the broader deep learning domain, driving algorithmic innovation through standardized datasets and reproducible evaluation protocols.

By Mohamed Elrefaie, Dule Shu, Matt Klenk, Faez Ahmed
arXiv Machine Learning
1d ago

Scientific Discovery under Validation Congestion via Multi-Fidelity Pairwise Rankings

The paper introduces PRISMS, a framework that uses expert pairwise rankings of varying fidelity to curate scientific designs without relying on data-intensive regression models. By escalating queries from lower- to higher-fidelity rankers based on Fisher-information, PRISMS improves discovery recall and reduces the number of screening rounds compared to regression-only and non‑escalated ranking methods. In optimization tasks, PRISMS outperforms Bayesian optimization by achieving higher hypervolume.

By Kevin Tirta Wijaya, Alston Lo, Michael Sun, Wojciech Matusik, Vahid Babaei
arXiv AI
Aug 14

PhysMaster: Building an Autonomous AI Physicist for Theoretical and Computational Physics Research

arXiv:2512. 19799v2 Announce Type: replace Abstract: Advances in LLM reasoning and tool use have enabled agentic science, yet frontier theoretical and computational physics remains challenging because research requires deep domain expertise, long-horizon reasoning, and reliable numerical computation.

By Tingjia Miao, Wenkai Jin, Jinxin Tan, Muhua Zhang, Xianghe Pang, Zexi Liu, Yuwen Du, Tian Jin, Tu Guo, Zhengliang Zhang, Jingkun Liu, Yuelin Hu, Jiejun Zhang, Yunjie Huang, Yuhan Wang, Wenbo Li, Yinuo Gao, Shuo Chen, Rui Ye, Yuzhi Zhang, Linfeng Zhang, Kun Chen, Wei Wang, Weinan E, Siheng Chen