The paper introduces two Hessian-based data augmentation techniques—UniAug and ModeAug—to improve machine‑learning interatomic potentials (MLIPs). These methods use simple Taylor expansions to generate augmented configurations without modifying training objectives or increasing computational overhead. Experiments on both non‑equilibrium and equilibrium datasets show that the augmentations enhance model accuracy and provide practical guidelines for specific tasks.
By Bumju Kwak, Jeonghee Jo
arXiv:2509.21624v4 Announce Type: replace
Abstract: Molecular Hessians, the second derivatives of the potential energy, are fundamental to many workflows in computational chemistry. Usually, accurate...
By Andreas Burger, Luca Thiede, Nikolaj R{\o}nne, Varinia Bernales, Nandita Vijaykumar, Tejs Vegge, Arghya Bhowmik, Alan Aspuru-Guzik
arXiv:2509. 21624v3 Announce Type: replace Abstract: Fundamental tasks in computational chemistry, from transition state search to vibrational analysis, rely on molecular Hessians, which are the second derivatives of the potential energy.
By Andreas Burger, Luca Thiede, Nikolaj R{\o}nne, Varinia Bernales, Nandita Vijaykumar, Tejs Vegge, Arghya Bhowmik, Alan Aspuru-Guzik
arXiv:2606. 04100v1 Announce Type: new Abstract: Machine learning interatomic potentials (MLIPs) enable efficient and accurate atomistic simulations but depend critically on the quality and diversity of the training data.
By Joanna Zou, Fraser Birks, Dallas Foster, Youssef Marzouk
arXiv:2606. 24999v1 Announce Type: new Abstract: High-dimensional partial differential equations (PDEs) with unknown coefficients arise widely in scientific machine learning, including continuous-time reinforcement learning, yet solving them efficiently in a data-driven way remains challenging.
By Yanwei Jia, Du Ouyang, Huy\^en Pham, Xun Yu Zhou
arXiv:2602. 04861v2 Announce Type: replace-cross Abstract: Machine Learning Interatomic Potentials (MLIPs) sometimes fail to reproduce the physical smoothness of the quantum potential energy surface (PES), leading to erroneous behavior in downstream simulations that standard energy and force regression evaluations can miss.
By Ryan Liu, Eric Qu, Tobias Kreiman, Samuel M. Blau, Aditi S. Krishnapriyan
arXiv:2606. 00401v1 Announce Type: cross Abstract: Simulating large molecular systems comprising thousands of atoms requires highly scalable methodologies.
By Abhiram Badrinarayanan, Davor Davidovic, Edoardo Di Napoli, Jurica Novak, Luigi Genovese, Gustavo Ramirez-Hidalgo, Xinzhe Wu
BOOM is a new benchmark for evaluating out‑of‑distribution (OOD) molecular property predictions in machine learning. It provides chemically‑informed tests across common property prediction tasks and assesses over 150 model‑task combinations. The study shows that current models, including chemical foundation models, struggle to generalize OOD, with the best model still exhibiting three times higher error than in‑distribution predictions.
By Evan R. Antoniuk, Shehtab Zaman, Tal Ben-Nun, Peggy Li, James Diffenderfer, Busra Sahin, Obadiah Smolenski, Everett Grethel, Tim Hsu, Anna M. Hiszpanski, Kenneth Chiu, Bhavya Kailkhura, Brian Van Essen
arXiv:2608. 11020v1 Announce Type: new Abstract: We systematically investigate finite-difference (FD) derivative computation in Physics-Informed Neural Networks (PINNs) as an alternative to automatic differentiation (AD).
By Maciej J. Mikulski, Tadeusz Uhl
arXiv:2609.33078v2 Announce Type: replace
Abstract: Automatic differentiation (AD) lets neural networks compute derivatives of governing equations to machine precision, and this precision has made it...
By Ameya D. Jagtap
The paper introduces ADAPT, a lightweight machine‑learning force field that replaces graph neural networks with a direct coordinates‑in‑space Transformer encoder to model all pairwise atomic interactions. Applied to silicon point defects, ADAPT reduces force prediction error by about 22% and energy prediction error by roughly 40% compared to a state‑of‑the‑art GNN model, while also cutting computational cost. This approach addresses common GNN issues such as oversmoothing, oversquashing, and poor long‑range interaction representation, which are especially problematic for point defect modeling.
By Evan Dramko, Yihuang Xiong, Yizhi Zhu, Geoffroy Hautier, Thomas Reps, Christopher Jermaine, Anastasios Kyrillidis
arXiv:2507.09001v4 Announce Type: replace-cross
Abstract: Machine learning (ML) models for electronic structure typically rely on large datasets generated by computationally expensive Kohn-Sham densi...
By Sazzad Hossain, Ponkrshnan Thiagarajan, Shashank Pathrudkar, Stephanie Taylor, Abhijeet S. Gangan, Amartya S. Banerjee, Susanta Ghosh