arXiv:2606. 30961v1 Announce Type: cross Abstract: Advances in deep learning architectures and representations have enabled ML-driven chemical property prediction, but state-of-the-art (SOTA) models have remained largely confined to independent codebases and lack support for diverse chemical species.
By Jacob W. Toney, Samir Darouich, Yiran Wang, Aaron G. Garrison, Johannes K\"astner, Heather J. Kulik
arXiv:2609.22663v1 Announce Type: cross
Abstract: Molecular systems have many degrees of freedom, but their metastable behavior can often be described by a few collective variables. Identifying these...
By Venkata Sai Sreyas Adury (Chemical Physics Program and Institute for Physical Science and Technology, University of Maryland), Pratyush Tiwary (Biophysics Program and Institute for Physical Science and Technology, University of Maryland, Department of Chemistry and Biochemistry and Institute for Physical Science and Technology, University of Maryland, University of Maryland Institute for Health Computing, Bethesda, USA)
BOOM is a new benchmark for evaluating out‑of‑distribution (OOD) molecular property predictions in machine learning. It provides chemically‑informed tests across common property prediction tasks and assesses over 150 model‑task combinations. The study shows that current models, including chemical foundation models, struggle to generalize OOD, with the best model still exhibiting three times higher error than in‑distribution predictions.
By Evan R. Antoniuk, Shehtab Zaman, Tal Ben-Nun, Peggy Li, James Diffenderfer, Busra Sahin, Obadiah Smolenski, Everett Grethel, Tim Hsu, Anna M. Hiszpanski, Kenneth Chiu, Bhavya Kailkhura, Brian Van Essen
arXiv:2602. 24007v3 Announce Type: replace-cross Abstract: Protein function relies on dynamic conformational ensembles, yet current generative models like AlphaFold3 often fail to produce ensembles that match experimental data.
By Advaith Maddipatla, Anar Rzayev, Marco Pegoraro, Martin Pacesa, Paul Schanda, Ailie Marx, Sanketh Vedula, Alex M. Bronstein
arXiv:2607. 10887v1 Announce Type: cross Abstract: Machine learning interatomic potentials (MLPs) have revolutionized atomistic modeling, offering the potential to replace traditional methods like Density Functional Theory (DFT).
By Jan Eckwert, Julija Zavadlav
arXiv:2608. 02642v1 Announce Type: cross Abstract: Accelerating scientific discovery is among the most consequential applications of AI, and computational biomolecular simulation stands out as a particularly promising target within this broader effort.
By Nithishwer Mouroug Anand, Wei-Tse Hsu, Kyle Vaccaro, Eden James Gage, Jonathan David Colburn, Linda Xi Phan, Minjoon Seo, Kevin Guan, Philip C. Biggin
arXiv:2606. 24983v1 Announce Type: cross Abstract: Implicit solvent machine learning potentials (MLPs) offer a powerful route to bridging the gap between accuracy and efficiency in molecular simulations.
By Linying Zhang, Julija Zavadlav
arXiv:2607. 19519v1 Announce Type: cross Abstract: Most 3D properties relevant to molecular design, including free energies and shape descriptors, are $\textit{expectations}$ over the Boltzmann distribution over 3D configurations of a molecular graph.
By Selma Moqvist, Richard Beckmann, Ross Irwin, Roc\'io Mercado, Simon Olsson
arXiv:2606. 30687v1 Announce Type: cross Abstract: Diffusion models are increasingly utilized for modeling molecular structures and conformational ensembles, yet the thermodynamic meaning of their learned representations and scores remains elusive.
By Wenjie Xi
The paper reports a large-scale, compute-controlled study of Chemical Language Models (CLMs) involving over 30,000 experiments across different molecular representations, tokenizations, model sizes, datasets, and architectures. It finds clear scaling trends in pretraining loss but shows that these improvements do not translate into proportional gains in goal-directed molecular design, with chemical syntax saturating early while semantic properties develop more slowly. The authors release a new suite of models, NovoMolGen, that achieves state-of-the-art results in drug discovery tasks, highlighting a disconnect between representation learning and downstream design and calling for new pretraining paradigms that target chemical semantics.
By Roshan Balaji, Kamran Chitsaz, Quentin Fournier, Nirav Pravinbhai Bhatt, Sarath Chandar
Mol-JEPA is a scalable multimodal framework that learns molecular world models by using modality masking instead of suboptimal perturbations. It incorporates diverse data such as molecular structures, cellular phenotypes, binding affinities, ADMET profiles, quantum chemistry simulations, and other drug‑discovery information. Benchmarks show that the representations it learns perform strongly, highlighting the benefit of embedding biochemical context via latent‑space prediction.
By Florian Rottach, Sebastian Schieferdecker, William Rudman, Randall Balestriero, Carsten Eickhoff
arXiv:2510.07589v2 Announce Type: replace-cross
Abstract: Understanding molecular structure, dynamics, and reactivity requires bridging processes that occur across widely separated time scales. Conve...
By Juan Viguera Diez, Mathias Schreiner, Simon Olsson