arXiv:2609.16527v1 Announce Type: cross
Abstract: Exploring the chemical space of flexible molecules remains challenging because the vast number of possible compounds and conformations, together with...
By Michael Hanna, Julian Cremer, Zekiye Erarslan, Leonardo Medrano Sandonas
arXiv:2606. 30170v1 Announce Type: cross Abstract: Generative molecular design is shaped by simple proxy benchmarks for drug-like properties and models pretrained on large pharmaceutical datasets.
By Matthias Blaschke, Daniel Kienzle, Zsuzsanna Koczor-Benda, Julian Lorenz, Rainer Lienhart, Fabian Pauly
arXiv:2606. 30961v1 Announce Type: cross Abstract: Advances in deep learning architectures and representations have enabled ML-driven chemical property prediction, but state-of-the-art (SOTA) models have remained largely confined to independent codebases and lack support for diverse chemical species.
By Jacob W. Toney, Samir Darouich, Yiran Wang, Aaron G. Garrison, Johannes K\"astner, Heather J. Kulik
arXiv:2606. 17077v1 Announce Type: cross Abstract: Proton dissociation constants (pKa) are critical for functional molecule discovery and molecular modeling.
By Wang Rui, Liu Dinghao
The paper presents a closed‑loop molecule generation pipeline that iteratively retrains on new quantum‑chemical simulation data, overcoming limitations of static generative models. This approach produces molecules whose properties extend up to 0.44 standard deviations beyond the training set and improves out‑of‑distribution classification accuracy by 79%. By conditioning on thermodynamic stability during the loop, the method yields a 3.5‑fold increase in the proportion of stable, potentially synthesizable molecules.
By Evan R. Antoniuk, Peggy Li, Nathan Keilbart, Stephen Weitzner, Bhavya Kailkhura, Anna M. Hiszpanski
arXiv:2606. 01220v1 Announce Type: cross Abstract: Generating molecules that simultaneously satisfy drug-like properties and conform to the 3D structure of a target protein is a core challenge in structure-based drug design (SBDD).
By Guang Lin, Shikui Tu, Lei Xu
arXiv:2607. 20551v1 Announce Type: cross Abstract: Effective molecular representation learning is crucial for accurate molecular property prediction.
By Tianming Han, Li Zhang, Qi Zhao
Mol-JEPA is a scalable multimodal framework that learns molecular world models by using modality masking instead of suboptimal perturbations. It incorporates diverse data such as molecular structures, cellular phenotypes, binding affinities, ADMET profiles, quantum chemistry simulations, and other drug‑discovery information. Benchmarks show that the representations it learns perform strongly, highlighting the benefit of embedding biochemical context via latent‑space prediction.
By Florian Rottach, Sebastian Schieferdecker, William Rudman, Randall Balestriero, Carsten Eickhoff
arXiv:2607. 19816v1 Announce Type: cross Abstract: Determining molecular structures from spectroscopic data remains fundamentally challenging because the inverse problem is intrinsically underdetermined: individual spectra are sparse, low-dimensional, and encode only partial structural evidence relative to the vast space of possible molecules.
By Chengchun Liu, Zhiyuan Yan, Li Yuan, Hao Li, Boxuan Zhao, Yonghong Tian, Bartosz A. Grzybowski, Fanyang Mo
arXiv:2606. 23856v1 Announce Type: new Abstract: Generative molecular models for drug design are a promising direction with much active research.
By Konstantin Yatsenko, Arvind Thiagarajan
arXiv:2606. 14498v1 Announce Type: cross Abstract: Predicting the Kohn-Sham Hamiltonian with machine learning can accelerate density functional theory while retaining access to molecular orbitals, energy levels, and electronic-structure observables that energy-only surrogates cannot resolve.
By Yunhong Lou, Xihang Yue, Xinran Wei, Tianqi Deng, Linchao Zhu
NEAT-POCKET is a pocket‑conditioned extension of the autoregressive NEAT model that generates 3D molecules atom by atom within protein binding pockets, maintaining atom permutation invariance and explicitly modeling hydrogen atoms. It outperforms existing baselines on the CrossDocked and SPINDR datasets, achieving competitive structure‑based generation performance while sampling significantly faster. The model also supports pocket‑conditioned fragment completion, a capability directly useful for lead optimization and scaffold elaboration in drug design.
By Roxane Axel Jacob, Daniel Rose, Thierry Langer, Johannes Kirchmair