arXiv:2610.08400v1 Announce Type: cross
Abstract: Large-scale self-supervised pretraining has reshaped modern machine learning, substantially advancing the ability of language and vision models to ge...
By Kasper Helverskov Petersen, Rasmus Hannibal Tirsgaard, Fran\c{c}ois R J Cornet, Mikkel Jordahn, Mikkel N. Schmidt
arXiv:2606. 30961v1 Announce Type: cross Abstract: Advances in deep learning architectures and representations have enabled ML-driven chemical property prediction, but state-of-the-art (SOTA) models have remained largely confined to independent codebases and lack support for diverse chemical species.
By Jacob W. Toney, Samir Darouich, Yiran Wang, Aaron G. Garrison, Johannes K\"astner, Heather J. Kulik
Large-scale self-supervised pretraining has reshaped modern machine learning, substantially advancing the ability of language and vision models to generalize across downstream tasks. While deep learni...
arXiv:2609.26402v1 Announce Type: new
Abstract: The discovery of novel inorganic materials drives technological breakthroughs in critical fields such as computing and energy storage. Generative AI ha...
By Thomas Egg, Harry Winston Sullivan, Ellad B. Tadmor, Stefano Martiniani
The paper reports a large-scale, compute-controlled study of Chemical Language Models (CLMs) involving over 30,000 experiments across different molecular representations, tokenizations, model sizes, datasets, and architectures. It finds clear scaling trends in pretraining loss but shows that these improvements do not translate into proportional gains in goal-directed molecular design, with chemical syntax saturating early while semantic properties develop more slowly. The authors release a new suite of models, NovoMolGen, that achieves state-of-the-art results in drug discovery tasks, highlighting a disconnect between representation learning and downstream design and calling for new pretraining paradigms that target chemical semantics.
By Roshan Balaji, Kamran Chitsaz, Quentin Fournier, Nirav Pravinbhai Bhatt, Sarath Chandar
The paper introduces CrystAF, an all‑atom crystal flow‑map generation model, and evaluates where physics should be incorporated into generative crystal structure models. By applying physics‑informed post‑training, the authors improve molecular validity and crystal packing without altering sampling speed, while inference‑time corrections further refine the structures. The study demonstrates that post‑training and inference‑time physics are complementary, and that the post‑training approach transfers to other generators such as Clari‑M and MolCrystalFlow.
By Haocheng Tang, Junmei Wang, Wengong Jin