arXiv AI By Miruna Cretu, Alex Abrudan, Antonia Panescu, Tynan Perez, Rishabh Anand, N. Benjamin Erichson, Michael W. Mahoney, Samuel Blau, Joseph Jacobson, Rafael G\'omez-Bombarelli, Rex Ying, Tuomas Knowles, Pietro Li\`o, Alex Morehead

Zatom-2: Multitask Pretraining on Atomistic Data for Generative Modeling across Domains

Read the original on arXiv AI →

Zatom-2 is a multitask generative model for atomistic data that has been pretrained on about five million structures from the OMol25 and OMat24 datasets. It uses a multiscale Transformer with conditional flow matching to support tasks such as generation, structure prediction, and energy/force prediction for both molecules and materials. The model outperforms its predecessor, Zatom-1, on molecular distribution fidelity and benchmark generation tasks, and improves protein backbone designability from 67.8% to 74.8% after finetuning on 2,000 protein domains.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 1

ElemeNet: Multiscale Molecular Machine Learning with Uncertainty Quantification Across the Periodic Table

arXiv:2606. 30961v1 Announce Type: cross Abstract: Advances in deep learning architectures and representations have enabled ML-driven chemical property prediction, but state-of-the-art (SOTA) models have remained largely confined to independent codebases and lack support for diverse chemical species.

By Jacob W. Toney, Samir Darouich, Yiran Wang, Aaron G. Garrison, Johannes K\"astner, Heather J. Kulik
arXiv Machine Learning
Oct 3

A Large Scale Investigation of Scaling Limits in Chemical Language Models

The paper reports a large-scale, compute-controlled study of Chemical Language Models (CLMs) involving over 30,000 experiments across different molecular representations, tokenizations, model sizes, datasets, and architectures. It finds clear scaling trends in pretraining loss but shows that these improvements do not translate into proportional gains in goal-directed molecular design, with chemical syntax saturating early while semantic properties develop more slowly. The authors release a new suite of models, NovoMolGen, that achieves state-of-the-art results in drug discovery tasks, highlighting a disconnect between representation learning and downstream design and calling for new pretraining paradigms that target chemical semantics.

By Roshan Balaji, Kamran Chitsaz, Quentin Fournier, Nirav Pravinbhai Bhatt, Sarath Chandar
arXiv AI
Sep 30

Where Should Physics Enter a Molecular Crystal Generator?

The paper introduces CrystAF, an all‑atom crystal flow‑map generation model, and evaluates where physics should be incorporated into generative crystal structure models. By applying physics‑informed post‑training, the authors improve molecular validity and crystal packing without altering sampling speed, while inference‑time corrections further refine the structures. The study demonstrates that post‑training and inference‑time physics are complementary, and that the post‑training approach transfers to other generators such as Clari‑M and MolCrystalFlow.

By Haocheng Tang, Junmei Wang, Wengong Jin