arXiv AI

Procedural Pretraining for Molecular Property Prediction

The paper proposes a three‑stage training pipeline that begins with procedural pretraining on abstract, procedurally generated data, followed by molecular pretraining on SMILES, and finally downstream fine‑tuning for molecular property prediction. Experiments show that procedural pretraining improves downstream performance—e.g., a 4.8% error reduction on Lipophilicity—especially when labeled data are scarce, and that the benefit peaks at an intermediate procedural training budget. Analysis indicates that transferable knowledge resides mainly in attention layers, while feed‑forward layers may over‑specialize.

arXiv Machine Learning
Jul 7

On the Design Space of Discrete Diffusion Online Adaptation for Molecular Optimization

arXiv:2607. 02834v1 Announce Type: new Abstract: Molecular optimization often starts from a pretrained generative model that captures a broad prior over valid molecular structures.

By Trevor Chen, Ariel Dai, Jason Yang, Riccardo De Santi, Daniel Khalil, Wenda Chu, Nate Gruver, Pranav Murugan, Alexander F. G. Goldberg, Maruan Al-Shedivat, Yisong Yue
arXiv Machine Learning
Aug 20

Monroe: A Molecular Foundation Model for In-Context Probabilistic Inference

Monroe is a new molecular foundation model that improves upon existing models by pre‑training on over 81 million molecules from the PM6 quantum chemistry dataset, enhancing stereochemistry representation, and introducing novel training losses such as conformer denoising and embedding decorrelation. It also incorporates a prior‑data‑fitted model (TabPFN) for downstream in‑context prediction and demonstrates superior performance on Polaris benchmarks and activity cliff tests. Ablation studies show that the PFN‑based downstream approach can upgrade other models, producing state‑of‑the‑art variants MiniMol_PFN and CheMeleon_PFN.

By Blazej Banaszewski, Andrew W. Fitzgibbon
arXiv Machine Learning
Sep 7

Training Large Language Models for Small-Molecule Design with Synthetic Task Scaling

The paper explores how large language models (LLMs) can be trained for small-molecule drug design by using synthetic tasks that are cheaper to evaluate. By employing a curriculum that gradually increases task difficulty, the authors demonstrate that LLMs can learn design strategies that outperform larger models on structure-based lead optimization. This approach shows that scaling post‑training with synthetic tasks can effectively adapt LLMs to high‑cost experimental scenarios that are otherwise infeasible to train on directly.

By Frank Hu, Shriram Chennakesavalu, Zichen Wang, Patricia Suriana, Bodhi Vani, Kirill Shmilovich, Kangway Chuang, Colin Grambow
arXiv Machine Learning
Aug 5

In-Context Molecular Property Prediction with LLMs: A Blinding Study on Memorization and Knowledge Conflicts

arXiv:2603. 25857v3 Announce Type: replace Abstract: The capabilities of large language models (LLMs) have expanded beyond natural language processing to scientific prediction tasks, including molecular property prediction.

By Matthias Busch, Marius Tacke, Sviatlana V. Lamaka, Mikhail L. Zheludkevich, Christian J. Cyron, Christian Feiler, Roland C. Aydin