arXiv Machine Learning

Controllable Molecular Generative Foundation Models

arXiv:2605. 15354v2 Announce Type: replace Abstract: Despite the success of foundation models in language and vision, molecular graph generation still lacks a unified framework for heterogeneous design tasks with reliable controllability.

arXiv Machine Learning
Sep 7

Training Large Language Models for Small-Molecule Design with Synthetic Task Scaling

The paper explores how large language models (LLMs) can be trained for small-molecule drug design by using synthetic tasks that are cheaper to evaluate. By employing a curriculum that gradually increases task difficulty, the authors demonstrate that LLMs can learn design strategies that outperform larger models on structure-based lead optimization. This approach shows that scaling post‑training with synthetic tasks can effectively adapt LLMs to high‑cost experimental scenarios that are otherwise infeasible to train on directly.

By Frank Hu, Shriram Chennakesavalu, Zichen Wang, Patricia Suriana, Bodhi Vani, Kirill Shmilovich, Kangway Chuang, Colin Grambow
arXiv AI
Aug 20

PGFS++: Molecular Property Improvement under Synthesis and Diversity Constraints

PGFS++ is a synthesis‑aware reinforcement learning framework that improves molecular properties while ensuring the resulting molecules can be synthesized and remain structurally similar to the input. It builds on PGFS+ by using trainable embedding lookup tables for reaction templates and second reactants, a more effective scoring function, and a refined RL algorithm. Experiments demonstrate that PGFS++ enhances target properties and preserves high output diversity, overcoming the reward‑hacking failure mode seen in earlier versions.

By Boqiao Zhang, Godbless James, Sai Krishna Gottipati, Andrew Fitzgibbon
arXiv Machine Learning
Jun 10

Synthesizable Molecular Generation via Soft-constrained GFlowNets with Rich Chemical Priors

arXiv:2602. 04119v2 Announce Type: replace Abstract: The application of generative models for experimental drug discovery campaigns is severely limited by the difficulty of designing molecules de novo that can be synthesized in practice.

By Hyeonah Kim, Minsu Kim, Celine Roget, Dionessa Biton, Louis Vaillancourt, Yves V. Brun, Yoshua Bengio, Alex Hernandez-Garcia
Hugging Face Trending Papers
Aug 19

PGFS++: Molecular Property Improvement under Synthesis and Diversity Constraints

PGFS++ is a synthesis‑aware reinforcement learning framework that improves molecular properties such as drug‑likeness or binding affinity while ensuring the resulting molecules can be synthesized and remain structurally similar to the input. It builds on PGFS+ by using trainable embedding lookup tables for reaction templates and second reactants, a more effective scoring function, and a refined RL algorithm. The method addresses a reward‑hacking failure mode by treating each input molecule as the start of a forward‑synthesis trajectory, applying learned reaction templates with in‑stock building blocks, and producing diverse, high‑quality outputs with explicit synthesis routes.

arXiv Machine Learning
Aug 26

Round-trip Reinforcement Learning: Self-Consistent Training for Better Chemical LLMs

The paper introduces Round-Trip Reinforcement Learning (RTRL), a framework that trains chemical language models to improve round‑trip consistency by rewarding successful forward and reverse transformations. By iteratively training forward and reverse mappings, RTRL leverages abundant unlabeled chemical data to enhance both consistency and overall performance across supervised, self‑supervised, and synthetic data regimes. Experiments show that RTRL outperforms strong baselines, demonstrating that round‑trip consistency can be treated as a trainable objective for more robust foundation models.

By Lecheng Kong, Xiyuan Wang, Yixin Chen, Muhan Zhang
arXiv Machine Learning
1d ago

A Large Scale Investigation of Scaling Limits in Chemical Language Models

The paper reports a large-scale, compute-controlled study of Chemical Language Models (CLMs) involving over 30,000 experiments across different molecular representations, tokenizations, model sizes, datasets, and architectures. It finds clear scaling trends in pretraining loss but shows that these improvements do not translate into proportional gains in goal-directed molecular design, with chemical syntax saturating early while semantic properties develop more slowly. The authors release a new suite of models, NovoMolGen, that achieves state-of-the-art results in drug discovery tasks, highlighting a disconnect between representation learning and downstream design and calling for new pretraining paradigms that target chemical semantics.

By Roshan Balaji, Kamran Chitsaz, Quentin Fournier, Nirav Pravinbhai Bhatt, Sarath Chandar