arXiv Machine Learning

De novo molecular generation with optical property preconditioning at the token level

arXiv:2606. 08221v1 Announce Type: new Abstract: Designing OLED molecules with targeted optical properties remains challenging due to the scarcity of high-quality data and the limited reliability of conditional control in generative models across chemical motifs.

arXiv Machine Learning
Jul 23

OLEDLM: A Unified Language Model for OLED Molecular Design

arXiv:2607. 20194v1 Announce Type: new Abstract: The development of organic light-emitting diode (OLED) materials faces the compounded challenges of an astronomically large chemical space, stringent quantum-chemical constraints, and a scarcity of labeled data.

By Fukang Wen, Yuchong Tang, Jingyuan Li, Beichen Wang, Yixuan Jiang, Xiaoyi Jiang, Yaxuan Liu, Shunyu Wang, Zuoqiang Shi, Yi Zhu, Yanan Zhu, Pipi Hu
arXiv Machine Learning
Sep 17

Active Learning Enables Generation of Molecules that Advance the Known Pareto Front

The paper presents a closed‑loop molecule generation pipeline that iteratively retrains on new quantum‑chemical simulation data, overcoming limitations of static generative models. This approach produces molecules whose properties extend up to 0.44 standard deviations beyond the training set and improves out‑of‑distribution classification accuracy by 79%. By conditioning on thermodynamic stability during the loop, the method yields a 3.5‑fold increase in the proportion of stable, potentially synthesizable molecules.

By Evan R. Antoniuk, Peggy Li, Nathan Keilbart, Stephen Weitzner, Bhavya Kailkhura, Anna M. Hiszpanski
arXiv Machine Learning
1d ago

A Large Scale Investigation of Scaling Limits in Chemical Language Models

The paper reports a large-scale, compute-controlled study of Chemical Language Models (CLMs) involving over 30,000 experiments across different molecular representations, tokenizations, model sizes, datasets, and architectures. It finds clear scaling trends in pretraining loss but shows that these improvements do not translate into proportional gains in goal-directed molecular design, with chemical syntax saturating early while semantic properties develop more slowly. The authors release a new suite of models, NovoMolGen, that achieves state-of-the-art results in drug discovery tasks, highlighting a disconnect between representation learning and downstream design and calling for new pretraining paradigms that target chemical semantics.

By Roshan Balaji, Kamran Chitsaz, Quentin Fournier, Nirav Pravinbhai Bhatt, Sarath Chandar
arXiv AI
Aug 19

Domain-Adapted Molecular Language Models for Efficient Search of Make-on-Demand Libraries

The study evaluates four pretrained molecular language models on six virtual libraries covering drug discovery, organic materials, and catalysis. It finds that native embeddings vary widely in performance, while molecular fingerprints remain consistently strong. Fine‑tuning the models on library‑specific data markedly improves sample efficiency, with several adapted encoders outperforming others across all tasks.

By Henrik Wille, Luis-Finley Sch\"utz, Felix Strieth-Kalthoff