arXiv:2608. 02688v1 Announce Type: cross Abstract: Phenotypic drug discovery enables the discovery of functional relationships between molecular structures and cellular responses.
By Xuan Lin, Jingyu Sheng, Tengfei Ma, Li Sun, Dapeng Xiong
arXiv:2609.37384v1 Announce Type: new
Abstract: Molecular representation learning is central to computer-aided drug discovery. Molecular graphs, SMILES strings, and 3D conformations provide complemen...
By Linqing Mo, Jiayu Zhou, Bin Chen
The paper reports a large-scale, compute-controlled study of Chemical Language Models (CLMs) involving over 30,000 experiments across different molecular representations, tokenizations, model sizes, datasets, and architectures. It finds clear scaling trends in pretraining loss but shows that these improvements do not translate into proportional gains in goal-directed molecular design, with chemical syntax saturating early while semantic properties develop more slowly. The authors release a new suite of models, NovoMolGen, that achieves state-of-the-art results in drug discovery tasks, highlighting a disconnect between representation learning and downstream design and calling for new pretraining paradigms that target chemical semantics.
By Roshan Balaji, Kamran Chitsaz, Quentin Fournier, Nirav Pravinbhai Bhatt, Sarath Chandar
The paper introduces Align-React, a chemical reaction representation learning framework that incorporates atomic correspondence between reactants and products, an adapter for embedding reaction conditions, and a Reaction-Center-Aware attention mechanism. These components enable the model to capture precise molecular transformations and focus on critical functional groups, leading to improved performance across a variety of organic reaction tasks. The framework outperforms existing architectures on most benchmark datasets.
By Kaipeng Zeng, Xianbin Liu, Yu Zhang, Xiaokang Yang, Yaohui Jin, Yanyan Xu
arXiv:2604. 06336v2 Announce Type: replace-cross Abstract: Fragment-level representations provide a natural way to capture recurring molecular substructures and reuse their learned representations across molecules.
By Yi Yang, Ovidiu Daescu
The paper investigates how explicitly supervising molecular embeddings with a molecule’s Bemis‑Murcko scaffold influences representation learning. Experiments compare Euclidean and Lorentz contrastive objectives under two augmentation strengths, showing that scaffold‑supervised models consistently group molecules by identical and related scaffolds. These embeddings also enhance property prediction on several tasks, though the magnitude of improvement varies with the target property and the geometry used.
By David Sulu, Lorenzo Di Fruscia, Jana M. Weber
The paper introduces ReGeoDTA, a framework that preserves chemical heterogeneity and continuous geometric relationships in drug and protein representations to improve drug–target affinity prediction. Experiments on three benchmark datasets show that maintaining representation fidelity consistently enhances predictive accuracy across various DTA architectures, while degrading representations harms performance and cannot be recovered by more complex downstream models. The study highlights representation fidelity as a key upstream design principle for accurate and generalizable affinity prediction.
By Yixiao Li, Yining Qian, Yefan Chen, Zenghui Chen, Jiayue Sun, Yuhai Zhao, Cheng Tan, An-Yang Lu
arXiv:2603. 25062v2 Announce Type: replace Abstract: Autoregressive molecular models assign probability to molecular serializations even though chemical identity is invariant to serialization.
By Xinyu Wang, Fei Dou, Jinbo Bi, Minghu Song
arXiv:2607. 02140v1 Announce Type: new Abstract: Chemical language models (CLMs) are trained with linearized representations such as SMILES, yet it remains unclear which chemically meaningful substructures they encode.
By Anna Karnysheva, Dietrich Klakow, Ji-Ung Lee
arXiv:2606. 03057v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for molecular tasks, but it remains unclear which molecular representation to use.
By Arun Raja, Garrett M. Morris, Kian Ming A. Chai
arXiv:2610.02186v1 Announce Type: cross
Abstract: Molecular learning models are strongly shaped by their underlying representations. Yet standard sequential and graph formalisms struggle to explicitl...
By Yiming Huang, Yujie Zeng, Vijay Prakash Dwivedi, Simone Foti, Jianmin Wang, Jure Leskovec, Tolga Birdal
arXiv:2510.07289v2 Announce Type: replace
Abstract: Molecular graph representation learning is widely used in chemical and biomedical research. While pre-trained 2D graph encoders have demonstrated s...
By Xingtong Yu, Chang Zhou, Xinming Zhang, Yuan Fang