arXiv:2607. 03007v1 Announce Type: cross Abstract: Recent advances in molecular large language models have led to strong performance on molecular understanding and generation tasks, yet these gains often come without reliable structural grounding.
By Wenda Wang, Jinjia Feng, Zhewei Wei
arXiv:2607. 02140v1 Announce Type: new Abstract: Chemical language models (CLMs) are trained with linearized representations such as SMILES, yet it remains unclear which chemically meaningful substructures they encode.
By Anna Karnysheva, Dietrich Klakow, Ji-Ung Lee
arXiv:2606. 03057v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for molecular tasks, but it remains unclear which molecular representation to use.
By Arun Raja, Garrett M. Morris, Kian Ming A. Chai
arXiv:2608. 11283v1 Announce Type: cross Abstract: Computation-ready metal-organic framework (MOF) databases are essential for high-throughput screening, yet many reported crystal structures remain chemically unreasonable or disordered, compromising simulation fidelity.
By Guobin Zhao, Xiao-Yan Li
HiPoly is a polymer-native AI framework that uses a three-level hierarchical graph architecture built on the G2RINS representation to process complete polymer descriptions. It encodes stochastic inter-monomer connectivity, composition, and molecular weight directly within its architecture, enabling end-to-end workflows from experimental data to property prediction, generative design, and physics-based validation. The framework achieves state-of-the-art accuracy for thermophysical properties of multi-component polymer systems and demonstrates generative design by discovering sustainable, PFAS-free alternatives with target surface-energy properties.
By Ge Sun, Gervasio Zaldivar, Yuan Tian, Gustavo Perez Lemus, Juhae Park, Dasha Safarian, Ming Han, Juan J. de Pablo
The paper reports a large-scale, compute-controlled study of Chemical Language Models (CLMs) involving over 30,000 experiments across different molecular representations, tokenizations, model sizes, datasets, and architectures. It finds clear scaling trends in pretraining loss but shows that these improvements do not translate into proportional gains in goal-directed molecular design, with chemical syntax saturating early while semantic properties develop more slowly. The authors release a new suite of models, NovoMolGen, that achieves state-of-the-art results in drug discovery tasks, highlighting a disconnect between representation learning and downstream design and calling for new pretraining paradigms that target chemical semantics.
By Roshan Balaji, Kamran Chitsaz, Quentin Fournier, Nirav Pravinbhai Bhatt, Sarath Chandar
arXiv:2607. 29256v1 Announce Type: new Abstract: Designing polyimide structures with specific glass transition temperatures (Tg) is highly challenging.
By Junquan Hu, Zhihui Wang, Peng Xu, Xinru Guo, Xintong Li, Kun Lu, Ben Fei
arXiv:2605. 16823v2 Announce Type: replace Abstract: Large language models succeed by combining large-scale pretraining with meaningful discrete tokens.
By Takayuki Kimura
arXiv:2502. 07027v4 Announce Type: replace-cross Abstract: Molecular Relational Learning (MRL) is widely applied in natural sciences to predict relationships between molecular pairs by extracting structural features.
By Peiliang Zhang, Jingling Yuan, Qing Xie, Yongjun Zhu, Chao Che, Lin Li
arXiv:2606. 05693v1 Announce Type: new Abstract: Large language models (LLMs) have shown promise for molecular property prediction, but their ability to reason over chemical structures remains limited, as molecular representations such as SMILES differ substantially from the natural language on which LLMs are primarily trained.
By Joey Chan, Wonbin Kweon, Ashley Shin, Niharika Bhattacharjee, Pengcheng Jiang, Yue Guo, Jiawei Han
The paper investigates how explicitly supervising molecular embeddings with a molecule’s Bemis‑Murcko scaffold influences representation learning. Experiments compare Euclidean and Lorentz contrastive objectives under two augmentation strengths, showing that scaffold‑supervised models consistently group molecules by identical and related scaffolds. These embeddings also enhance property prediction on several tasks, though the magnitude of improvement varies with the target property and the geometry used.
By David Sulu, Lorenzo Di Fruscia, Jana M. Weber
arXiv:2510. 16023v2 Announce Type: replace Abstract: Linear polymers, macromolecules formed from monomers covalently bonded into continuous chains, underpin countless technologies and are indispensable to modern life.
By Fanmeng Wang, Ruochao Wang, Shan Mei, Wentao Guo, Hongshuai Wang, Qi Ou, Zhifeng Gao, Hongteng Xu