arXiv:2606. 03057v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for molecular tasks, but it remains unclear which molecular representation to use.
By Arun Raja, Garrett M. Morris, Kian Ming A. Chai
MolSC is a new dataset of 181,000 substituent-level examples that captures how attaching specific substituents to molecular scaffolds changes properties such as bioactivity and physicochemical descriptors. The authors also provide MolSC-Bench, a held‑out benchmark of 1,541 examples that are disjoint from MolSC at scaffold, substituent, and molecule levels. Experiments show that training molecular large language models on MolSC markedly improves their ability to predict substituent contributions, outperforming existing models on a range of downstream chemistry tasks.
By Hyuntae Park, Sooyeon Kim, Jiwon Park, SangKeun Lee
arXiv:2609.38744v1 Announce Type: new
Abstract: Predicting molecular properties for compounds that differ structurally from labeled training molecules is important for drug discovery and materials de...
By Jinmo Lee, Dooho Lee, Minho Jeong, Jaemin Yoo
arXiv:2608. 03855v1 Announce Type: new Abstract: Transformer models have revolutionized natural language processing (NLP), and text-based molecular representations like SMILES have successfully extended these architectures to chemistry.
By David Ming Segura, Jeremy Goumaz, Joshua W. Sin, Bojana Rankovi\'c, Philippe Schwaller
arXiv:2602. 02320v4 Announce Type: replace-cross Abstract: Molecular function is largely determined by structure.
By Feiyang Cai, Guijuan He, Yi Hu, Jingjing Wang, Joshua Luo, Tianyu Zhu, Srikanth Pilla, Gang Li, Ling Liu, Feng Luo
MolEmb is a lightweight framework that adapts multimodal large language models (MLLMs) to serve as general molecular embedding models. By aligning molecular profiles with textual descriptions in a shared embedding space using a bidirectional contrastive objective, MolEmb produces embeddings conditioned on both a molecular profile and a natural‑language semantic context. The model performs competitively on molecular property prediction and enables cross‑modal molecule‑text retrieval, while the newly introduced MolCAR benchmark demonstrates that context‑aware molecular embedding is largely a data property of the supervision.
By Xinjian Zhao, Xiangru Jian, Yaoyao Xu, Xiaozhuang Song, Wei Pang, Lei Bai, Tianshu Yu
arXiv:2608. 10480v1 Announce Type: new Abstract: Large language models (LLMs) are widely applied across chemical tasks, such as molecular property prediction, which underpins drug discovery.
By Junwoo Park, Minyoung Shin, Cheol Soon Lee, Sujee Lee
arXiv:2603. 25857v3 Announce Type: replace Abstract: The capabilities of large language models (LLMs) have expanded beyond natural language processing to scientific prediction tasks, including molecular property prediction.
By Matthias Busch, Marius Tacke, Sviatlana V. Lamaka, Mikhail L. Zheludkevich, Christian J. Cyron, Christian Feiler, Roland C. Aydin
Fraglingo is an autoregressive fragment-based molecular generator that jointly models fragment identity and attachment in a continuous latent space. It predicts attachment-aware fragment embeddings using a wildcard-anchored readout that captures the growing molecule’s active attachment site, then retrieves the next fragment via latent-space nearest-neighbor search. This approach allows new fragments to be added at inference time without retraining and achieves stronger joint property control on benchmarks while maintaining high validity, uniqueness, and novelty.
By Thao Nguyen, Jeonghwan Kim, Zhenhailong Wang, Heng Ji
Large language models (LLMs) are widely applied across chemical tasks, such as molecular property prediction, which underpins drug discovery. Molecular LLMs represent a molecule through several modalities, notably a 1D SMILES sequence or a 2D molecular graph.
The study evaluates four pretrained molecular language models on six virtual libraries covering drug discovery, organic materials, and catalysis. It finds that native embeddings vary widely in performance, while molecular fingerprints remain consistently strong. Fine‑tuning the models on library‑specific data markedly improves sample efficiency, with several adapted encoders outperforming others across all tasks.
By Henrik Wille, Luis-Finley Sch\"utz, Felix Strieth-Kalthoff
arXiv:2603. 25062v2 Announce Type: replace Abstract: Autoregressive molecular models assign probability to molecular serializations even though chemical identity is invariant to serialization.
By Xinyu Wang, Fei Dou, Jinbo Bi, Minghu Song