arXiv:2606. 03057v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for molecular tasks, but it remains unclear which molecular representation to use.
By Arun Raja, Garrett M. Morris, Kian Ming A. Chai
MolSC is a new dataset of 181,000 substituent-level examples that captures how attaching specific substituents to molecular scaffolds changes properties such as bioactivity and physicochemical descriptors. The authors also provide MolSC-Bench, a held‑out benchmark of 1,541 examples that are disjoint from MolSC at scaffold, substituent, and molecule levels. Experiments show that training molecular large language models on MolSC markedly improves their ability to predict substituent contributions, outperforming existing models on a range of downstream chemistry tasks.
By Hyuntae Park, Sooyeon Kim, Jiwon Park, SangKeun Lee
arXiv:2609.38744v1 Announce Type: new
Abstract: Predicting molecular properties for compounds that differ structurally from labeled training molecules is important for drug discovery and materials de...
By Jinmo Lee, Dooho Lee, Minho Jeong, Jaemin Yoo
arXiv:2608. 03855v1 Announce Type: new Abstract: Transformer models have revolutionized natural language processing (NLP), and text-based molecular representations like SMILES have successfully extended these architectures to chemistry.
By David Ming Segura, Jeremy Goumaz, Joshua W. Sin, Bojana Rankovi\'c, Philippe Schwaller
arXiv:2602. 02320v4 Announce Type: replace-cross Abstract: Molecular function is largely determined by structure.
By Feiyang Cai, Guijuan He, Yi Hu, Jingjing Wang, Joshua Luo, Tianyu Zhu, Srikanth Pilla, Gang Li, Ling Liu, Feng Luo
MolEmb is a lightweight framework that adapts multimodal large language models (MLLMs) to serve as general molecular embedding models. By aligning molecular profiles with textual descriptions in a shared embedding space using a bidirectional contrastive objective, MolEmb produces embeddings conditioned on both a molecular profile and a natural‑language semantic context. The model performs competitively on molecular property prediction and enables cross‑modal molecule‑text retrieval, while the newly introduced MolCAR benchmark demonstrates that context‑aware molecular embedding is largely a data property of the supervision.
By Xinjian Zhao, Xiangru Jian, Yaoyao Xu, Xiaozhuang Song, Wei Pang, Lei Bai, Tianshu Yu