arXiv AI By Arun Raja, Garrett M. Morris, Kian Ming A. Chai

Rethinking Molecular Text Representations for LLMs: An Empirical Study

Read the original on arXiv AI →

arXiv:2606. 03057v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for molecular tasks, but it remains unclear which molecular representation to use.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jun 5

MolE-RAG: Molecular Structure-Enhanced Retrieval-Augmented Generation for Chemistry

arXiv:2606. 05693v1 Announce Type: new Abstract: Large language models (LLMs) have shown promise for molecular property prediction, but their ability to reason over chemical structures remains limited, as molecular representations such as SMILES differ substantially from the natural language on which LLMs are primarily trained.

By Joey Chan, Wonbin Kweon, Ashley Shin, Niharika Bhattacharjee, Pengcheng Jiang, Yue Guo, Jiawei Han
arXiv Machine Learning
Sep 22

MolSC: Leveraging Substituent Contributions to Enhance Fine-grained Molecular Understanding in LLMs

MolSC is a new dataset of 181,000 substituent-level examples that captures how attaching specific substituents to molecular scaffolds changes properties such as bioactivity and physicochemical descriptors. The authors also provide MolSC-Bench, a held‑out benchmark of 1,541 examples that are disjoint from MolSC at scaffold, substituent, and molecule levels. Experiments show that training molecular large language models on MolSC markedly improves their ability to predict substituent contributions, outperforming existing models on a range of downstream chemistry tasks.

By Hyuntae Park, Sooyeon Kim, Jiwon Park, SangKeun Lee
arXiv Machine Learning
Sep 17

A Multitask Large Reasoning Model for Molecular Science

The paper introduces a multitask large reasoning model for molecular science that incorporates chemical knowledge via a multispecialist architecture, chain-of-thought supervision, and molecule-informed reinforcement learning. It coordinates prediction and inference specialists across ten molecular tasks—including description, generation, nomenclature translation, property prediction, and reaction prediction—using task-conditioned routing. The model surpasses more than 20 general-purpose and molecular large language models, improving aggregate performance by 50.3% and outperforming leading multitask baselines on most tasks, while maintaining interpretable chemical inference and demonstrating a workflow for CNS candidate generation and retrosynthetic planning.

By Pengfei Liu, Shuang Ge, Xiaobo Wang, Xin Liu, Jun Tao, Yan Li, Chao Liu, Ling Chen, Zhixiang Ren