arXiv:2607. 03007v1 Announce Type: cross Abstract: Recent advances in molecular large language models have led to strong performance on molecular understanding and generation tasks, yet these gains often come without reliable structural grounding.
By Wenda Wang, Jinjia Feng, Zhewei Wei
WOMBAT is a benchmark comprising 14 whitebox graph neural networks (GNNs) whose message‑passing weights are manually set to detect specific SMARTS motifs. Each model’s decision rule is explicitly known, providing a ground truth for attribution that allows researchers to identify and study errors in post‑hoc explainers such as GNNExplainer, PGExplainer, and Integrated Gradients. The authors validate the models on millions of PubChem molecules, demonstrate how Integrated Gradients can be misled to spread attribution, and release the dataset, models, and evaluation code for future XAI tool development.
By Dominik Matuszek, Bartosz Zieli\'nski, Tomasz Danel, Dawid Rymarczyk
Large language models (LLMs) are widely applied across chemical tasks, such as molecular property prediction, which underpins drug discovery. Molecular LLMs represent a molecule through several modalities, notably a 1D SMILES sequence or a 2D molecular graph.
arXiv:2606. 05693v1 Announce Type: new Abstract: Large language models (LLMs) have shown promise for molecular property prediction, but their ability to reason over chemical structures remains limited, as molecular representations such as SMILES differ substantially from the natural language on which LLMs are primarily trained.
By Joey Chan, Wonbin Kweon, Ashley Shin, Niharika Bhattacharjee, Pengcheng Jiang, Yue Guo, Jiawei Han
arXiv:2608. 10480v1 Announce Type: new Abstract: Large language models (LLMs) are widely applied across chemical tasks, such as molecular property prediction, which underpins drug discovery.
By Junwoo Park, Minyoung Shin, Cheol Soon Lee, Sujee Lee
arXiv:2606. 12113v1 Announce Type: cross Abstract: Transformer-based language models for SMILES strings suffer from a locality gap: standard character-level tokenization fragments chemically meaningful motifs, forcing models to repeatedly learn local syntax at the expense of long-range dependencies.
By Xinni Zhang, Zijing Liu, He Cao, Yu Li, Irwin King
arXiv:2607. 08996v1 Announce Type: cross Abstract: Graph Neural Networks have emerged as a powerful tool for the fast and accurate prediction of various crystal properties.
By Shrimon Mukherjee, Kishalay Das, Partha Basuchowdhuri, Pawan Goyal, Niloy Ganguly
arXiv:2606. 03057v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for molecular tasks, but it remains unclear which molecular representation to use.
By Arun Raja, Garrett M. Morris, Kian Ming A. Chai
arXiv:2607. 01800v1 Announce Type: new Abstract: Large Language Models (LLMs) have recently shown promise in molecular discovery, yet a gap remains between their probabilistic nature over discrete sequential tokens and the rigid topological constraints of chemical space.
By Jiatong Li, Weida Wang, Changmeng Zheng, Shufei Zhang, Yatao Bian, Xiao-yong Wei, Qing Li
arXiv:2510.07289v2 Announce Type: replace
Abstract: Molecular graph representation learning is widely used in chemical and biomedical research. While pre-trained 2D graph encoders have demonstrated s...
By Xingtong Yu, Chang Zhou, Xinming Zhang, Yuan Fang
arXiv:2603. 25857v3 Announce Type: replace Abstract: The capabilities of large language models (LLMs) have expanded beyond natural language processing to scientific prediction tasks, including molecular property prediction.
By Matthias Busch, Marius Tacke, Sviatlana V. Lamaka, Mikhail L. Zheludkevich, Christian J. Cyron, Christian Feiler, Roland C. Aydin
arXiv:2608.30674v1 Announce Type: cross
Abstract: Accurate molecular property prediction requires both statistical reliability and chemical reasoning. Graph neural networks can be calibrated directly...
By Wentao Li, Jiangjie Qiu, Yijun Li, Leyi Zhao, Xiaonan Wang