arXiv:2607. 03007v1 Announce Type: cross Abstract: Recent advances in molecular large language models have led to strong performance on molecular understanding and generation tasks, yet these gains often come without reliable structural grounding.
By Wenda Wang, Jinjia Feng, Zhewei Wei
WOMBAT is a benchmark comprising 14 whitebox graph neural networks (GNNs) whose message‑passing weights are manually set to detect specific SMARTS motifs. Each model’s decision rule is explicitly known, providing a ground truth for attribution that allows researchers to identify and study errors in post‑hoc explainers such as GNNExplainer, PGExplainer, and Integrated Gradients. The authors validate the models on millions of PubChem molecules, demonstrate how Integrated Gradients can be misled to spread attribution, and release the dataset, models, and evaluation code for future XAI tool development.
By Dominik Matuszek, Bartosz Zieli\'nski, Tomasz Danel, Dawid Rymarczyk
Large language models (LLMs) are widely applied across chemical tasks, such as molecular property prediction, which underpins drug discovery. Molecular LLMs represent a molecule through several modalities, notably a 1D SMILES sequence or a 2D molecular graph.
arXiv:2606. 05693v1 Announce Type: new Abstract: Large language models (LLMs) have shown promise for molecular property prediction, but their ability to reason over chemical structures remains limited, as molecular representations such as SMILES differ substantially from the natural language on which LLMs are primarily trained.
By Joey Chan, Wonbin Kweon, Ashley Shin, Niharika Bhattacharjee, Pengcheng Jiang, Yue Guo, Jiawei Han
arXiv:2608. 10480v1 Announce Type: new Abstract: Large language models (LLMs) are widely applied across chemical tasks, such as molecular property prediction, which underpins drug discovery.
By Junwoo Park, Minyoung Shin, Cheol Soon Lee, Sujee Lee
arXiv:2606. 12113v1 Announce Type: cross Abstract: Transformer-based language models for SMILES strings suffer from a locality gap: standard character-level tokenization fragments chemically meaningful motifs, forcing models to repeatedly learn local syntax at the expense of long-range dependencies.
By Xinni Zhang, Zijing Liu, He Cao, Yu Li, Irwin King