The paper evaluates 22 state‑of‑the‑art large language models on 12 molecular regression datasets in a zero‑shot setting, comparing their predictions to a molecule‑blind reference derived from the datasets’ labels. It finds that many models retrieve published values rather than truly predicting properties, with significant retrieval concentrated on five datasets and additional flagged instances elsewhere. Increasing the reasoning setting doubles the number of flagged model–dataset pairs, and an in‑context blinding experiment reduces but does not eliminate retrieval, also altering model rankings and increasing errors.
By Matthias Busch, Marius Tacke, Sviatlana V. Lamaka, Mikhail L. Zheludkevich, Christian J. Cyron, Roland C. Aydin, Christian Feiler
arXiv:2606. 05693v1 Announce Type: new Abstract: Large language models (LLMs) have shown promise for molecular property prediction, but their ability to reason over chemical structures remains limited, as molecular representations such as SMILES differ substantially from the natural language on which LLMs are primarily trained.
By Joey Chan, Wonbin Kweon, Ashley Shin, Niharika Bhattacharjee, Pengcheng Jiang, Yue Guo, Jiawei Han
The paper audits 22 frontier language models on 12 molecular property regression benchmarks to assess verbatim retrieval of published values. It finds widespread but benchmark‑specific retrieval, with over 50% of models retrieving exact values on five datasets and isolated occurrences on others. Experiments at different reasoning levels show that higher reasoning increases retrieval flags, and attempts to interrupt retrieval reveal that top models can still recognize transformed SMILES and original labels. Suppressing retrieval reduces prediction error variance, indicating that predictive performance is not solely due to memorized values.
By Matthias Busch, Marius Tacke, Sviatlana V. Lamaka, Mikhail L. Zheludkevich, Christian J. Cyron, Roland C. Aydin, Christian Feiler
arXiv:2609.38744v1 Announce Type: new
Abstract: Predicting molecular properties for compounds that differ structurally from labeled training molecules is important for drug discovery and materials de...
By Jinmo Lee, Dooho Lee, Minho Jeong, Jaemin Yoo
MolSC is a new dataset of 181,000 substituent-level examples that captures how attaching specific substituents to molecular scaffolds changes properties such as bioactivity and physicochemical descriptors. The authors also provide MolSC-Bench, a held‑out benchmark of 1,541 examples that are disjoint from MolSC at scaffold, substituent, and molecule levels. Experiments show that training molecular large language models on MolSC markedly improves their ability to predict substituent contributions, outperforming existing models on a range of downstream chemistry tasks.
By Hyuntae Park, Sooyeon Kim, Jiwon Park, SangKeun Lee
The paper proposes a three‑stage training pipeline that begins with procedural pretraining on abstract, procedurally generated data, followed by molecular pretraining on SMILES, and finally downstream fine‑tuning for molecular property prediction. Experiments show that procedural pretraining improves downstream performance—e.g., a 4.8% error reduction on Lipophilicity—especially when labeled data are scarce, and that the benefit peaks at an intermediate procedural training budget. Analysis indicates that transferable knowledge resides mainly in attention layers, while feed‑forward layers may over‑specialize.
By Moritz Friedemann, Zachary Shinnick, Philip Torr, Bruno Andreis