arXiv:2603. 25857v3 Announce Type: replace Abstract: The capabilities of large language models (LLMs) have expanded beyond natural language processing to scientific prediction tasks, including molecular property prediction.
By Matthias Busch, Marius Tacke, Sviatlana V. Lamaka, Mikhail L. Zheludkevich, Christian J. Cyron, Christian Feiler, Roland C. Aydin
The paper audits 22 frontier language models on 12 molecular property regression benchmarks to assess verbatim retrieval of published values. It finds widespread but benchmark‑specific retrieval, with over 50% of models retrieving exact values on five datasets and isolated occurrences on others. Experiments at different reasoning levels show that higher reasoning increases retrieval flags, and attempts to interrupt retrieval reveal that top models can still recognize transformed SMILES and original labels. Suppressing retrieval reduces prediction error variance, indicating that predictive performance is not solely due to memorized values.
By Matthias Busch, Marius Tacke, Sviatlana V. Lamaka, Mikhail L. Zheludkevich, Christian J. Cyron, Roland C. Aydin, Christian Feiler
MolSC is a new dataset of 181,000 substituent-level examples that captures how attaching specific substituents to molecular scaffolds changes properties such as bioactivity and physicochemical descriptors. The authors also provide MolSC-Bench, a held‑out benchmark of 1,541 examples that are disjoint from MolSC at scaffold, substituent, and molecule levels. Experiments show that training molecular large language models on MolSC markedly improves their ability to predict substituent contributions, outperforming existing models on a range of downstream chemistry tasks.
By Hyuntae Park, Sooyeon Kim, Jiwon Park, SangKeun Lee
arXiv:2606. 05693v1 Announce Type: new Abstract: Large language models (LLMs) have shown promise for molecular property prediction, but their ability to reason over chemical structures remains limited, as molecular representations such as SMILES differ substantially from the natural language on which LLMs are primarily trained.
By Joey Chan, Wonbin Kweon, Ashley Shin, Niharika Bhattacharjee, Pengcheng Jiang, Yue Guo, Jiawei Han
arXiv:2608. 11444v1 Announce Type: cross Abstract: Drug response prediction (DRP) models are an active area of research in pharmacogenomics, with growing potential to accelerate the identification of effective anticancer drugs.
By Vincent Lavelle, Yitan Zhu, Kaitlyn Marlor, Thomas Brettin, Rick Stevens
arXiv:2608. 01734v1 Announce Type: new Abstract: Predicting transcriptomic responses to small-molecule perturbations across cell lines is central to drug discovery, but exhaustive profiling of drug-cell combinations is infeasible.
By Betty Xiong, Jan-Christian Huetter, Gabriele Scalia, Tommaso Biancalani, Sepideh Maleki
The study evaluates four pretrained molecular language models on six virtual libraries covering drug discovery, organic materials, and catalysis. It finds that native embeddings vary widely in performance, while molecular fingerprints remain consistently strong. Fine‑tuning the models on library‑specific data markedly improves sample efficiency, with several adapted encoders outperforming others across all tasks.
By Henrik Wille, Luis-Finley Sch\"utz, Felix Strieth-Kalthoff
arXiv:2603. 25062v2 Announce Type: replace Abstract: Autoregressive molecular models assign probability to molecular serializations even though chemical identity is invariant to serialization.
By Xinyu Wang, Fei Dou, Jinbo Bi, Minghu Song
Large language models (LLMs) are widely applied across chemical tasks, such as molecular property prediction, which underpins drug discovery. Molecular LLMs represent a molecule through several modalities, notably a 1D SMILES sequence or a 2D molecular graph.
The study introduces Malaria-Instruct, a curated instruction-following dataset for malaria virtual screening, and evaluates five open-source large language models (Gemma-2, TxGemma, and LlaSMol-Mistral) against classical machine learning baselines and proprietary models. Fine‑tuned LLMs outperform all baselines, with TxGemma-9B achieving the highest ROC‑AUC (0.731 ± 0.005) and LlaSMol-Mistral-7B delivering the best enrichment factor (EF@1% ≈ 4.99). The results demonstrate that domain‑specific fine‑tuning and chemistry‑aware pretraining are essential for reliable discrimination, positioning fine‑tuned open‑source LLMs as a resource‑efficient alternative for antimalarial virtual screening.
By Marvellous O. Ajala (Magami Open Sciences Initiative), Zainab Ashimiyu-Abdusalam (Magami Open Sciences Initiative), Comfort Adesina (Magami Open Sciences Initiative)
arXiv:2607. 13115v1 Announce Type: new Abstract: Small language models (SLMs) have shown promise for zero-shot molecular property prediction from SMILES strings, yet they often suffer from structural blindness because sequence representations under-specify key graph-topological cues.
By Konstantinos Bougiatiotis, Dimitrios Kelesis, Georgios Paliouras
arXiv:2608. 10480v1 Announce Type: new Abstract: Large language models (LLMs) are widely applied across chemical tasks, such as molecular property prediction, which underpins drug discovery.
By Junwoo Park, Minyoung Shin, Cheol Soon Lee, Sujee Lee