arXiv:2606. 05693v1 Announce Type: new Abstract: Large language models (LLMs) have shown promise for molecular property prediction, but their ability to reason over chemical structures remains limited, as molecular representations such as SMILES differ substantially from the natural language on which LLMs are primarily trained.
By Joey Chan, Wonbin Kweon, Ashley Shin, Niharika Bhattacharjee, Pengcheng Jiang, Yue Guo, Jiawei Han
arXiv:2606. 03057v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for molecular tasks, but it remains unclear which molecular representation to use.
By Arun Raja, Garrett M. Morris, Kian Ming A. Chai
arXiv:2606. 03660v1 Announce Type: new Abstract: Large language models are increasingly used as chemistry assistants, yet most chemistry benchmarks still score only final answers.
By Hongyu Guo, Hao Li, He Cao, Gongbo Zhang, Li Yuan
R-GroundBench is a new diagnostic benchmark for evaluating AI models on R‑group grounding in Markush molecular editing, derived from real pharmaceutical patents. It includes a Multiple‑Choice VQA track with varying difficulty and modality splits, as well as an open‑ended Generation track. Experiments show a large performance gap: models score over 90% on easy VQA but drop to 56–66% on hard VQA, and generation exact match stays below 20% (and under 8% with visual input).
By Xin Wang, Zichuan Ying, Xinna Lin, Junqi Zhang, Hanyi Xiong, Tianyu Gao, Hairong Zhang, Qixiang Hua, Botian Shi, Zhenhailong Wang, Kaicheng Yu
ChemVTS-Bench is a domain-authentic benchmark that evaluates Visual‑Textual‑Symbolic reasoning in multimodal large language models for chemistry. It presents diverse chemical problems—organic molecules, inorganic materials, and 3D crystal structures—in three input modes: visual-only, visual‑text hybrid, and SMILES-based symbolic. The benchmark includes an automated agent workflow for inference, answer verification, and failure diagnosis, and shows that visual-only inputs and structural chemistry remain challenging for current models.
By Zhiyuan Huang, Baichuan Yang, Zikun He, Yanhong Wu, Fang Hongyu, Zhenhe Liu, Lin Dongsheng, Bing Su
The study evaluates four pretrained molecular language models on six virtual libraries covering drug discovery, organic materials, and catalysis. It finds that native embeddings vary widely in performance, while molecular fingerprints remain consistently strong. Fine‑tuning the models on library‑specific data markedly improves sample efficiency, with several adapted encoders outperforming others across all tasks.
By Henrik Wille, Luis-Finley Sch\"utz, Felix Strieth-Kalthoff