oMeBench is a large-scale, expert-curated benchmark designed to evaluate large language models (LLMs) on organic mechanism reasoning. It contains over 10,000 annotated mechanistic steps, including reaction type labels, intermediate structures, and difficulty ratings, and introduces the oMeS scoring framework to assess logical consistency and chemical structural similarity. Evaluation shows that while current LLMs display promising chemical intuition, they often fail to produce correct and consistent multi-step reasoning, though prompting and fine-tuning can bring smaller models up to the level of closed‑source frontier models.
By Ruiling Xu, Yifan Zhang
MolDesignBench is a new benchmark for evaluating large language model (LLM)-based agents in scenario‑grounded molecular design. It contains 2,000 generation and optimization tasks that blend implicit narrative requirements with explicit property and functional‑group constraints, including infeasible cases, and require the use of 17 specialized chemistry tools. Experiments with leading LLMs show low success rates (best ~43%) and highlight failures in implicit‑constraint reasoning, infeasibility detection, and tool usage, underscoring the benchmark’s role in identifying key bottlenecks for future research.
By Yongjun Jeong, Hanbum Ko, Ye Rin Kim, Chanhui Lee, Rodrigo Hormazabal, Jaewan Lee, Sehui Han, Sungbin Lim, Sungwoong Kim
arXiv:2606. 03057v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for molecular tasks, but it remains unclear which molecular representation to use.
By Arun Raja, Garrett M. Morris, Kian Ming A. Chai
arXiv:2608. 02595v1 Announce Type: new Abstract: Language models are playing an increasingly important role in laboratory science, performing tasks such as experiment planning, execution, and post-hoc analysis.
By Brandon Wang, Andrei S. Tyrin, Daniil A. Boiko
The paper introduces a multitask large reasoning model for molecular science that incorporates chemical knowledge via a multispecialist architecture, chain-of-thought supervision, and molecule-informed reinforcement learning. It coordinates prediction and inference specialists across ten molecular tasks—including description, generation, nomenclature translation, property prediction, and reaction prediction—using task-conditioned routing. The model surpasses more than 20 general-purpose and molecular large language models, improving aggregate performance by 50.3% and outperforming leading multitask baselines on most tasks, while maintaining interpretable chemical inference and demonstrating a workflow for CNS candidate generation and retrosynthetic planning.
By Pengfei Liu, Shuang Ge, Xiaobo Wang, Xin Liu, Jun Tao, Yan Li, Chao Liu, Ling Chen, Zhixiang Ren
ChemVTS-Bench is a domain-authentic benchmark that evaluates Visual‑Textual‑Symbolic reasoning in multimodal large language models for chemistry. It presents diverse chemical problems—organic molecules, inorganic materials, and 3D crystal structures—in three input modes: visual-only, visual‑text hybrid, and SMILES-based symbolic. The benchmark includes an automated agent workflow for inference, answer verification, and failure diagnosis, and shows that visual-only inputs and structural chemistry remain challenging for current models.
By Zhiyuan Huang, Baichuan Yang, Zikun He, Yanhong Wu, Fang Hongyu, Zhenhe Liu, Lin Dongsheng, Bing Su
arXiv:2607. 19935v1 Announce Type: new Abstract: Large metal-organic framework (MOF) databases support simulation, screening, and machine learning through crystallographic information files (CIFs).
By Yu Liu, Zhiwei Yang, Diandian Guo, Kun Peng, Fangfang Yuan, Cong Cao, Chaozhuo Li, Zhiyuan Ma, Yanbing Liu, Guobin Zhao
arXiv:2508. 10967v3 Announce Type: replace-cross Abstract: Retrosynthesis prediction aims to infer the reactant molecules based on a given product molecule, which is a fundamental task in chemical synthesis.
By Xinyi Li, Sai Wang, Yutian Lin, Yu Wu
Chemical reasoning language models are expected to produce faithful chain-of-thought (CoT) explanations when answering chemistry tasks, but across four model families and twelve tasks, hallucinations are widespread and largely independent of answer correctness. Attribution analyses reveal that these models use a shared scratchpad function: Chem‑R and ether‑0 rely on fragmented SMILES drafts, while ChemDFM‑R emphasizes scaffold, positional, and naming cues. Perturbing Chem‑R’s SMILES sketches degrades generation, indicating that structural drafts can be causally load‑bearing even when verbal structural claims are largely inert.
By Jiatong Li, Yuxuan Ren, Weida Wang, Xiaoyong Wei, Yatao Bian
R-GroundBench is a new diagnostic benchmark for evaluating AI models on R‑group grounding in Markush molecular editing, derived from real pharmaceutical patents. It includes a Multiple‑Choice VQA track with varying difficulty and modality splits, as well as an open‑ended Generation track. Experiments show a large performance gap: models score over 90% on easy VQA but drop to 56–66% on hard VQA, and generation exact match stays below 20% (and under 8% with visual input).
By Xin Wang, Zichuan Ying, Xinna Lin, Junqi Zhang, Hanyi Xiong, Tianyu Gao, Hairong Zhang, Qixiang Hua, Botian Shi, Zhenhailong Wang, Kaicheng Yu
arXiv:2606. 11256v1 Announce Type: cross Abstract: Designing molecules with target properties is most useful when candidate structures are accompanied by feasible synthetic routes.
By C\'esar Ojeda, Darius A. Faroughy, Maryam Karimi, Payam Zarrintaj, Mir Mehdi Seyedebrahimi, Mart\'in Carballo-Pacheco
arXiv:2607. 12771v1 Announce Type: new Abstract: Reaction mechanisms consist of the step-by-step sequences of elementary reactions that explain chemical transformations.
By Xingyu Dang, Haocheng Tang, Junmei Wang, Yanjun Li