oMeBench is a large-scale, expert-curated benchmark designed to evaluate large language models (LLMs) on organic mechanism reasoning. It contains over 10,000 annotated mechanistic steps, including reaction type labels, intermediate structures, and difficulty ratings, and introduces the oMeS scoring framework to assess logical consistency and chemical structural similarity. Evaluation shows that while current LLMs display promising chemical intuition, they often fail to produce correct and consistent multi-step reasoning, though prompting and fine-tuning can bring smaller models up to the level of closed‑source frontier models.
By Ruiling Xu, Yifan Zhang
The paper introduces Top‑K prompting as a training and inference strategy to better capture the diverse, plausible predictions inherent in single‑step retrosynthesis. Using an ultra‑large dataset (CREED‑CCV‑2+USPTO‑XL) of ~45.6 million verified reactions, the authors train the Chemistry Constraint‑Consistent Language Model (C3LM). With fine‑tuning that incorporates ChemCensor‑based and novelty‑oriented rewards, C3LM achieves state‑of‑the‑art performance on the OOD URSA‑expert‑2026 benchmark and demonstrates complementary reaction space exploration compared to conventional models, suggesting benefits for ensemble‑based retrosynthesis systems.
By Bogdan Zagribelnyy, Ivan Ilin, Nikita Bondarev, Maksim Kuznetsov, Mathieu Reymond, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov
arXiv:2508. 10967v3 Announce Type: replace-cross Abstract: Retrosynthesis prediction aims to infer the reactant molecules based on a given product molecule, which is a fundamental task in chemical synthesis.
By Xinyi Li, Sai Wang, Yutian Lin, Yu Wu
arXiv:2603. 12666v2 Announce Type: replace-cross Abstract: Retrosynthesis prediction aims to identify reactants that can synthesize a given product molecule.
By Hanbum Ko, Chanhui Lee, Ye Rin Kim, Rodrigo Hormazabal, Sehui Han, Sungbin Lim, Sungwoong Kim
arXiv:2507.17448v2 Announce Type: replace-cross
Abstract: Retrosynthetic planning is a cornerstone of organic synthesis and drug discovery. Yet existing AI methods often rely on pattern matching rath...
By Situo Zhang, Hanqi Li, Lu Chen, Zihan Zhao, Xuanze Lin, Zichen Zhu, Danyu Luo, Bo Chen, Xin Chen, Kai Yu
The paper introduces a multitask large reasoning model for molecular science that incorporates chemical knowledge via a multispecialist architecture, chain-of-thought supervision, and molecule-informed reinforcement learning. It coordinates prediction and inference specialists across ten molecular tasks—including description, generation, nomenclature translation, property prediction, and reaction prediction—using task-conditioned routing. The model surpasses more than 20 general-purpose and molecular large language models, improving aggregate performance by 50.3% and outperforming leading multitask baselines on most tasks, while maintaining interpretable chemical inference and demonstrating a workflow for CNS candidate generation and retrosynthetic planning.
By Pengfei Liu, Shuang Ge, Xiaobo Wang, Xin Liu, Jun Tao, Yan Li, Chao Liu, Ling Chen, Zhixiang Ren
The paper introduces Align-React, a chemical reaction representation learning framework that incorporates atomic correspondence between reactants and products, an adapter for embedding reaction conditions, and a Reaction-Center-Aware attention mechanism. These components enable the model to capture precise molecular transformations and focus on critical functional groups, leading to improved performance across a variety of organic reaction tasks. The framework outperforms existing architectures on most benchmark datasets.
By Kaipeng Zeng, Xianbin Liu, Yu Zhang, Xiaokang Yang, Yaohui Jin, Yanyan Xu
arXiv:2602. 03554v2 Announce Type: replace-cross Abstract: Recent progress has expanded the use of large language models (LLMs) in drug discovery, including synthesis planning.
By Bogdan Zagribelnyy, Ivan Ilin, Maksim Kuznetsov, Nikita Bondarev, Mathieu Reymond, Roman Schutski, Thomas MacDougall, Rim Shayakhmetov, Zulfat Miftakhutdinov, Mikolaj Mizera, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov
ReactBench is a benchmark designed to evaluate the structural reasoning abilities of multimodal large language models (MLLMs) using chemical reaction diagrams. The dataset contains 1,618 expert‑annotated question‑answer pairs that test reasoning across four hierarchical task dimensions, from simple endpoint counting to complex topological analysis. Evaluation of 24 MLLMs shows a performance gap of more than 30% between anchor‑based tasks and holistic structural reasoning tasks, indicating that current models struggle with reasoning over branching, converging, and cyclic structures.
By Qiang Xu, Shengyuan Bai, Yu Wang, He Cao, Leqing Chen, Yuanyuan Liu, Bin Feng, Zijing Liu, Yu Li
arXiv:2608. 16111v1 Announce Type: cross Abstract: Retrosynthesis is a cornerstone of drug discovery and organic synthesis.
By Mianzhi Liu, Fan Xiao, Zhiliang Yu, Huayang Huang, Yuke Li, Yi Yang, Wenbo Liu, Yu Wu
arXiv:2602. 13136v2 Announce Type: replace Abstract: Template-free retrosynthesis methods treat the task as black-box sequence generation, limiting learning efficiency, while semi-template approaches rely on rigid reaction libraries that constrain generalization.
By Chenguang Wang, Zihan Zhou, Lei Bai, Tianshu Yu
arXiv:2607. 01061v1 Announce Type: new Abstract: Computer-assisted synthesis planning breaks target molecules into accessible precursors using large libraries of reaction rules that assign each transformation a deterministic, interpretable label.
By Daniel Armstrong, Maarten Dobbelaere, Valentas Olikauskas, Helena Avila, Octavian Susanu, J\'er\^ome Waser, Philippe Schwaller