oMeBench is a large-scale, expert-curated benchmark designed to evaluate large language models (LLMs) on organic mechanism reasoning. It contains over 10,000 annotated mechanistic steps, including reaction type labels, intermediate structures, and difficulty ratings, and introduces the oMeS scoring framework to assess logical consistency and chemical structural similarity. Evaluation shows that while current LLMs display promising chemical intuition, they often fail to produce correct and consistent multi-step reasoning, though prompting and fine-tuning can bring smaller models up to the level of closed‑source frontier models.
By Ruiling Xu, Yifan Zhang
The paper introduces Top‑K prompting as a training and inference strategy to better capture the diverse, plausible predictions inherent in single‑step retrosynthesis. Using an ultra‑large dataset (CREED‑CCV‑2+USPTO‑XL) of ~45.6 million verified reactions, the authors train the Chemistry Constraint‑Consistent Language Model (C3LM). With fine‑tuning that incorporates ChemCensor‑based and novelty‑oriented rewards, C3LM achieves state‑of‑the‑art performance on the OOD URSA‑expert‑2026 benchmark and demonstrates complementary reaction space exploration compared to conventional models, suggesting benefits for ensemble‑based retrosynthesis systems.
By Bogdan Zagribelnyy, Ivan Ilin, Nikita Bondarev, Maksim Kuznetsov, Mathieu Reymond, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov
arXiv:2508. 10967v3 Announce Type: replace-cross Abstract: Retrosynthesis prediction aims to infer the reactant molecules based on a given product molecule, which is a fundamental task in chemical synthesis.
By Xinyi Li, Sai Wang, Yutian Lin, Yu Wu
arXiv:2603. 12666v2 Announce Type: replace-cross Abstract: Retrosynthesis prediction aims to identify reactants that can synthesize a given product molecule.
By Hanbum Ko, Chanhui Lee, Ye Rin Kim, Rodrigo Hormazabal, Sehui Han, Sungbin Lim, Sungwoong Kim
arXiv:2507.17448v2 Announce Type: replace-cross
Abstract: Retrosynthetic planning is a cornerstone of organic synthesis and drug discovery. Yet existing AI methods often rely on pattern matching rath...
By Situo Zhang, Hanqi Li, Lu Chen, Zihan Zhao, Xuanze Lin, Zichen Zhu, Danyu Luo, Bo Chen, Xin Chen, Kai Yu
The paper introduces a multitask large reasoning model for molecular science that incorporates chemical knowledge via a multispecialist architecture, chain-of-thought supervision, and molecule-informed reinforcement learning. It coordinates prediction and inference specialists across ten molecular tasks—including description, generation, nomenclature translation, property prediction, and reaction prediction—using task-conditioned routing. The model surpasses more than 20 general-purpose and molecular large language models, improving aggregate performance by 50.3% and outperforming leading multitask baselines on most tasks, while maintaining interpretable chemical inference and demonstrating a workflow for CNS candidate generation and retrosynthetic planning.
By Pengfei Liu, Shuang Ge, Xiaobo Wang, Xin Liu, Jun Tao, Yan Li, Chao Liu, Ling Chen, Zhixiang Ren