arXiv AI

RxnOptBench: Benchmarking LLMs for Reaction-Condition Optimization in Organic Methodology

arXiv AI
Sep 24

MolDesignBench: Evaluating LLM-based Agent for Scenario-grounded Molecular Design

MolDesignBench is a new benchmark for evaluating large language model (LLM)-based agents in scenario‑grounded molecular design. It contains 2,000 generation and optimization tasks that blend implicit narrative requirements with explicit property and functional‑group constraints, including infeasible cases, and require the use of 17 specialized chemistry tools. Experiments with leading LLMs show low success rates (best ~43%) and highlight failures in implicit‑constraint reasoning, infeasibility detection, and tool usage, underscoring the benchmark’s role in identifying key bottlenecks for future research.

By Yongjun Jeong, Hanbum Ko, Ye Rin Kim, Chanhui Lee, Rodrigo Hormazabal, Jaewan Lee, Sehui Han, Sungbin Lim, Sungwoong Kim
arXiv Machine Learning
Sep 11

Dynamic language model representations for multi-objective reaction optimisation

The paper introduces a method that learns dynamic reaction representations directly from textual descriptions using a fine‑tuned language model coupled with Gaussian process surrogates. This approach enables multi‑objective Bayesian optimisation for chemical reactions, achieving faster convergence than traditional descriptor libraries or one‑hot encodings across nickel‑, palladium‑, and iridium‑catalysed systems. Prospective experiments on a palladium‑catalysed cyanation and an asymmetric hydrogenation produced high‑yield, high‑enantiomeric‑excess conditions after only two rounds of high‑throughput testing, translating directly to gram‑scale synthesis.

By Joshua W. Sin, David Ming Segura, Bojana Rankovi\'c, Siu Lun Chau, Marius D. R. Lutz, Andrea Anelli, Ryan P. Burwood, Kurt P\"untener, Maximilian J. Notheis, Raphael Bigler, Philippe Schwaller
arXiv AI
Aug 20

Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis

The paper introduces Top‑K prompting as a training and inference strategy to better capture the diverse, plausible predictions inherent in single‑step retrosynthesis. Using an ultra‑large dataset (CREED‑CCV‑2+USPTO‑XL) of ~45.6 million verified reactions, the authors train the Chemistry Constraint‑Consistent Language Model (C3LM). With fine‑tuning that incorporates ChemCensor‑based and novelty‑oriented rewards, C3LM achieves state‑of‑the‑art performance on the OOD URSA‑expert‑2026 benchmark and demonstrates complementary reaction space exploration compared to conventional models, suggesting benefits for ensemble‑based retrosynthesis systems.

By Bogdan Zagribelnyy, Ivan Ilin, Nikita Bondarev, Maksim Kuznetsov, Mathieu Reymond, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov
arXiv AI
Sep 18

oMeBench: Towards Robust Benchmarking of LLMs in Organic Mechanism Elucidation and Reasoning

oMeBench is a large-scale, expert-curated benchmark designed to evaluate large language models (LLMs) on organic mechanism reasoning. It contains over 10,000 annotated mechanistic steps, including reaction type labels, intermediate structures, and difficulty ratings, and introduces the oMeS scoring framework to assess logical consistency and chemical structural similarity. Evaluation shows that while current LLMs display promising chemical intuition, they often fail to produce correct and consistent multi-step reasoning, though prompting and fine-tuning can bring smaller models up to the level of closed‑source frontier models.

By Ruiling Xu, Yifan Zhang
arXiv AI
Jun 2

When Single Answer Is Not Enough: Rethinking Single-Step Retrosynthesis Benchmarks for LLMs

arXiv:2602. 03554v2 Announce Type: replace-cross Abstract: Recent progress has expanded the use of large language models (LLMs) in drug discovery, including synthesis planning.

By Bogdan Zagribelnyy, Ivan Ilin, Maksim Kuznetsov, Nikita Bondarev, Mathieu Reymond, Roman Schutski, Thomas MacDougall, Rim Shayakhmetov, Zulfat Miftakhutdinov, Mikolaj Mizera, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov
arXiv AI
4d ago

R-GroundBench: A Diagnostic Benchmark for R-Group Groundingin Markush Molecular Editing

R-GroundBench is a new diagnostic benchmark for evaluating AI models on R‑group grounding in Markush molecular editing, derived from real pharmaceutical patents. It includes a Multiple‑Choice VQA track with varying difficulty and modality splits, as well as an open‑ended Generation track. Experiments show a large performance gap: models score over 90% on easy VQA but drop to 56–66% on hard VQA, and generation exact match stays below 20% (and under 8% with visual input).

By Xin Wang, Zichuan Ying, Xinna Lin, Junqi Zhang, Hanyi Xiong, Tianyu Gao, Hairong Zhang, Qixiang Hua, Botian Shi, Zhenhailong Wang, Kaicheng Yu
arXiv Machine Learning
Sep 7

Training Large Language Models for Small-Molecule Design with Synthetic Task Scaling

The paper explores how large language models (LLMs) can be trained for small-molecule drug design by using synthetic tasks that are cheaper to evaluate. By employing a curriculum that gradually increases task difficulty, the authors demonstrate that LLMs can learn design strategies that outperform larger models on structure-based lead optimization. This approach shows that scaling post‑training with synthetic tasks can effectively adapt LLMs to high‑cost experimental scenarios that are otherwise infeasible to train on directly.

By Frank Hu, Shriram Chennakesavalu, Zichen Wang, Patricia Suriana, Bodhi Vani, Kirill Shmilovich, Kangway Chuang, Colin Grambow