arXiv:2606. 03660v1 Announce Type: new Abstract: Large language models are increasingly used as chemistry assistants, yet most chemistry benchmarks still score only final answers.
By Hongyu Guo, Hao Li, He Cao, Gongbo Zhang, Li Yuan
MolDesignBench is a new benchmark for evaluating large language model (LLM)-based agents in scenario‑grounded molecular design. It contains 2,000 generation and optimization tasks that blend implicit narrative requirements with explicit property and functional‑group constraints, including infeasible cases, and require the use of 17 specialized chemistry tools. Experiments with leading LLMs show low success rates (best ~43%) and highlight failures in implicit‑constraint reasoning, infeasibility detection, and tool usage, underscoring the benchmark’s role in identifying key bottlenecks for future research.
By Yongjun Jeong, Hanbum Ko, Ye Rin Kim, Chanhui Lee, Rodrigo Hormazabal, Jaewan Lee, Sehui Han, Sungbin Lim, Sungwoong Kim
The paper introduces a method that learns dynamic reaction representations directly from textual descriptions using a fine‑tuned language model coupled with Gaussian process surrogates. This approach enables multi‑objective Bayesian optimisation for chemical reactions, achieving faster convergence than traditional descriptor libraries or one‑hot encodings across nickel‑, palladium‑, and iridium‑catalysed systems. Prospective experiments on a palladium‑catalysed cyanation and an asymmetric hydrogenation produced high‑yield, high‑enantiomeric‑excess conditions after only two rounds of high‑throughput testing, translating directly to gram‑scale synthesis.
By Joshua W. Sin, David Ming Segura, Bojana Rankovi\'c, Siu Lun Chau, Marius D. R. Lutz, Andrea Anelli, Ryan P. Burwood, Kurt P\"untener, Maximilian J. Notheis, Raphael Bigler, Philippe Schwaller
The paper introduces Top‑K prompting as a training and inference strategy to better capture the diverse, plausible predictions inherent in single‑step retrosynthesis. Using an ultra‑large dataset (CREED‑CCV‑2+USPTO‑XL) of ~45.6 million verified reactions, the authors train the Chemistry Constraint‑Consistent Language Model (C3LM). With fine‑tuning that incorporates ChemCensor‑based and novelty‑oriented rewards, C3LM achieves state‑of‑the‑art performance on the OOD URSA‑expert‑2026 benchmark and demonstrates complementary reaction space exploration compared to conventional models, suggesting benefits for ensemble‑based retrosynthesis systems.
By Bogdan Zagribelnyy, Ivan Ilin, Nikita Bondarev, Maksim Kuznetsov, Mathieu Reymond, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov
arXiv:2604.07669v3 Announce Type: replace-cross
Abstract: Synthesizable molecular optimization seeks to improve target properties while ensuring that molecular modifications follow feasible synthetic...
By Tao Li, Kaiyuan Hou, Tuan Vinh, Fanglei Xue, Monika Raj, Zhichun Guo, Carl Yang
oMeBench is a large-scale, expert-curated benchmark designed to evaluate large language models (LLMs) on organic mechanism reasoning. It contains over 10,000 annotated mechanistic steps, including reaction type labels, intermediate structures, and difficulty ratings, and introduces the oMeS scoring framework to assess logical consistency and chemical structural similarity. Evaluation shows that while current LLMs display promising chemical intuition, they often fail to produce correct and consistent multi-step reasoning, though prompting and fine-tuning can bring smaller models up to the level of closed‑source frontier models.
By Ruiling Xu, Yifan Zhang
arXiv:2608.22967v1 Announce Type: new
Abstract: Practical molecular inverse design is rarely a one-shot generation problem; it often takes the form of closed-loop candidate-pool enrichment, where und...
By Yaoyao Xu, Xinjian Zhao, Xiaozhuang Song, Lei Bai, Tianshu Yu
arXiv:2606. 11256v1 Announce Type: cross Abstract: Designing molecules with target properties is most useful when candidate structures are accompanied by feasible synthetic routes.
By C\'esar Ojeda, Darius A. Faroughy, Maryam Karimi, Payam Zarrintaj, Mir Mehdi Seyedebrahimi, Mart\'in Carballo-Pacheco
arXiv:2602. 03554v2 Announce Type: replace-cross Abstract: Recent progress has expanded the use of large language models (LLMs) in drug discovery, including synthesis planning.
By Bogdan Zagribelnyy, Ivan Ilin, Maksim Kuznetsov, Nikita Bondarev, Mathieu Reymond, Roman Schutski, Thomas MacDougall, Rim Shayakhmetov, Zulfat Miftakhutdinov, Mikolaj Mizera, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov
arXiv:2602.00663v3 Announce Type: replace
Abstract: Optimizing molecules to achieve desired properties is a central bottleneck across the chemical sciences, particularly in the pharmaceutical industr...
By Fabian P. Kr\"uger, Andrea Hunklinger, Adrian Wolny, Tim J. Adler, Igor Tetko, Santiago David Villalba
R-GroundBench is a new diagnostic benchmark for evaluating AI models on R‑group grounding in Markush molecular editing, derived from real pharmaceutical patents. It includes a Multiple‑Choice VQA track with varying difficulty and modality splits, as well as an open‑ended Generation track. Experiments show a large performance gap: models score over 90% on easy VQA but drop to 56–66% on hard VQA, and generation exact match stays below 20% (and under 8% with visual input).
By Xin Wang, Zichuan Ying, Xinna Lin, Junqi Zhang, Hanyi Xiong, Tianyu Gao, Hairong Zhang, Qixiang Hua, Botian Shi, Zhenhailong Wang, Kaicheng Yu
The paper explores how large language models (LLMs) can be trained for small-molecule drug design by using synthetic tasks that are cheaper to evaluate. By employing a curriculum that gradually increases task difficulty, the authors demonstrate that LLMs can learn design strategies that outperform larger models on structure-based lead optimization. This approach shows that scaling post‑training with synthetic tasks can effectively adapt LLMs to high‑cost experimental scenarios that are otherwise infeasible to train on directly.
By Frank Hu, Shriram Chennakesavalu, Zichen Wang, Patricia Suriana, Bodhi Vani, Kirill Shmilovich, Kangway Chuang, Colin Grambow