RxnOptBench: Benchmarking LLMs for Reaction-Condition Optimization in Organic Methodology
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2606. 03660v1 Announce Type: new Abstract: Large language models are increasingly used as chemistry assistants, yet most chemistry benchmarks still score only final answers.
MolDesignBench is a new benchmark for evaluating large language model (LLM)-based agents in scenario‑grounded molecular design. It contains 2,000 generation and optimization tasks that blend implicit narrative requirements with explicit property and functional‑group constraints, including infeasible cases, and require the use of 17 specialized chemistry tools. Experiments with leading LLMs show low success rates (best ~43%) and highlight failures in implicit‑constraint reasoning, infeasibility detection, and tool usage, underscoring the benchmark’s role in identifying key bottlenecks for future research.
The paper introduces a method that learns dynamic reaction representations directly from textual descriptions using a fine‑tuned language model coupled with Gaussian process surrogates. This approach enables multi‑objective Bayesian optimisation for chemical reactions, achieving faster convergence than traditional descriptor libraries or one‑hot encodings across nickel‑, palladium‑, and iridium‑catalysed systems. Prospective experiments on a palladium‑catalysed cyanation and an asymmetric hydrogenation produced high‑yield, high‑enantiomeric‑excess conditions after only two rounds of high‑throughput testing, translating directly to gram‑scale synthesis.
The paper introduces Top‑K prompting as a training and inference strategy to better capture the diverse, plausible predictions inherent in single‑step retrosynthesis. Using an ultra‑large dataset (CREED‑CCV‑2+USPTO‑XL) of ~45.6 million verified reactions, the authors train the Chemistry Constraint‑Consistent Language Model (C3LM). With fine‑tuning that incorporates ChemCensor‑based and novelty‑oriented rewards, C3LM achieves state‑of‑the‑art performance on the OOD URSA‑expert‑2026 benchmark and demonstrates complementary reaction space exploration compared to conventional models, suggesting benefits for ensemble‑based retrosynthesis systems.
arXiv:2604.07669v3 Announce Type: replace-cross Abstract: Synthesizable molecular optimization seeks to improve target properties while ensuring that molecular modifications follow feasible synthetic...
oMeBench is a large-scale, expert-curated benchmark designed to evaluate large language models (LLMs) on organic mechanism reasoning. It contains over 10,000 annotated mechanistic steps, including reaction type labels, intermediate structures, and difficulty ratings, and introduces the oMeS scoring framework to assess logical consistency and chemical structural similarity. Evaluation shows that while current LLMs display promising chemical intuition, they often fail to produce correct and consistent multi-step reasoning, though prompting and fine-tuning can bring smaller models up to the level of closed‑source frontier models.