arXiv Machine Learning

RxnCLF: Contrastive Transformation-Aware Reaction Foundation Model for Improved Reactivity Prediction

arXiv:2608. 06259v1 Announce Type: new Abstract: Reaction yield prediction remains challenging because labeled data are scarce and reaction space is both combinatorially large and sparsely populated, limiting the generalization of existing reaction representations.

arXiv Machine Learning
Aug 27

A General-Purpose Framework for Chemical Reaction Representation with Atomic Correspondence and Flexible Condition Adaptation

The paper introduces Align-React, a chemical reaction representation learning framework that incorporates atomic correspondence between reactants and products, an adapter for embedding reaction conditions, and a Reaction-Center-Aware attention mechanism. These components enable the model to capture precise molecular transformations and focus on critical functional groups, leading to improved performance across a variety of organic reaction tasks. The framework outperforms existing architectures on most benchmark datasets.

By Kaipeng Zeng, Xianbin Liu, Yu Zhang, Xiaokang Yang, Yaohui Jin, Yanyan Xu
arXiv Machine Learning
Sep 22

Role-Aware Morgan Fingerprints for Reaction Yield Prediction

The paper introduces MFP, a reaction yield prediction method that uses role-aware Morgan fingerprints. It computes count-based circular fingerprints for each reaction component, aggregates them by chemical role, and combines them with transformation-sensitive difference features into a fixed-length descriptor for a feed-forward neural regressor. On the Suzuki‑Miyaura and Buchwald‑Hartwig benchmarks, MFP achieves R² scores of 0.878 and 0.969 respectively, while training an order of magnitude faster than graph or Transformer-based alternatives.

By Chinmay Mirji, Prashant Shekhar, Foram Madiyar, Hao Peng
arXiv Machine Learning
Aug 12

Order Matters in Retrosynthesis: Structure-aware Generation via Reaction-Center-Guided Discrete Flow Matching

arXiv:2602. 13136v2 Announce Type: replace Abstract: Template-free retrosynthesis methods treat the task as black-box sequence generation, limiting learning efficiency, while semi-template approaches rely on rigid reaction libraries that constrain generalization.

By Chenguang Wang, Zihan Zhou, Lei Bai, Tianshu Yu
arXiv AI
Aug 17

Reaction-Transformation-Aware Flow Matching for Generalizable Transition State Generation

arXiv:2608. 14076v1 Announce Type: cross Abstract: Transition-state (TS) structures define the energetic barriers and mechanistic pathways of elementary chemical reactions, yet their identification remains computationally demanding because conventional saddle-point searches require expensive quantum-mechanical calculations.

By Kaipeng Zeng, Wenxi Zhai, Shengrui Xu, Jie Zhao, Bowen Li, Shiyue Wang, Junchi Yan, Tong Zhu
arXiv AI
Aug 20

PGFS++: Molecular Property Improvement under Synthesis and Diversity Constraints

PGFS++ is a synthesis‑aware reinforcement learning framework that improves molecular properties while ensuring the resulting molecules can be synthesized and remain structurally similar to the input. It builds on PGFS+ by using trainable embedding lookup tables for reaction templates and second reactants, a more effective scoring function, and a refined RL algorithm. Experiments demonstrate that PGFS++ enhances target properties and preserves high output diversity, overcoming the reward‑hacking failure mode seen in earlier versions.

By Boqiao Zhang, Godbless James, Sai Krishna Gottipati, Andrew Fitzgibbon
Hugging Face Trending Papers
Aug 19

PGFS++: Molecular Property Improvement under Synthesis and Diversity Constraints

PGFS++ is a synthesis‑aware reinforcement learning framework that improves molecular properties such as drug‑likeness or binding affinity while ensuring the resulting molecules can be synthesized and remain structurally similar to the input. It builds on PGFS+ by using trainable embedding lookup tables for reaction templates and second reactants, a more effective scoring function, and a refined RL algorithm. The method addresses a reward‑hacking failure mode by treating each input molecule as the start of a forward‑synthesis trajectory, applying learned reaction templates with in‑stock building blocks, and producing diverse, high‑quality outputs with explicit synthesis routes.

arXiv AI
Jul 10

Reaction-network reasoning with frontier models for experimentally confirmed catalyst-selectivity hypotheses

arXiv:2607. 08003v1 Announce Type: cross Abstract: Catalysts are essential for sustainable chemical manufacturing, yet discovering novel architectures remains a bottleneck dominated by trial-and-error experimentation and computationally intensive screening.

By Sutanay Choudhury, Anwesha Banerjee, Udishnu Sanyal, Jorin Dawidowicz, Chiezugolum Ijeoma Odilinye, Jesun Firoz, Liney Arnadottir, Simone Raugei, Johannes Lercher, Arnab Dutta
arXiv Machine Learning
Sep 11

Dynamic language model representations for multi-objective reaction optimisation

The paper introduces a method that learns dynamic reaction representations directly from textual descriptions using a fine‑tuned language model coupled with Gaussian process surrogates. This approach enables multi‑objective Bayesian optimisation for chemical reactions, achieving faster convergence than traditional descriptor libraries or one‑hot encodings across nickel‑, palladium‑, and iridium‑catalysed systems. Prospective experiments on a palladium‑catalysed cyanation and an asymmetric hydrogenation produced high‑yield, high‑enantiomeric‑excess conditions after only two rounds of high‑throughput testing, translating directly to gram‑scale synthesis.

By Joshua W. Sin, David Ming Segura, Bojana Rankovi\'c, Siu Lun Chau, Marius D. R. Lutz, Andrea Anelli, Ryan P. Burwood, Kurt P\"untener, Maximilian J. Notheis, Raphael Bigler, Philippe Schwaller
arXiv AI
Aug 28

Mechanistic Reaction Prediction via Discrete Flow Matching on Graph-Structured Electron Occupation

MAELLE is a mechanistic reaction prediction framework that models chemical reactions as discrete flow matching over graph-structured electron occupation vectors. It formulates the reactant-to-product mapping as a Continuous-time Markov Chain on electron sites and uses Optimal Transport to generate mechanistically interpretable edit trajectories without elementary step annotations. The method achieves competitive accuracy on the USPTO-480K benchmark, remains robust in out-of-distribution scenarios, and can recover mechanistic pathways that align with known chemistry and predict side products.

By Nguyen Xuan-Vu, Octavian Susanu, Daniel Armstrong, Philippe Schwaller