arXiv:2602. 13136v2 Announce Type: replace Abstract: Template-free retrosynthesis methods treat the task as black-box sequence generation, limiting learning efficiency, while semi-template approaches rely on rigid reaction libraries that constrain generalization.
By Chenguang Wang, Zihan Zhou, Lei Bai, Tianshu Yu
MAELLE is a mechanistic reaction prediction framework that models chemical reactions as discrete flow matching over graph-structured electron occupation vectors. It formulates the reactant-to-product mapping as a Continuous-time Markov Chain on electron sites and uses Optimal Transport to generate mechanistically interpretable edit trajectories without elementary step annotations. The method achieves competitive accuracy on the USPTO-480K benchmark, remains robust in out-of-distribution scenarios, and can recover mechanistic pathways that align with known chemistry and predict side products.
By Nguyen Xuan-Vu, Octavian Susanu, Daniel Armstrong, Philippe Schwaller
ReCurveflow is a flow‑matching framework that learns to predict transition state geometries by training on continuously curved reference paths derived from full NEB bands, rather than straight linear paths. It introduces an off‑path correction mechanism that generates corrective velocity fields when the model encounters geometries off the training path, improving resistance to exposure bias and TS prediction accuracy. Across multiple data splits and evaluation metrics, ReCurveflow outperforms seven baselines and produces reaction trajectories whose energy profiles closely follow the reference NEB path, aiding NEB optimization and demonstrating effective corrective behavior.
By Seungheun Baek, Mogan Gim, Jaewoo Kang
arXiv:2607. 17033v1 Announce Type: new Abstract: Forecasting the outcomes of transition-metal-catalyzed reactions is notoriously complex due to the interplay of diverse physical and chemical variables.
By Qiwei Han, Chi Zhou
The paper introduces Align-React, a chemical reaction representation learning framework that incorporates atomic correspondence between reactants and products, an adapter for embedding reaction conditions, and a Reaction-Center-Aware attention mechanism. These components enable the model to capture precise molecular transformations and focus on critical functional groups, leading to improved performance across a variety of organic reaction tasks. The framework outperforms existing architectures on most benchmark datasets.
By Kaipeng Zeng, Xianbin Liu, Yu Zhang, Xiaokang Yang, Yaohui Jin, Yanyan Xu
arXiv:2608. 06259v1 Announce Type: new Abstract: Reaction yield prediction remains challenging because labeled data are scarce and reaction space is both combinatorially large and sparsely populated, limiting the generalization of existing reaction representations.
By Yiting Zheng, Cheng Fang, Anthony Donofrio, Haote Li
The paper introduces Equivariant-Free Transformer-Autoencoded Latent Flow Matching (EF‑TALFM), a two‑stage generative framework that uses a single fixed‑dimensional latent vector to produce variable‑size 3D molecules. The first stage samples the latent vector via flow matching, and the second stage employs an autoregressive Transformer decoder that determines molecule size while generating atom types, coordinates, and chemical states. EF‑TALFM outperforms prior methods on the PCQM4Mv2 benchmark, achieving higher uniqueness, novelty, and computational throughput, and its internal ranking improves the hit rate for target HOMO–LUMO gaps while maintaining novelty.
The paper introduces Top‑K prompting as a training and inference strategy to better capture the diverse, plausible predictions inherent in single‑step retrosynthesis. Using an ultra‑large dataset (CREED‑CCV‑2+USPTO‑XL) of ~45.6 million verified reactions, the authors train the Chemistry Constraint‑Consistent Language Model (C3LM). With fine‑tuning that incorporates ChemCensor‑based and novelty‑oriented rewards, C3LM achieves state‑of‑the‑art performance on the OOD URSA‑expert‑2026 benchmark and demonstrates complementary reaction space exploration compared to conventional models, suggesting benefits for ensemble‑based retrosynthesis systems.
By Bogdan Zagribelnyy, Ivan Ilin, Nikita Bondarev, Maksim Kuznetsov, Mathieu Reymond, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov
PGFS++ is a synthesis‑aware reinforcement learning framework that improves molecular properties while ensuring the resulting molecules can be synthesized and remain structurally similar to the input. It builds on PGFS+ by using trainable embedding lookup tables for reaction templates and second reactants, a more effective scoring function, and a refined RL algorithm. Experiments demonstrate that PGFS++ enhances target properties and preserves high output diversity, overcoming the reward‑hacking failure mode seen in earlier versions.
By Boqiao Zhang, Godbless James, Sai Krishna Gottipati, Andrew Fitzgibbon
The paper introduces MFP, a reaction yield prediction method that uses role-aware Morgan fingerprints. It computes count-based circular fingerprints for each reaction component, aggregates them by chemical role, and combines them with transformation-sensitive difference features into a fixed-length descriptor for a feed-forward neural regressor. On the Suzuki‑Miyaura and Buchwald‑Hartwig benchmarks, MFP achieves R² scores of 0.878 and 0.969 respectively, while training an order of magnitude faster than graph or Transformer-based alternatives.
By Chinmay Mirji, Prashant Shekhar, Foram Madiyar, Hao Peng
The paper presents a closed‑loop molecule generation pipeline that iteratively retrains on new quantum‑chemical simulation data, overcoming limitations of static generative models. This approach produces molecules whose properties extend up to 0.44 standard deviations beyond the training set and improves out‑of‑distribution classification accuracy by 79%. By conditioning on thermodynamic stability during the loop, the method yields a 3.5‑fold increase in the proportion of stable, potentially synthesizable molecules.
By Evan R. Antoniuk, Peggy Li, Nathan Keilbart, Stephen Weitzner, Bhavya Kailkhura, Anna M. Hiszpanski
arXiv:2609.08333v1 Announce Type: cross
Abstract: In molecular discovery, molecule size is coupled to composition, structure, and other target properties. Yet most 3D generators require molecule size...
By Weichi Yao, Cameron Gruich, Bryan R. Goldsmith, Yixin Wang